Infrastructure Engineer (GPU & Compute)

lightningai

Deutschland

Hybrid

EUR 155.000 - 189.000

Vollzeit

Vor 6 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Erhalte eine Antwort von diesem Arbeitgeber — ein Lebenslauf und ein Anschreiben, die genau auf die Eigenschaften eingehen, die gesucht werden.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Discretionary bonus
Equity
Comprehensive medical coverage
Unlimited PTO
Sabbatical after 4 years
Hybrid work model

Zusammenfassung

lightningai is seeking a senior infrastructure engineer to bring up, validate, and operate large-scale bare-metal GPU environments. You will own diagnostics, tooling, and automation that keep clusters ready for demanding AI/ML and HPC workloads.

This role blends hardware, systems software, and automation, requiring strong Linux production experience, hands-on GPU tooling, and Python scripting for provisioning and validation. Hybrid remote work with US hubs is available.

Qualifikationen

  • 5+ years in infrastructure engineering, systems engineering, or a closely related role.
  • Strong Linux systems experience in production environments.
  • Hands-on experience with GPU-enabled systems and diagnostic tooling such as NVIDIA DCGM.
  • Familiarity with bare-metal provisioning and system bring-up workflows.
  • Proficiency in Python or comparable scripting/programming languages for automation.
  • Ability to debug complex issues across hardware, OS, GPUs, and system software.

Aufgaben

  • Own and evolve image management, deployment, and validation pipelines across bare-metal infrastructure, including firmware, driver, and OS qualification for GPU-enabled systems
  • Operate and maintain test clusters used for bring-up and diagnostics, supporting hardware qualification efforts for next-generation platforms
  • Diagnose and resolve complex issues spanning GPUs, drivers, OS, and underlying hardware; analyze performance using tools such as NVIDIA DCGM
  • Build Python-based automation for provisioning, validation, and system bring-up, improving reliability, repeatability, and scalability
  • Manage Linux-based production and validation environments, including virtualization and PXE/image-based bare-metal provisioning workflows
  • Partner with infrastructure, hardware, and data center teams, plus platform and ML stakeholders, to ensure systems meet workload requirements and contribute to provisioning and lifecycle best practices

Kenntnisse

Linux systems
Python
GPU systems
NVIDIA DCGM
Bare-metal provisioning
Automation
Diagnostics
Hardware-software debugging

Tools

iDRAC
IPMI
Redfish
PXE boot
LiveCD provisioning

Jobbeschreibung

Role overview

This position focuses on bringing up, validating, and operating large-scale bare-metal compute environments with an emphasis on GPU-enabled systems. The engineer will sit at the intersection of hardware, systems software, and automation, owning diagnostics, qualification, and the tooling that keeps clusters ready for demanding AI/ML and HPC workloads.

Responsibilities
  • Own and evolve image management, deployment, and validation pipelines across bare-metal infrastructure, including firmware, driver, and OS qualification for GPU-enabled systems
  • Operate and maintain test clusters used for bring-up and diagnostics, supporting hardware qualification efforts for next-generation platforms
  • Diagnose and resolve complex issues spanning GPUs, drivers, OS, and underlying hardware; analyze performance using tools such as NVIDIA DCGM
  • Build Python-based automation for provisioning, validation, and system bring-up, improving reliability, repeatability, and scalability
  • Manage Linux-based production and validation environments, including virtualization and PXE/image-based bare-metal provisioning workflows
  • Partner with infrastructure, hardware, and data center teams, plus platform and ML stakeholders, to ensure systems meet workload requirements and contribute to provisioning and lifecycle best practices
Requirements
  • 5+ years in infrastructure engineering, systems engineering, or a closely related role
  • Strong Linux systems experience in production environments
  • Hands-on experience with GPU-enabled systems and diagnostic tooling such as NVIDIA DCGM
  • Familiarity with bare-metal provisioning and system bring-up workflows
  • Proficiency in Python or comparable scripting/programming languages for automation
  • Ability to debug complex issues across hardware, OS, GPUs, and system software
Nice to have
  • Experience with high-performance interconnects such as InfiniBand or NVLink
  • Familiarity with PXE boot environments, LiveCD systems, or image-based provisioning workflows
  • Working knowledge of hardware management interfaces (iDRAC, IPMI, Redfish)
  • Data center operations experience with physical hardware
  • Background supporting AI/ML or HPC workloads at scale
  • Experience with GPU validation frameworks or large-scale hardware qualification processes
Benefits and work setup
  • Anticipated annual base salary range of $180,000–$220,000 USD, plus discretionary bonus and equity
  • Comprehensive medical, dental, and vision coverage for employees and eligible dependents
  • Retirement savings support (401(k) matching in the U.S., pension contributions in the U.K.)
  • Unlimited PTO, company holidays, floating holidays, and a two-week company-wide winter break
  • Paid parental and family leave, plus a four-week paid sabbatical after four years of service
  • Annual professional development allowance, plus wellness and work-from-home stipends
  • Flexible schedules with a hybrid model for office-based teams and complimentary in-office meals at hubs
  • Work may be fully remote within the U.S. or hybrid out of office hubs in NYC, SF, Seattle, or London, with occasional team and company offsites; visa sponsorship is not available for this role
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Infrastructure Operations Engineer
Infrastructure Operations Engineer

lightningai • Deutschland

Hybrid
EUR 138.000 - 172.000
Discretionary bonus
Equity
401(k) matching
+1
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 140.000
HPC Cluster Architect
HPC Cluster Architect

nexgencloud • Deutschland

Hybrid
EUR 120.000 - 180.000
Annual discretionary bonus
25 days holiday
Remote or hybrid options
+1
Software Python Engineer (GPU Cloud)
Software Python Engineer (GPU Cloud)

Talanto • Deutschland

Hybrid
EUR 90.000 - 140.000
Remote/Hybrid work
Private medical insurance
Paid sick leave
+4
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
Principal Software Engineer - Compute Infrastructure
Principal Software Engineer - Compute Infrastructure

NVIDIA • Deutschland

Hybrid
EUR 214.000 - 337.000
Equity
Hybrid work model
Senior Technical Marketing Engineer - DSX AI Infrastructure Software
Senior Technical Marketing Engineer - DSX AI Infrastructure Software

NVIDIA • Deutschland

Hybrid
EUR 138.000 - 278.000
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Deutschland

Remote
EUR 103.000 - 147.000
Base salary + equity
Remote-first team
Annual reviews and progression plan
+1
Senior Data Center Infrastructure Engineer
Senior Data Center Infrastructure Engineer

NVIDIA • Deutschland

Remote
EUR 145.000 - 230.000
Equity
Benefits
Technical Program Manager, Infrastructure Delivery
Technical Program Manager, Infrastructure Delivery

lightningai • Deutschland

Hybrid
EUR 138.000 - 189.000
Discretionary bonus
Equity
Medical, dental, vision
+5