Infrastructure Operations Engineer

lightningai

Deutschland

Hybrid

EUR 138.000 - 172.000

Vollzeit

Vor 6 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Hebe dich für diese Rolle von der Masse ab — erstelle in etwa einer Minute einen maßgeschneiderten Lebenslauf und ein Anschreiben.

Schaffe es an den ATS-Filtern vorbei

Benefits dieser Stelle

Discretionary bonus
Equity
401(k) matching
Unlimited PTO

Zusammenfassung

Lightning AI is seeking an Infrastructure Operations Engineer to scale and operate a cutting-edge AI infrastructure platform supporting large-scale training and inference workloads. The role focuses on GPU infrastructure reliability, automation, and platform performance, spanning Linux servers, bare metal and virtualized environments, and provisioning workflows.

You will own on-call rotations, deploy updates, and collaborate with cross-functional teams to reduce toil, improve observability, and

Qualifikationen

  • 8+ years Linux server administration experience.
  • 5+ years AWS experience.
  • 2+ years Kubernetes experience.
  • 2+ years Terraform and Ansible.
  • 2+ years managing NFS/ceph storage.
  • Hands-on GPU bare metal or VM servers.
  • Strong networking (switches, routers, firewalls).
  • Software automation using Python/Go/Bash.
  • Familiarity with Prometheus and the ELK stack and gitops.

Aufgaben

  • Design, build, and roll out platforms and patterns to minimize incidents.
  • Deploy updates supporting internal and end-customer use cases.
  • Troubleshoot with cross-team collaboration to improve efficiency.
  • Participate in an on-call rotation with primary/secondary pattern.
  • Build automation to reduce manual toil across the stack.

Kenntnisse

Linux server management
Automation scripting
On-call experience

Tools

AWS
Kubernetes
Terraform
Ansible
NFS/Ceph storage
Python
Go
Bash
Prometheus
ELK stack
GitOps
Git

Jobbeschreibung

Role overview

An Infrastructure Operations Engineer is needed to help scale and operate a next-generation AI infrastructure platform that supports training and inference workloads at large scale. The InfraOps team sits at the center of reliability, automation, and operational scale for GPU infrastructure, owning break/fix operations, incident response, customer provisioning, observability, and the automation systems that keep complex infrastructure running efficiently. The role is hands-on, working across large-scale GPU environments, Linux systems, bare metal infrastructure, provisioning workflows, and platform reliability.

Responsibilities
  • Design, build, and roll out new platforms and patterns that minimize incidents and enable customer-facing and internal features
  • Deploy updates and improvements supporting both internal and end-customer use cases
  • Partner with Infrastructure Engineering, Network Operations, Customer Success, and Software/Platform Development teams to troubleshoot issues and improve operational efficiency
  • Participate in an evenly distributed on-call rotation following a primary/secondary pattern
  • Build automation that reduces manual toil over time across the infrastructure stack
Requirements
  • 8+ years working with Linux as a server/hosting platform, with Ubuntu experience a plus
  • 5+ years of experience with AWS
  • 2+ years of experience with Kubernetes and strong container fundamentals
  • 2+ years of experience with Terraform and Ansible
  • 2+ years managing network-attached storage via NFS, ceph, or similar protocols
  • Hands-on experience with GPU servers in bare metal or virtualized form
  • Deep experience with network switches, routers, and firewalls
  • Software development experience using Python, Go, bash, or similar for automation and integrating systems and APIs
  • Familiarity with monitoring systems such as Prometheus and the ELK stack, plus gitops workflows
Nice to have
  • Experience with VAST storage systems
  • Dell hardware troubleshooting and provisioning experience
  • Datacenter-level networking, 400Gb ethernet, and Infiniband experience
  • Experience with SONiC switches, Palo Alto firewalls, or Juniper Networks
Benefits and work setup
  • Anticipated annual base salary range of $160,000–$200,000 USD, plus discretionary bonus and meaningful equity
  • Comprehensive medical, dental, and vision coverage for employees and eligible dependents
  • 401(k) matching
  • Unlimited PTO, company holidays, and floating holidays
  • Two-week company-wide winter closure
  • Paid parental and family leave
  • Annual learning and development allowance
  • Wellness and work-from-home stipends
  • Four weeks of paid sabbatical leave after four years of service
  • Flexible schedules with a hybrid model for office-based teams
  • Fully remote within the U.S., or hybrid from NYC, SF, Seattle, or London, with occasional team and company offsites
  • Visa sponsorship is not available for this role
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

lightningai • Deutschland

Hybrid
EUR 155.000 - 189.000
Discretionary bonus
Equity
Comprehensive medical coverage
+3
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)
Senior Technical Operations & Deployment Engineer (GPU Cloud Infrastructure)

Jobgether • Deutschland

Vor Ort
EUR 90.000 - 140.000
Software Python Engineer (GPU Cloud)
Software Python Engineer (GPU Cloud)

Talanto • Deutschland

Hybrid
EUR 90.000 - 140.000
Remote/Hybrid work
Private medical insurance
Paid sick leave
+4
Member of Technical Staff - Research Infrastructure Engineer
Member of Technical Staff - Research Infrastructure Engineer

Black Forest Labs • Freiburg im Breisgau

Vor Ort
EUR 100.000 - 230.000
Senior Infrastructure Support Engineer
Senior Infrastructure Support Engineer

Nscale • Deutschland

Remote
EUR 103.000 - 147.000
Base salary + equity
Remote-first team
Annual reviews and progression plan
+1
Technical Lead – GPU Infrastructure
Technical Lead – GPU Infrastructure

Jobtailor • Deutschland

Remote
EUR 120.000 - 160.000
Senior Technical Marketing Engineer - DSX AI Infrastructure Software
Senior Technical Marketing Engineer - DSX AI Infrastructure Software

NVIDIA • Deutschland

Hybrid
EUR 138.000 - 278.000
Senior System Engineer (Munich, Germany)
Senior System Engineer (Munich, Germany)

Remotestar • München

Hybrid
EUR 80.000 - 110.000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+8
HPC Cluster Architect
HPC Cluster Architect

nexgencloud • Deutschland

Hybrid
EUR 120.000 - 180.000
Annual discretionary bonus
25 days holiday
Remote or hybrid options
+1
Principal Software Engineer - Compute Infrastructure
Principal Software Engineer - Compute Infrastructure

NVIDIA • Deutschland

Hybrid
EUR 214.000 - 337.000
Equity
Hybrid work model