InfraOps Engineer - AI GPU Platform

Lightning-Ai

San Francisco (CA)

Hybrid

USD 160,000 - 200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental and vision coverage (U
Pension contribution
Generous paid time off
Paid parental leave
Wellness stipend
Flexible work environment

Job summary

Lightning AI is seeking an experienced Infrastructure Operations Engineer to scale and operate our AI infrastructure platform. You will own reliability, automation, and incident response for large-scale GPU environments, partnering with Infrastructure Engineering, Network Operations and Software Platform teams to improve efficiency and robustness.

This role is based in one of our hubs (NYC, SF, Seattle, or London) with at least two in-office days per week.

Qualifications

  • 8+ years Linux server/hosting experience, Ubuntu a bonus.
  • 5+ years AWS experience.
  • 2+ years Kubernetes experience and container fundamentals.
  • 2+ years Terraform and Ansible experience.
  • 2+ years NAS management (NFS/Ceph)
  • Experience with monitoring systems (Prometheus/ELK).
  • Familiarity with GitOps workflow.
  • Automation programming in Python/Go/Bash for system integration.
  • Strong networking fundamentals, including data center networks.
  • Experience delivering complex systems and balancing design, risk, cost and outcomes.
  • Excellent written and verbal communication.

Responsibilities

  • Design, build, and roll out platforms to minimize incidents and enable features.
  • Deploy updates to support internal and end-user use cases.
  • Collaborate with Infra Eng, Network Ops, Customer Success and Software/Platform teams.
  • Participate in on-call rotation with primary/secondary scheduling.

Skills

Linux server administration
AWS
Networking fundamentals
Python/Go scripting
Communication skills
Cloud infrastructure design

Tools

Kubernetes
Terraform
Ansible
Prometheus/ELK
GitOps workflow
NFS/Ceph storage

Job description

Lightning AI is seeking an experienced Infrastructure Operations Engineer to scale and operate our AI infrastructure platform. You will own reliability, automation, and incident response for large-scale GPU environments, partnering with Infrastructure Engineering, Network Operations and Software Platform teams to improve efficiency and robustness.

This role is based in one of our hubs (NYC, SF, Seattle, or London) with at least two in-office days per week.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU & Compute Infra Engineer - Bare-Metal & AI
GPU & Compute Infra Engineer - Bare-Metal & AI

Lightning-Ai • New York (NY)

Hybrid
USD 180,000 - 200,000
Medical coverage (US)
Dental coverage (US)
Vision coverage (US)
+6
Infrastructure Operations Engineer
Infrastructure Operations Engineer

Lightning-Ai • San Francisco (CA)

Hybrid
USD 160,000 - 200,000
Medical, dental and vision coverage (U
Pension contribution
Generous paid time off
+3
Senior Backend Engineer, GPU Cloud Infra & Kubernetes
Senior Backend Engineer, GPU Cloud Infra & Kubernetes

Socket.dev • New York (NY)

Hybrid
USD 180,000 - 250,000
Health insurance
Equity
401(k) matching
+6
Hybrid Storage Infrastructure Engineer for AI/ML
Hybrid Storage Infrastructure Engineer for AI/ML

Lightningai • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 220,000
Health coverage
Equity
401(k) match
+8
Forward-Deployed Platform Engineer for AI Systems
Forward-Deployed Platform Engineer for AI Systems

Lightningai • San Francisco (CA), New York (NY)

Hybrid
USD 120,000 - 250,000
Health coverage
Equity
Retirement savings
+8
Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

Lightning-Ai • New York (NY)

Hybrid
USD 180,000 - 200,000
Medical coverage (US)
Dental coverage (US)
Vision coverage (US)
+6
Backend Engineer (Go) for Scalable AI Platform
Backend Engineer (Go) for Scalable AI Platform

Lightningai • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 250,000
Comprehensive Health Coverage
Meaningful Equity
401(k) matching
+8
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
AI Platform Engineer: Production ML & Kubernetes
AI Platform Engineer: Production ML & Kubernetes

Lightning AI • New York (NY)

Hybrid
USD 115,000 - 140,000
Health coverage
Equity (RSUs)
401(k) matching
+3
AI Infrastructure & Platform Operations Engineer (remote in the US)
AI Infrastructure & Platform Operations Engineer (remote in the US)

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events