InfraOps Engineer – AI Platform & GPU Infra (Hybrid)

Lightning AI

United States

Hybrid

USD 160,000 - 200,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health coverage
Equity/RSU
401(k) matching
Unlimited PTO
Winter break
Parental leave
Learning allowance
Wellness stipend
Sabbatical
Hybrid work
In-office meals

Job summary

Lightning AI in the United States is seeking an experienced Infrastructure Operations Engineer to scale and operate our next-generation GPU infrastructure platform. You will own break/fix operations, incident response, provisioning, observability, and automation to reduce manual toil.

Based in hubs like NYC, SF, Seattle, or London, this role requires at least 2 in-office days per week and collaboration across Infrastructure Engineering, Network Operations, and Software Platform teams to keep

Qualifications

  • 8+ years Linux server experience.
  • 5+ years AWS experience.
  • 2+ years Kubernetes and container fundamentals.
  • 2+ years Terraform and Ansible.
  • 2+ years storage management (NFS/Ceph).
  • Experience with monitoring systems (Prometheus/ELK).
  • Familiarity with GitOps workflows.
  • Software development for automation (Python/Go).
  • Strong networking fundamentals, datacenter-level networks.
  • Able to balance design, risk, cost, outcomes.

Responsibilities

  • Design, build, and roll out new platforms and patterns to minimize incidents and enable customer facing and internal features.
  • Deploy updates and improvements to support both Voltage Park’s internal and end customer use cases.
  • Collaborate with colleagues in Infrastructure Engineering, Network Operations, Customer Success and Software and Platform Development Teams.
  • Participate in the on-call rotation which is evenly distributed across all team members in a primary / secondary pattern where you are primary then move to a secondary position.

Skills

Linux
AWS
Kubernetes
Terraform
Ansible
Networking
Python/Go
GitOps
Monitoring
Automation

Tools

Prometheus
ELK Stack
NFS/Ceph
GitOps tooling
Dell hardware
Infiniband

Job description

Lightning AI in the United States is seeking an experienced Infrastructure Operations Engineer to scale and operate our next-generation GPU infrastructure platform. You will own break/fix operations, incident response, provisioning, observability, and automation to reduce manual toil.

Based in hubs like NYC, SF, Seattle, or London, this role requires at least 2 in-office days per week and collaboration across Infrastructure Engineering, Network Operations, and Software Platform teams to keep

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior InfraOps Engineer — GPU Infra, Hybrid, Equity
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity

Lightning AI • New York (NY)

Hybrid
USD 160,000 - 200,000
Health coverage
Equity/RSUs
401(k) matching
+1
Senior Infrastructure Software Engineer — Remote/Hybrid
Senior Infrastructure Software Engineer — Remote/Hybrid

Lightning AI • New York (NY)

On-site
USD 180,000 - 220,000
Health coverage
Equity (RSUs)
401(k) matching
+7
Infrastructure Operations Engineer
Infrastructure Operations Engineer

Lightning AI • United States

Hybrid
USD 160,000 - 200,000
Health coverage
Equity/RSU
401(k) matching
+8
AI Infrastructure Platform Operations Engineer remote in the US
AI Infrastructure Platform Operations Engineer remote in the US

Mirantis • United States

Remote
USD 110,000 - 150,000
Professional development
Conferences attendance
Team events
Infrastructure Operations Engineer
Infrastructure Operations Engineer

Lightning AI • New York (NY)

Hybrid
USD 160,000 - 200,000
Health coverage
Equity/RSUs
401(k) matching
+1
Product Lead, GPU Cloud & AI Infrastructure
Product Lead, GPU Cloud & AI Infrastructure

Lightning AI • New York (NY), San Francisco (CA)

Hybrid
USD 200,000 - 250,000
Health coverage
Equity opportunity
401(k) + pension
Senior GPU Data Center Engineer
Senior GPU Data Center Engineer

Prime Intellect AI • San Francisco (CA)

On-site
USD 150,000 - 300,000
Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

Lightning AI • San Francisco (CA)

On-site
USD 180,000 - 200,000
Comprehensive medical, dental, and vision coverage
Generous paid time off
Flexible work environment
Infrastructure Engineer (GPU & Compute)
Infrastructure Engineer (GPU & Compute)

Lightning AI • Town of Tonawanda (NY)

On-site
USD 180,000 - 220,000
Health Coverage
Equity
401(k) matching
+8
Hybrid GPU Data Center Engineer: Automation & AI Infra
Hybrid GPU Data Center Engineer: Automation & AI Infra

Blue Signal Search • United States

Hybrid
USD 120,000 - 180,000
Competitive compensation
Equity opportunity
Comprehensive benefits