AI Platform Engineer: Production ML & Kubernetes

Lightning AI

New York (NY)

Hybrid

USD 115,000 - 140,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health coverage
Equity (RSUs)
401(k) matching
Unlimited PTO
Hybrid work model
Office meals

Job summary

Lightning AI is hiring an AI Platform Support Engineer to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production.

You will be a technical partner to ML teams, diagnosing failures, improving reliability, and guiding customers through complex distributed systems problems in a hybrid role out of Seattle, SF, or NY with 2 days in-office per week.

Qualifications

  • Strong software engineering and systems troubleshooting background.
  • Experience with Kubernetes and containerized environments.
  • Linux knowledge including networking, storage, and performance tuning.
  • Cloud infrastructure and distributed systems experience.
  • Observability and debugging tools such as Prometheus, Grafana, or OpenTelemetry.
  • Hands-on experience operating machine learning workloads in production or research environments.
  • Experience with PyTorch, CUDA, or NCCL; familiarity with GPU infrastructure.

Responsibilities

  • Partner with customer engineering teams running training and inference workloads in production.
  • Diagnose and resolve complex distributed systems and ML infrastructure issues.
  • Act as a technical advisor during high impact incidents and platform degradation events.
  • Translate infrastructure issues into actionable guidance for ML engineers.
  • Build credibility with customers through strong technical reasoning and clear communication.
  • Investigate failures involving distributed training, Kubernetes, GPU allocation, networking, and storage.
  • Troubleshoot PyTorch, CUDA, NCCL, and inference serving issues; analyze logs and traces.
  • Debug containerized workloads across Kubernetes and bare metal GPU environments.
  • Support customers scaling workloads across multi-node GPU systems.
  • Identify recurring patterns and drive long-term reliability improvements.

Skills

Kubernetes
Linux
Python
Distributed systems
Cloud infrastructure
Observability
PyTorch/CUDA
Networking

Education

Bachelor's degree in CS or related field

Tools

Kubeflow
Ray
Slurm
Prometheus
Grafana
OpenTelemetry

Job description

Lightning AI is hiring an AI Platform Support Engineer to join our US Customer Experience team, supporting ML engineers running large-scale training and inference workloads across cloud infrastructure, Kubernetes, and GPU platforms in production.

You will be a technical partner to ML teams, diagnosing failures, improving reliability, and guiding customers through complex distributed systems problems in a hybrid role out of Seattle, SF, or NY with 2 days in-office per week.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Support Engineer: Production ML Infra - Hybrid | Unlimited PTO
AI Platform Support Engineer: Production ML Infra - Hybrid | Unlimited PTO

Lightning-Ai • Seattle (WA)

Hybrid
USD 115,000 - 140,000
Health coverage
Meaningful equity
401(k) matching
+5
AI Platform Support Engineer (US)
AI Platform Support Engineer (US)

Lightning-Ai • Seattle (WA)

Hybrid
USD 115,000 - 140,000
Health coverage
Meaningful equity
401(k) matching
+5
AI Platform Support Engineer (US)
AI Platform Support Engineer (US)

Lightning AI • New York (NY)

Hybrid
USD 115,000 - 140,000
Health coverage
Equity (RSUs)
401(k) matching
+3
Fullstack AI Platform Engineer (React/Go)
Fullstack AI Platform Engineer (React/Go)

Lightning AI • San Francisco (CA), Seattle (WA), New York (NY)

Hybrid
USD 120,000 - 250,000
Health coverage
Equity (RSUs)
401(k) Matching
+8
Customer-Facing Platform Engineer for AI Deployments
Customer-Facing Platform Engineer for AI Deployments

Lightning-Ai • New York (NY)

Hybrid
USD 150,000 - 250,000
Health coverage
Equity
401(k) matching
+8
AI Platform Engineer — Scalable ML Infra (Hybrid NYC)
AI Platform Engineer — Scalable ML Infra (Hybrid NYC)

Accrete • New York (NY)

Hybrid
USD 150,000 - 180,000
Health & wellness benefits
Flexible PTO
Catered lunch
+1
AI Platform Engineer — Build & Run Production Systems
AI Platform Engineer — Build & Run Production Systems

Long Lake • San Francisco (CA), New York (NY)

On-site
USD 120,000 - 150,000
ML Platform Engineer: Scalable AI Infrastructure
ML Platform Engineer: Scalable AI Infrastructure

Bjak • Germany (OH)

On-site
USD 120,000 - 180,000
ML Platform Engineer: Scale AI Infra, Deploy & Optimize
ML Platform Engineer: Scale AI Infra, Deploy & Optimize

United States Digital Space LLC • United States

Remote
USD 120,000 - 180,000
Software Engineer, AI Platform — Production ML
Software Engineer, AI Platform — Production ML

Genios AI, Inc. • Santa Clara (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Competitive Compensation
Unlimited PTO
AI Assistants for work