Senior AI Infra Engineer - Scale ML Platforms

NVIDIA

Santa Clara (CA)

On-site

USD 184,000 - 357,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA's DGX Cloud Lepton Team is seeking an AI infrastructure software engineer to design, build, and maintain scalable AI platforms enabling large-scale AI training, inference, and Agentic AI in production. This role emphasizes reliability, API co-design, and end-to-end triage from application to hardware.

You will work on platform tooling, metrics, and tooling improvements with a culture of blameless postmortems and iterative improvement.

Qualifications

  • Minimum 8+ years of software infrastructure experience for large-scale AI systems.
  • Bachelor's degree in Computer Science or related technical field (or equivalent).
  • Strong debugging and triage skills from application to hardware level.
  • Proven track record in building and scaling distributed systems.

Responsibilities

  • Develop platform and tools for large-scale AI, LLM, and GenAI infrastructure.
  • Develop and optimize tools to improve AI/ML workload efficiency and resiliency.
  • Root cause analysis and triage failures from application to hardware level.
  • Enhance infrastructure and products underpinning NVIDIA's AI platforms.
  • Co-design and implement APIs for integration with resiliency stacks.
  • Define meaningful reliability metrics to track system and service reliability.
  • Show strong problem-solving, root-cause analysis, and optimization skills.

Skills

Large-scale AI systems
Distributed systems
Kubernetes
Observability platforms
Golang
Python
C/C++
Troubleshooting
APIs
Data infrastructure

Education

Bachelor's degree in Computer Science or related field

Tools

ELK
Prometheus
Loki

Job description

NVIDIA's DGX Cloud Lepton Team is seeking an AI infrastructure software engineer to design, build, and maintain scalable AI platforms enabling large-scale AI training, inference, and Agentic AI in production. This role emphasizes reliability, API co-design, and end-to-end triage from application to hardware.

You will work on platform tooling, metrics, and tooling improvements with a culture of blameless postmortems and iterative improvement.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer — Scale Cloud AI Platforms
Senior AI Infra Engineer — Scale Cloud AI Platforms

NVIDIA • United States

On-site
USD 184,000 - 357,000
Senior AI Infrastructure Software Engineer - DGX Cloud
Senior AI Infrastructure Software Engineer - DGX Cloud

NVIDIA • United States

On-site
USD 184,000 - 357,000
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Senior AI Infrastructure Software Engineer - DGX Cloud
Senior AI Infrastructure Software Engineer - DGX Cloud

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infra Engineer — Scale AI Workloads & Equity
Senior AI Infra Engineer — Scale AI Workloads & Equity

Socket.dev • Santa Clara (UT)

On-site
USD 184,000 - 357,000
Equity
Health benefits
Senior AI Infrastructure Software Engineer - DGX Cloud
Senior AI Infrastructure Software Engineer - DGX Cloud

Socket.dev • Santa Clara (UT)

On-site
USD 184,000 - 357,000
Equity
Health benefits
Senior Full-Stack Lead Engineer: AI Infra Platform
Senior Full-Stack Lead Engineer: AI Infra Platform

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 224,000 - 357,000
Senior AI Infrastructure Software Engineer - DGX Cloud
Senior AI Infrastructure Software Engineer - DGX Cloud

NVIDIA AI • Seattle (WA)

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior Performance Engineer — AI Scale & Efficiency
Senior Performance Engineer — AI Scale & Efficiency

NVIDIA • Washington

On-site
USD 224,000 - 432,000
Equity
Benefits
Senior AI Infra Engineer — Telemetry & ML Ops Equity
Senior AI Infra Engineer — Telemetry & ML Ops Equity

NVIDIA • Durham (NC)

On-site
USD 184,000 - 357,000