Machine Learning Engineer - Distributed AI & GPU Pipelines

NVIDIA Corporation

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

4 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a talented Machine Learning Engineer to drive development, evaluation, deployment and end-to-end lifecycle management of AI-powered systems. You will apply AI agents, build automated testing frameworks, and implement secure CI/CD pipelines with GitLab to ensure code quality and system resilience.

A core component is deploying and scaling models across distributed infrastructure, managing GPU orchestration, and building advanced AI workflows with Kubernetes, Ray, or Slurm.

Qualifications

  • Master's or PhD in CS/EE or related field or equivalent experience.
  • 3+ years of professional experience writing production-grade, asynchronous Python.
  • Deep experience with LangChain, Hugging Face libraries, vLLM, and SGLang.
  • Experience with ML frameworks like TensorFlow, PyTorch and Scikit-learn.
  • Data analysis in Python (pandas/NumPy) to communicate results clearly.
  • Deployment and orchestration with Kubernetes, Ray, or Slurm.

Responsibilities

  • Architect, deploy, and scale open-source models using Kubernetes, Ray, or Slurm.
  • Design ML systems and data pipelines; evaluate and deploy production models and agents.
  • Run benchmarks and perform error and gap analysis; build performance dashboards.
  • Own features from ideation to production across multiple repositories.

Skills

Python
Kubernetes
Ray
Slurm
GitLab CI/CD
PyTorch

Education

Master's or PhD in CS/EE or related

Tools

LangChain
Hugging Face
vLLM
SGLang
TensorFlow
PyTorch

Job description

NVIDIA is seeking a talented Machine Learning Engineer to drive development, evaluation, deployment and end-to-end lifecycle management of AI-powered systems. You will apply AI agents, build automated testing frameworks, and implement secure CI/CD pipelines with GitLab to ensure code quality and system resilience.

A core component is deploying and scaling models across distributed infrastructure, managing GPU orchestration, and building advanced AI workflows with Kubernetes, Ray, or Slurm.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer - Scalable AI Systems & GPU Orchestration
ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Machine Learning Engineer
Machine Learning Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior AI Infra Engineer: Kubernetes & Distributed Systems
Senior AI Infra Engineer: Kubernetes & Distributed Systems

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
AI Systems Engineer - Distributed, Multi-GPU (Equity)
AI Systems Engineer - Distributed, Multi-GPU (Equity)

NVIDIA AI • Eugene (OR)

On-site
USD 120,000 - 180,000
Equity
Health Insurance
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)
Senior AI Infra Engineer-Distributed GPU Clusters (Equity)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Senior DL Infra Engineer: Scalable GPU AI Training
Senior DL Infra Engineer: Scalable GPU AI Training

NVIDIA • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Software Engineer - AI Research Clusters
Software Engineer - AI Research Clusters

NVIDIA • Austin (TX)

On-site
USD 124,000 - 196,000
Equity
Benefits