ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a talented Machine Learning Engineer to drive end-to-end AI system development, evaluation, deployment and lifecycle management. You will deploy and scale models across distributed infrastructure, manage GPU orchestration, prompt-tune models, and build advanced AI workflows using Kubernetes, Ray, or Slurm.

You will work on designing experiments, benchmarking, and building testing frameworks within secure CI/CD pipelines with GitLab, ensuring code quality and system resilience.

Qualifications

  • Master’s or PhD in CS/EE or equivalent
  • 3+ years of production-grade Python experience
  • Experience with LangChain and ML frameworks (TF/PyTorch/Scikit-learn)
  • Data analysis in Python (pandas/NumPy)
  • Production-grade model deployment and monitoring
  • CI/CD with GitLab and security automation
  • Advanced Git workflows and collaboration
  • GPU memory management for high-throughput AI inference
  • Container orchestration with Kubernetes
  • Benchmarking, analytics dashboards for stakeholders

Responsibilities

  • Architect, deploy, and scale AI models with Kubernetes, Ray, or Slurm
  • Design ML systems and data pipelines; prompt-tune and deploy production models
  • Run model benchmarks and build analytics dashboards for stakeholders
  • Own features from ideation to production across repos

Skills

Python development
Asynchronous Python
GitLab CI/CD
PyTest
Data analysis Python

Education

Master's/PhD in CS/EE

Tools

Kubernetes
Ray
Slurm
LangChain
Hugging Face
vLLM
SGLang
TensorFlow
PyTorch
Scikit-learn

Job description

NVIDIA is seeking a talented Machine Learning Engineer to drive end-to-end AI system development, evaluation, deployment and lifecycle management. You will deploy and scale models across distributed infrastructure, manage GPU orchestration, prompt-tune models, and build advanced AI workflows using Kubernetes, Ray, or Slurm.

You will work on designing experiments, benchmarking, and building testing frameworks within secure CI/CD pipelines with GitLab, ensuring code quality and system resilience.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer - Distributed AI & GPU Pipelines
Machine Learning Engineer - Distributed AI & GPU Pipelines

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Machine Learning Engineer
Machine Learning Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Senior Systems Software Engineer: AI Infra & Kubernetes
Senior Systems Software Engineer: AI Infra & Kubernetes

NVIDIA • Seattle (WA)

Hybrid
USD 184,000 - 357,000
Equity
Health benefits
Flexible work arrangement
+1
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
ML Infra Engineer - GPU Clusters & AIOps
ML Infra Engineer - GPU Clusters & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 124,000 - 196,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]
Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distrib[...]

Intelliswift - An LTTS Company • Sunnyvale (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Health insurance
Flexible work hours
Senior AI Infra Engineer: Kubernetes & Distributed Systems
Senior AI Infra Engineer: Kubernetes & Distributed Systems

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Senior DL Infra Engineer: Scalable GPU AI Training
Senior DL Infra Engineer: Scalable GPU AI Training

NVIDIA • California (MO)

On-site
USD 224,000 - 431,000
Equity
Benefits
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1