ML Engineer: Scalable AI Systems & GPU Orchestration

NVIDIA AI

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

6 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Benefits

Job summary

NVIDIA is seeking a Machine Learning Engineer to drive end-to-end AI system development, deployment, and lifecycle management. You will deploy and scale models across distributed infrastructure, manage GPU orchestration, and build advanced AI workflows with Kubernetes, Ray, or Slurm.

You will design experiments, prompt-tune models, and benchmark performance while collaborating across teams to ensure robust, production-grade AI solutions.

Qualifications

  • Master's/PhD or equivalent experience in CS/EE or related field.
  • 3+ years producing production-grade asynchronous Python with clean architecture.
  • Deep experience with LangChain, Hugging Face, vLLM, and SGLang.
  • Data analysis in Python (pandas, NumPy) to derive insights.
  • Production deployment and scaling with Kubernetes, Ray, or Slurm.

Responsibilities

  • Architect, deploy, and scale models using Kubernetes, Ray, or Slurm for robust AI workloads.
  • Design ML systems and data pipelines; prompt-tune and deploy production models and agents.
  • Benchmark, analyze errors, and build dashboards to communicate performance to stakeholders.
  • Own features from ideation to production across multiple repos and teams.

Skills

Python & Systems Engineering
AI Tools
Data Analysis
Deployment & Orchestration
Hardware & Scaling
GitLab CI/CD & Security
Testing Toolchains
Advanced Git Workflows

Education

Master's or PhD in CS/EE or related field

Tools

LangChain
Hugging Face
vLLM
SGLang
TensorFlow
PyTorch
Scikit-learn

Job description

NVIDIA is seeking a Machine Learning Engineer to drive end-to-end AI system development, deployment, and lifecycle management. You will deploy and scale models across distributed infrastructure, manage GPU orchestration, and build advanced AI workflows with Kubernetes, Ray, or Slurm.

You will design experiments, prompt-tune models, and benchmark performance while collaborating across teams to ensure robust, production-grade AI solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer - Scalable AI Systems & GPU Orchestration
ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
ML Engineer: AI Systems & Distributed Deployment
ML Engineer: AI Systems & Distributed Deployment

NVIDIA • California (MO)

On-site
USD 152,000 - 242,000
ML Systems Engineer - Distributed AI & GPU
ML Systems Engineer - Distributed AI & GPU

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Machine Learning Engineer: AI & GPU Orchestration
Machine Learning Engineer: AI & GPU Orchestration

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 242,000
Machine Learning Engineer
Machine Learning Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Machine Learning Engineer
Machine Learning Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Machine Learning Engineer
Machine Learning Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits package