ML Systems Engineer - Distributed AI & GPU

NVIDIA Gruppe

Santa Clara (CA)

On-site

USD 152,000 - 242,000

Full time

11 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Equity
Benefits package

Job summary

NVIDIA is seeking a talented Machine Learning Engineer to drive end-to-end lifecycle management of AI-powered systems across distributed infrastructure. You will deploy and scale models, manage GPU orchestration, and build automated testing and CI/CD pipelines using GitLab.

The role demands strong Python engineering skills, experience with Kubernetes, Ray, Slurm, and ML frameworks, and a track record of delivering production-grade AI workflows at scale in a fast-paced environment.

Qualifications

  • Master’s or PhD in Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • 3+ years of production-grade Python and distributed systems design.
  • Experience with LangChain, Hugging Face, vLLM, and SGLang; TensorFlow, PyTorch, Scikit-learn.
  • Proficient in Python data analysis with pandas/NumPy.
  • Production-grade model deployment and multi-node scaling with Kubernetes, Ray, Slurm.
  • GPU memory management and infrastructure tuning.
  • Expert with GitLab CI/CD and security automation.
  • Testing with PyTest and automated test generation.
  • Advanced Git workflows and repository mirroring.

Responsibilities

  • Architect, deploy, and scale open-source models using distributed frameworks (Kubernetes, Ray, Slurm).
  • Design ML systems and data pipelines; run experiments and benchmark performance.
  • Build analytics dashboards to communicate model performance to stakeholders.
  • Own features from ideation to production across code repositories and teams.

Skills

Advanced degree in CS/EE
Python & Systems Eng
LangChain
HuggingFace
TensorFlow
PyTorch
Scikit-learn
Data analysis (Python)
GPU memory management
GitLab CI/CD
PyTest
Git workflows

Education

Master’s or PhD in Computer Science / Electrical Engineering

Tools

Kubernetes
Ray
Slurm

Job description

NVIDIA is seeking a talented Machine Learning Engineer to drive end-to-end lifecycle management of AI-powered systems across distributed infrastructure. You will deploy and scale models, manage GPU orchestration, and build automated testing and CI/CD pipelines using GitLab.

The role demands strong Python engineering skills, experience with Kubernetes, Ray, Slurm, and ML frameworks, and a track record of delivering production-grade AI workflows at scale in a fast-paced environment.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Engineer - Scalable AI Systems & GPU Orchestration
ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
ML Engineer: AI Systems & Distributed Deployment
ML Engineer: AI Systems & Distributed Deployment

NVIDIA • California (MO)

On-site
USD 152,000 - 242,000
ML Engineer: Scalable AI Systems & GPU Orchestration
ML Engineer: Scalable AI Systems & GPU Orchestration

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Machine Learning Engineer: AI & GPU Orchestration
Machine Learning Engineer: AI & GPU Orchestration

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 242,000
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads
ML Systems Engineer: Inference & GPU-Driven Distributed Workloads

Bake AI • San Mateo (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Realtime ML Systems Engineer, Networking & AIOps
Realtime ML Systems Engineer, Networking & AIOps

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 152,000 - 288,000
AI Systems Engineer - Distributed, Multi-GPU (Equity)
AI Systems Engineer - Distributed, Multi-GPU (Equity)

NVIDIA AI • Eugene (OR)

On-site
USD 120,000 - 180,000
Equity
Health Insurance