ML Engineer: AI Systems & Distributed Deployment

NVIDIA

California (MO)

On-site

USD 152,000 - 242,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

NVIDIA is seeking a talented Machine Learning Engineer to lead end-to-end AI system development, deployment, and automation. You will architect scalable models, build AI workflows with Kubernetes, Ray, and Slurm, and implement secure CI/CD pipelines with GitLab.

Expect ownership from ideation to production and cross-team collaboration on high-performance GPU workloads. The role requires a Master’s or PhD in a related field, 3+ years of Python experience, and deep expertise with LangChain,

Qualifications

  • Master's or PhD in Computer Science, Electrical Engineering, or related field, or equivalent experience.
  • 3+ years of professional experience writing production-grade, asynchronous Python and robust system design.
  • Deep experience with LangChain, Hugging Face libraries, vLLM, and SGLang; TensorFlow, PyTorch and Scikit-learn.
  • Proficient in data analysis using Python (pandas, NumPy) and communicating results to technical and non-technical audiences.
  • Production-grade model deployment, monitoring, and multi-node scaling with Kubernetes, Ray, or Slurm.
  • Strong understanding of GPU memory management and infrastructure tuning for high-throughput AI inference.
  • GitLab CI/CD and security automation integrated into MR workflows.
  • Familiarity with PyTest, mocking, and automated test generation for AI workloads.
  • Advanced Git workflows, including rebase strategies and signed commits.

Responsibilities

  • Architect, deploy, and scale open-source models using Kubernetes, Ray, or Slurm for robust AI workloads.
  • Design experiments, prompt-tune, evaluate, and deploy production-grade models and AI agents with scalable data pipelines.
  • Run model benchmarks, perform error analysis, and build dashboards to communicate performance to stakeholders.
  • Take ownership of features from ideation to production across multiple repositories and teams.

Skills

Python
Data analysis
CI/CD automation
Git workflows
Testing frameworks
Distributed AI systems

Education

Master's degree in CS/EE or related
PhD in CS/EE or related

Tools

Kubernetes
Ray
Slurm
LangChain
Hugging Face
vLLM
SGLang
TensorFlow
PyTorch
Scikit-learn
GitLab
PyTest

Job description

NVIDIA is seeking a talented Machine Learning Engineer to lead end-to-end AI system development, deployment, and automation. You will architect scalable models, build AI workflows with Kubernetes, Ray, and Slurm, and implement secure CI/CD pipelines with GitLab.

Expect ownership from ideation to production and cross-team collaboration on high-performance GPU workloads. The role requires a Master’s or PhD in a related field, 3+ years of Python experience, and deep expertise with LangChain,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer - Distributed AI & GPU
ML Systems Engineer - Distributed AI & GPU

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits package
ML Engineer: Scalable AI Systems & GPU Orchestration
ML Engineer: Scalable AI Systems & GPU Orchestration

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
ML Engineer - Scalable AI Systems & GPU Orchestration
ML Engineer - Scalable AI Systems & GPU Orchestration

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Machine Learning Engineer: AI & GPU Orchestration
Machine Learning Engineer: AI & GPU Orchestration

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • California (MO)

On-site
USD 152,000 - 242,000
Machine Learning Engineer
Machine Learning Engineer

NVIDIA AI • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000
Machine Learning Engineer
Machine Learning Engineer

Nvidia Corporation in • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Machine Learning Engineer
Machine Learning Engineer

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits package
Machine Learning Engineer
Machine Learning Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 152,000 - 242,000
Equity
Benefits