Senior AI Systems and Algorithms Engineer

Jobtailor

California (MO)

On-site

USD 180,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NVIDIA is seeking a highly skilled researcher/engineer to advance foundation models, train and deploy multimodal data pipelines, and optimize large-scale AI systems. The role emphasizes building scalable infrastructure, reducing costs, and enabling efficient inference across cloud and edge environments, with extensive work on Megatron-LM, Megatron Bridge, and NeMo-RL.

Ideal candidates have 5+ years of experience, strong ML/DL foundations, and hands-on PyTorch development in distributed training

Qualifications

  • MS or PhD in Computer Science or related field (or equivalent experience).
  • 5+ years of relevant industry experience.
  • Strong foundation in machine learning, deep learning, and optimization.
  • Excellent software engineering skills, including Python and PyTorch.
  • Experience building high-performance software for large-scale AI systems.
  • Experience with distributed training at scale, such as Megatron-LM, Megatron Bridge, FSDP, and TP/PP/CP/DP.
  • Experience with optimizer research and efficient sparse or long-context attention.
  • Experience with supervised fine-tuning, reinforcement learning for LLMs, PPO, GRPO, asynchronous RL, or NeMo-RL.
  • Experience with model compression, quantization, pruning, knowledge distillation, neural architecture search, or diffusion/non-autoregressive language models.
  • Experience contributing to open-source AI frameworks such as Megatron-LM, Megatron Bridge, NeMo-RL, or Hugging Face Transformers
  • Experience with GPU performance optimization, distributed systems, latency/throughput analysis, and profiling of large-scale AI workloads

Responsibilities

  • Advance the state of the art in foundation model development, training, and deployment.
  • Design scalable systems for preparing high-quality multimodal datasets for frontier foundation model training.
  • Develop algorithms and systems that improve scalability, efficiency, and cost of large-scale pre-training and post-training.
  • Advance techniques that improve inference performance, reduce deployment cost, and enable efficient serving across cloud and edge platforms.
  • Develop reusable infrastructure and contribute new model support to NVIDIA's open-source GenAI training platform.
  • Collaborate with research, product, and infrastructure teams to design new algorithms and optimize existing systems.
  • Contribute to NVIDIA's open-source AI stack, including Megatron-LM, Megatron Bridge, and NeMo-RL

Skills

Foundation Model Development
Machine Learning
Deep Learning
Python Programming
PyTorch Framework

Education

MS or PhD in Computer Science or related field

Tools

Megatron-LM
Megatron Bridge
NeMo-RL
Hugging Face Transformers

Job description

  • Advance the state of the art in foundation model development, training, and deployment
  • Design scalable systems for preparing high-quality multimodal datasets for frontier foundation model training
  • Develop algorithms and systems that improve the scalability, efficiency, and cost of large-scale pre-training and post-training
  • Advance techniques that improve inference performance, reduce deployment cost, and enable efficient serving across cloud and edge platforms
  • Develop reusable infrastructure and contribute new model support to NVIDIA's open-source GenAI training platform
  • Collaborate with research, product, and infrastructure teams to design new algorithms and optimize existing systems
  • Contribute to NVIDIA's open-source AI stack, including Megatron-LM, Megatron Bridge, and NeMo-RL
Requirements
  • MS or Ph.D in Computer Science, AI, Applied Mathematics, or a related field (or equivalent experience)
  • 5+ years of relevant industry experience
  • Strong foundation in machine learning, deep learning, and optimization
  • Excellent software engineering skills, including Python and PyTorch
  • Experience building high-performance software for large-scale AI systems
  • Strong analytical, debugging, and performance optimization skills
  • Experience with distributed training at scale, such as Megatron-LM, Megatron Bridge, FSDP, and TP/PP/CP/DP
  • Experience with optimizer research and efficient sparse or long-context attention
  • Experience with supervised fine-tuning, reinforcement learning for LLMs, PPO, GRPO, asynchronous RL, or NeMo-RL
  • Experience with model compression, quantization, pruning, knowledge distillation, neural architecture search, or diffusion/non-autoregressive language models
  • Experience contributing to open-source AI frameworks such as Megatron-LM, Megatron Bridge, NeMo-RL, or Hugging Face Transformers
  • Experience with GPU performance optimization, distributed systems, latency/throughput analysis, and profiling of large-scale AI workloads
Core Competencies

Demonstrates expertise in foundation model development, including machine learning and deep learning, with a strong focus on optimizing large-scale AI systems. Proficient in Python and PyTorch, with experience in distributed training and contributing to open-source AI frameworks.

Highest-signal resume keywords
  • Foundation Model Development
  • Machine Learning
  • Deep Learning
  • Python Programming
  • PyTorch Framework
ATS Optimization Keywords
Hard Skills
  • Machine Learning
  • Deep Learning
  • Optimization
  • Distributed Training
  • Model Compression
  • Quantization
  • Pruning
  • Neural Architecture Search
  • Reinforcement Learning
  • Performance Optimization
Soft Skills
  • Analytical Skills
  • Debugging Skills
Certifications & Qualifications
  • MS or Ph.D in Computer Science
  • AI
  • Applied Mathematics
Industry Keywords
  • AI Systems
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer – Local AI
Senior Software Engineer – Local AI

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • United States

On-site
USD 180,000 - 240,000
Senior AI Systems and Algorithms Engineer
Senior AI Systems and Algorithms Engineer

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 152,000 - 288,000
Equity grants
Benefits
Software Engineer, CUDA Deep Learning Systems
Software Engineer, CUDA Deep Learning Systems

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Applied AI Scientist, Senior/Staff
Applied AI Scientist, Senior/Staff

Jobtailor • United States

On-site
USD 120,000 - 180,000
Member of Technical Staff, ML Engineer
Member of Technical Staff, ML Engineer

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
AI and ML Infra Software Engineer, GPU Clusters
AI and ML Infra Software Engineer, GPU Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
Applied AI Scientist, Senior/Staff
Applied AI Scientist, Senior/Staff

Jobtailor • New York (NY)

On-site
USD 150,000 - 230,000
Applied Researcher I, AI Foundations, LLM Core, Agentic AI
Applied Researcher I, AI Foundations, LLM Core, Agentic AI

Jobtailor • New York (NY)

On-site
USD 150,000 - 210,000