Distributed ML Engineer for High-Performance AI Training

Ifm Us

Sunnyvale (CA)

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical benefits
Dental benefits
Vision benefits
Bonus
401K Plan
Paid time off
Parental Leave
Employee Assistance Program
Life insurance
Disability insurance

Job summary

MBZUAI in Sunnyvale, CA seeks a Distributed ML Engineer to optimize ML software stacks for training and inference on state-of-the-art hardware. You will prototype kernels, build benchmarks, and help scale large‑scale models while collaborating with researchers and engineers.

The role emphasizes parallel computing, system‑level coding and hands-on ML experience, with visa sponsorship available. You will contribute to cutting‑edge foundational AI research and advanced HPC capabilities.

Qualifications

  • Ph.D. in CS, EE or CSEE with 1+ years working experience, OR Master’s in CS, EE or CSEE or equivalent experience with 2+ year working experience.

Responsibilities

  • Understand, analyze, profile, optimize, and provide guidance on deep learning workloads to improve efficiency.
  • Design and implement performance benchmarks and testing methodologies.
  • Build tools to automate workload analysis and optimization.
  • Triage system issues and identify bottlenecks to enhance GPU utilization.
  • Support kernels and systems for new model architectures and algorithms.
  • Participate in or lead design reviews and decide among technologies.
  • Review code for best practices and quality.
  • Contribute to documentation and educational content.
  • Represent MBZUAI at conferences and events to showcase HPC and DL capabilities.
  • Perform other duties as directed by the line manager.

Skills

Parallel computing
System level coding
GPU utilization
Performance benchmarking
Code review

Education

Ph.D. in CS, EE or CSEE
Master’s in CS, EE or CSEE

Job description

MBZUAI in Sunnyvale, CA seeks a Distributed ML Engineer to optimize ML software stacks for training and inference on state-of-the-art hardware. You will prototype kernels, build benchmarks, and help scale large‑scale models while collaborating with researchers and engineers.

The role emphasizes parallel computing, system‑level coding and hands-on ML experience, with visa sponsorship available. You will contribute to cutting‑edge foundational AI research and advanced HPC capabilities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Distributed ML Research Scientist — Frontier-Scale Training
Distributed ML Research Scientist — Frontier-Scale Training

Ifm Us • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Health benefits
401K Plan
Paid time off
+2
Distributed Machine Learning Engineer
Distributed Machine Learning Engineer

Ifm Us • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
Medical benefits
Dental benefits
Vision benefits
+7
ML Performance Engineer: Scale Training & Throughput
ML Performance Engineer: Scale Training & Throughput

Applied Intuition • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Distributed Machine Learning Engineer
Distributed Machine Learning Engineer

Institute of Foundation Models • Sunnyvale (CA)

On-site
USD 150,000 - 450,000
Comprehensive medical, dental, and vision benefits
Bonus
401K Plan
+4
Distributed AI Systems Architect - High-Performance ML
Distributed AI Systems Architect - High-Performance ML

Baidu USA • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Senior ML Performance Engineer: Scale & Throughput
Senior ML Performance Engineer: Scale & Throughput

NLP PEOPLE • Sunnyvale (CA)

On-site
USD 215,000 - 285,000
Distributed ML Training Systems Engineer
Distributed ML Training Systems Engineer

Mind Robotics • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior ML Systems Scientist — High-Performance Inference
Senior ML Systems Scientist — High-Performance Inference

ByteDance • San Jose (CA)

On-site
USD 212,800 - 387,600
Medical insurance
Dental insurance
Vision insurance
+5
Infrastructure Kernel Engineer for Scalable AI Training
Infrastructure Kernel Engineer for Scalable AI Training

Thinkingmachines • San Francisco (CA)

On-site
USD 350,000 - 475,000
Generous health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
High-Performance AI Research Engineer (CUDA/ML)
High-Performance AI Research Engineer (CUDA/ML)

Metamorphic • Palo Alto (CA)

On-site
USD 200,000 - 280,000
Visa sponsorship
Competitive compensation
Mentorship and career development
+1