Machine Learning Engineer, Inference Infrastructure

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 190,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
Lunch & wellness stipends

Job summary

AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.

The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco.

Qualifications

  • BS/MS/PhD in Computer Science or related technical field.
  • 3+ years of industry experience building large-scale distributed systems or ML infrastructure.
  • Deep proficiency in C++ and Python.
  • Extensive experience with CUDA, Triton, or deep learning hardware accelerators.
  • Familiarity with distributed training and inference frameworks (vLLM, TensorRT-LLM, Megatron).

Responsibilities

  • Architect, build, and scale low-latency distributed inference serving systems for massive generative models.
  • Optimize GPU utilization, memory management, and kernel execution for state-of-the-art transformer architectures.
  • Collaborate with research teams to ensure smooth transition of new model architectures into production.
  • Monitor system performance, troubleshoot bottlenecks, and implement robust reliability measures.

Skills

C++
Python
CUDA
Triton
Megatron

Education

BS/MS/PhD in Computer Science or related field

Tools

vLLM
TensorRT-LLM
Megatron

Job description

About the Role

We are seeking a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.

Responsibilities
  • Architect, build, and scale low-latency distributed inference serving systems for massive generative models.
  • Optimize GPU utilization, memory management, and kernel execution for state-of-the-art transformer architectures.
  • Collaborate with research teams to ensure smooth transition of new model architectures into production.
  • Monitor system performance, troubleshoot bottlenecks, and implement robust reliability measures.
Requirements
  • BS, MS, or Ph.D. in Computer Science or related technical field.
  • 3+ years of industry experience building large-scale distributed systems or ML infrastructure.
  • Deep proficiency in C++ and Python.
  • Extensive experience with CUDA, Triton, or deep learning hardware accelerators.
  • Familiarity with distributed training and inference frameworks (vLLM, TensorRT-LLM, Megatron).
Benefits
  • Top-tier compensation including equity.
  • Full medical, dental, and vision coverage with zero employee contribution options.
  • Unlimited paid time off and flexible hybrid work policy.
  • Catered daily lunches and wellness stipends.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Inference Engineer
Inference Engineer

techire.® • San Francisco (CA)

On-site
USD 140,000 - 210,000
Medical insurance (including dental &视
Dental insurance
Vision insurance
+4
Member of Technical Staff - Mid-Training Infra
Member of Technical Staff - Mid-Training Infra

Reflection • New York (NY)

On-site
USD 150,000 - 200,000
Comprehensive medical, dental, vision, life, and disability insurance
Fully paid parental leave
Daily meals provided
+2
Scale Low-Latency ML Inference Engineer (GPU/CUDA)
Scale Low-Latency ML Inference Engineer (GPU/CUDA)

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
+1
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

inference.net • San Francisco (CA)

Hybrid
USD 220,000 - 320,000
Equity in a high-growth startup
Comprehensive benefits
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian • Palo Alto (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
+1
Inference Performance Engineer
Inference Performance Engineer

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1
Member of Technical Staff — Inference Infrastructure
Member of Technical Staff — Inference Infrastructure

Uncover • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Senior Software Engineer - Model Performance
Senior Software Engineer - Model Performance

Inference • San Francisco (CA)

On-site
USD 220,000 - 320,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits