Scale Low-Latency ML Inference Engineer (GPU/CUDA)

AI Breaking Wire

San Francisco, Northern (CA, KY)

Hybrid

USD 140,000 - 190,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
Lunch & wellness stipends

Job summary

AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.

The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco.

Qualifications

  • BS/MS/PhD in Computer Science or related technical field.
  • 3+ years of industry experience building large-scale distributed systems or ML infrastructure.
  • Deep proficiency in C++ and Python.
  • Extensive experience with CUDA, Triton, or deep learning hardware accelerators.
  • Familiarity with distributed training and inference frameworks (vLLM, TensorRT-LLM, Megatron).

Responsibilities

  • Architect, build, and scale low-latency distributed inference serving systems for massive generative models.
  • Optimize GPU utilization, memory management, and kernel execution for state-of-the-art transformer architectures.
  • Collaborate with research teams to ensure smooth transition of new model architectures into production.
  • Monitor system performance, troubleshoot bottlenecks, and implement robust reliability measures.

Skills

C++
Python
CUDA
Triton
Megatron

Education

BS/MS/PhD in Computer Science or related field

Tools

vLLM
TensorRT-LLM
Megatron

Job description

AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.

The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer, Inference Infrastructure
Machine Learning Engineer, Inference Infrastructure

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Equity
Medical benefits (full coverage)
PTO & Hybrid work policy
+1
Inference Optimization Engineer: Fast, Cost-Effective ML
Inference Optimization Engineer: Fast, Cost-Effective ML

Build AI • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive pay
Medical, dental, and vision packages
Housing subsidy $2k/month near SF offi
+6
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Low-Latency ML Inference Engineer
Low-Latency ML Inference Engineer

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
AI Inference Engineer – High-Performance GPU Systems
AI Inference Engineer – High-Performance GPU Systems

Perplexity • California (MO)

On-site
USD 120,000 - 170,000
Low-Latency AI Inference Engineer
Low-Latency AI Inference Engineer

OpenAI • California (MO)

On-site
USD 180,000 - 240,000
High-Performance ML Inference Engineer
High-Performance ML Inference Engineer

Reactor • San Francisco (CA)

On-site
USD 180,000 - 240,000
Competitive SF salary
Early equity
Visa sponsorship
+2
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8
Low-Latency Inference Systems Engineer (Multi-GPU)
Low-Latency Inference Systems Engineer (Multi-GPU)

Digital Waffle • San Francisco (CA)

Hybrid
USD 180,000 - 280,000
AI Inference Systems Engineer (High-Throughput, Low-Latency)
AI Inference Systems Engineer (High-Throughput, Low-Latency)

SpaceX • Palo Alto (CA)

On-site
USD 135,000 - 210,000
Stock options
Excellent medical coverage
401(k) plan