ML Systems Scientist — High-Throughput Inference & GPUs

Jobtailor

California (MO)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior AI systems engineer to design and optimize high-throughput, low-latency inference systems, write GPU kernels, and mentor teams. You will benchmark with PyTorch Profiler, collaborate with infrastructure to build scalable distributed systems, and translate measurements into performance improvements.

The role requires a Master’s/PhD in CS or EE and deep knowledge of GPU architecture, Python, C++, CUDA, and Triton.

Qualifications

  • Master’s or PhD in CS, EE, or related field.
  • Hands-on experience scaling large-scale inference or serving workloads.
  • Strong understanding of GPU architecture.
  • Proficiency in Python and C++.
  • Familiarity with CUDA or Triton for performance‑critical workloads.
  • Demonstrated ability to make engineering decisions based on rigorous measurement and benchmarking.

Responsibilities

  • Design and optimize high‑throughput, low‑latency inference systems
  • Write and maintain high‑performance GPU kernels
  • Conduct deep performance analysis using tools such as PyTorch Profiler
  • Partner with infrastructure teams to design scalable distributed systems
  • Establish and track efficiency metrics and build benchmarking frameworks
  • Serve as a trusted technical advisor to research and product teams

Skills

GPU kernels
Python
C++
GPU architecture
high-throughput systems
low-latency systems
benchmarking
measurement
Triton
CUDA
PyTorch

Education

Master's in CS
PhD in CS
Master's in EE
PhD in EE

Tools

CUDA
Triton
PyTorch

Job description

Jobtailor is seeking a senior AI systems engineer to design and optimize high-throughput, low-latency inference systems, write GPU kernels, and mentor teams. You will benchmark with PyTorch Profiler, collaborate with infrastructure to build scalable distributed systems, and translate measurements into performance improvements.

The role requires a Master’s/PhD in CS or EE and deep knowledge of GPU architecture, Python, C++, CUDA, and Triton.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Systems Engineer: Optimizing Training & GPU Kernels
ML Systems Engineer: Optimizing Training & GPU Kernels

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Senior Inference Systems Engineer — Low-Latency ML Serving
Senior Inference Systems Engineer — Low-Latency ML Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000
Senior Research Scientist – Machine Learning Systems, Efficiency Engineer
Senior Research Scientist – Machine Learning Systems, Efficiency Engineer

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Senior AI Systems Engineer: GPU Kernels & Inference Equity
Senior AI Systems Engineer: GPU Kernels & Inference Equity

NVIDIA AI • Michigan

On-site
USD 150,000 - 190,000
Equity
Health Insurance
GPU Systems Engineer — Distributed Training & Inference
GPU Systems Engineer — Distributed Training & Inference

TensorScale AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
ML Systems Engineer: AI Infra & GPU Acceleration
ML Systems Engineer: AI Infra & GPU Acceleration

Meta • San Francisco (CA)

On-site
USD 180,000 - 240,000
Bonus
Equity
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
GPU AI/ML Infra Engineer for HPC Clusters
GPU AI/ML Infra Engineer for HPC Clusters

Jobtailor • California (MO)

On-site
USD 120,000 - 190,000
Senior AI Systems Engineer - Inference & GPU Kernels
Senior AI Systems Engineer - Inference & GPU Kernels

NVIDIA • Westford (MA)

On-site
USD 184,000 - 288,000
Equity
Benefits