Senior Research Scientist – Machine Learning Systems, Efficiency Engineer

Jobtailor

California (MO)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a senior AI systems engineer to design and optimize high-throughput, low-latency inference systems, write GPU kernels, and mentor teams. You will benchmark with PyTorch Profiler, collaborate with infrastructure to build scalable distributed systems, and translate measurements into performance improvements.

The role requires a Master’s/PhD in CS or EE and deep knowledge of GPU architecture, Python, C++, CUDA, and Triton.

Qualifications

  • Master’s or PhD in CS, EE, or related field.
  • Hands-on experience scaling large-scale inference or serving workloads.
  • Strong understanding of GPU architecture.
  • Proficiency in Python and C++.
  • Familiarity with CUDA or Triton for performance‑critical workloads.
  • Demonstrated ability to make engineering decisions based on rigorous measurement and benchmarking.

Responsibilities

  • Design and optimize high‑throughput, low‑latency inference systems
  • Write and maintain high‑performance GPU kernels
  • Conduct deep performance analysis using tools such as PyTorch Profiler
  • Partner with infrastructure teams to design scalable distributed systems
  • Establish and track efficiency metrics and build benchmarking frameworks
  • Serve as a trusted technical advisor to research and product teams

Skills

GPU kernels
Python
C++
GPU architecture
high-throughput systems
low-latency systems
benchmarking
measurement
Triton
CUDA
PyTorch

Education

Master's in CS
PhD in CS
Master's in EE
PhD in EE

Tools

CUDA
Triton
PyTorch

Job description

Responsibilities
  • Design and optimize high‑throughput, low‑latency inference systems
  • Write and maintain high‑performance GPU kernels
  • Conduct deep performance analysis using tools such as PyTorch Profiler
  • Partner with infrastructure teams to design scalable distributed systems
  • Establish and track efficiency metrics and build benchmarking frameworks
  • Serve as a trusted technical advisor to research and product teams
Requirements
  • Master’s or PhD in Computer Science, Electrical Engineering, or related field
  • Hands‑on experience implementing and scaling large‑scale inference or serving workloads
  • Strong understanding of GPU architecture
  • Proficiency in Python and C++
  • Familiarity with CUDA or Triton for performance‑critical workloads
  • Demonstrated ability to make engineering decisions based on rigorous measurement and benchmarking
  • Hard skills: GPU kernels, high‑throughput systems, low‑latency systems, performance analysis, Python, C++, CUDA, Triton, benchmarking frameworks, inference workloads
  • Soft skills: technical advisor, collaboration, decision making, measurement
  • Certifications: Master’s in Computer Science, PhD in Computer Science, Master’s in Electrical Engineering, PhD in Electrical Engineering
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

Susquehanna International Group • Bala Cynwyd (PA)

On-site
USD 100,000 - 130,000
GPU Performance Engineer | Experienced Hire
GPU Performance Engineer | Experienced Hire

SIG Susquehanna • Pennsylvania

On-site
USD 120,000 - 150,000
Machine Learning Systems Engineer
Machine Learning Systems Engineer

Jobtailor • Massachusetts

On-site
USD 120,000 - 180,000
Machine Learning Engineer- Inference Optimization | Experienced Hire
Machine Learning Engineer- Inference Optimization | Experienced Hire

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Research Engineer, GPU Performance
Research Engineer, GPU Performance

Harnham • California (MO)

On-site
USD 120,000 - 160,000
Engineering Manager, Deep Learning Inference
Engineering Manager, Deep Learning Inference

Jobtailor • California (MO)

On-site
USD 180,000 - 260,000
ML Systems Scientist — High-Throughput Inference & GPUs
ML Systems Scientist — High-Throughput Inference & GPUs

Jobtailor • California (MO)

On-site
USD 150,000 - 210,000
Research Engineer
Research Engineer

Harnham • United States

On-site
USD 120,000 - 150,000
Machine Learning Performance Engineer (Inference)
Machine Learning Performance Engineer (Inference)

Career Techniques • New York (NY)

Hybrid
USD 200,000 - 300,000
Training / AI Infrastructure
Training / AI Infrastructure

Genesis AI • San Francisco (CA)

On-site
USD 180,000 - 260,000