ML Performance Engineer: Low-Level Systems & GPUs

Trading Interview

New York (NY)

On-site

USD 170,000 - 210,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Jane Street is seeking an engineer with expertise in low-level systems programming and optimisation to join our ML team. The role focuses on optimizing model performance for training and inference, with attention to scalable training, low-latency real-time inference, and high-throughput research workloads.

You will engage with CUDA, GPU memory hierarchies, and distributed systems, ensuring throughput aligns with real throughput expectations.

Qualifications

  • Understanding of modern ML techniques and toolsets.
  • Experience debugging a training run's performance end to end.
  • Low-level GPU knowledge including PTX, SASS, warps, Tensor Cores.
  • Familiar with CUDA graph launch, memory hierarchy and synchronization.
  • Experience with GPU networking concepts (Infiniband, NVLink, RoCE).

Responsibilities

  • Optimize performance of ML models for training and inference.
  • Improve throughput in research and real-time systems; reduce latency.
  • Work across storage, networking and host/GPU levels for end-to-end efficiency.
  • Debug performance bottlenecks using CUDA tools and profiling suites.

Skills

ML techniques
Performance debugging
GPU architecture

Tools

CUDA GDB
NSight Systems
NSight Compute
PTX
SASS
Tensor Cores
Triton
CUTLASS
CUB
Thrust
cuDNN
cuBLAS
Infiniband
NVLink

Job description

Jane Street is seeking an engineer with expertise in low-level systems programming and optimisation to join our ML team. The role focuses on optimizing model performance for training and inference, with attention to scalable training, low-latency real-time inference, and high-throughput research workloads.

You will engage with CUDA, GPU memory hierarchies, and distributed systems, ensuring throughput aligns with real throughput expectations.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Performance Engineer
Machine Learning Performance Engineer

Trading Interview • New York (NY)

On-site
USD 170,000 - 210,000
Staff ML Systems Engineer - GPU & Performance
Staff ML Systems Engineer - GPU & Performance

Google • Sunnyvale (CA)

On-site
USD 207,000 - 300,000
Equity
Performance bonus
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
Staff ML Performance Engineer - GPU Optimization
Staff ML Performance Engineer - GPU Optimization

Google LLC • Sunnyvale (CA)

On-site
USD 186,000 - 228,000
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
ML Systems Performance Engineer
ML Systems Performance Engineer

Foundation Capital • United States

On-site
USD 100,000 - 130,000
ML Research Engineer: From Research to Production
ML Research Engineer: From Research to Production

Trading Interview • Northern (KY), New York (NY)

Hybrid
USD 150,000 - 230,000
Lead GPU ML Training Performance Engineer
Lead GPU ML Training Performance Engineer

Advanced Micro Devices, Inc. • San Jose (CA)

On-site
USD 120,000 - 160,000
ML Systems Performance Engineer — Hardware Co-Design
ML Systems Performance Engineer — Hardware Co-Design

Cerebras Systems • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
LLM Performance Engineer — GPU & HPC
LLM Performance Engineer — GPU & HPC

Baseten • San Francisco (CA)

On-site
USD 160,000 - 200,000
Competitive compensation
Equity
Medical/dental/vision insurance
+4