ML Performance Engineer: Low-Level Systems & GPUs

Trading Interview

New York (NY)

On-site

USD 170,000 - 210,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jane Street is seeking an engineer with expertise in low-level systems programming and optimisation to join our ML team. The role focuses on optimizing model performance for training and inference, with attention to scalable training, low-latency real-time inference, and high-throughput research workloads.

You will engage with CUDA, GPU memory hierarchies, and distributed systems, ensuring throughput aligns with real throughput expectations.

Qualifications

  • Understanding of modern ML techniques and toolsets.
  • Experience debugging a training run's performance end to end.
  • Low-level GPU knowledge including PTX, SASS, warps, Tensor Cores.
  • Familiar with CUDA graph launch, memory hierarchy and synchronization.
  • Experience with GPU networking concepts (Infiniband, NVLink, RoCE).

Responsibilities

  • Optimize performance of ML models for training and inference.
  • Improve throughput in research and real-time systems; reduce latency.
  • Work across storage, networking and host/GPU levels for end-to-end efficiency.
  • Debug performance bottlenecks using CUDA tools and profiling suites.

Skills

ML techniques
Performance debugging
GPU architecture

Tools

CUDA GDB
NSight Systems
NSight Compute
PTX
SASS
Tensor Cores
Triton
CUTLASS
CUB
Thrust
cuDNN
cuBLAS
Infiniband
NVLink

Job description

Jane Street is seeking an engineer with expertise in low-level systems programming and optimisation to join our ML team. The role focuses on optimizing model performance for training and inference, with attention to scalable training, low-latency real-time inference, and high-throughput research workloads.

You will engage with CUDA, GPU memory hierarchies, and distributed systems, ensuring throughput aligns with real throughput expectations.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Performance Engineer
Machine Learning Performance Engineer

Trading Interview • New York (NY)

On-site
USD 170,000 - 210,000
ML Performance Engineer: GPU/CUDA at Scale
ML Performance Engineer: GPU/CUDA at Scale

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000
ML Platform Engineer: Production-Driven & Research-Focused
ML Platform Engineer: Production-Driven & Research-Focused

Jane Street • United States

On-site
USD 180,000 - 320,000
ML Performance Engineer: Scale GPU-Driven Training
ML Performance Engineer: Scale GPU-Driven Training

Decisive Point • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
GPU Performance Engineer: Scale ML Inference & Systems
GPU Performance Engineer: Scale ML Inference & Systems

Anthropic • New York (NY)

Hybrid
USD 280,000 - 850,000
ML Inference Engineer - Low-Latency GPU Systems
ML Inference Engineer - Low-Latency GPU Systems

Susquehanna International Group, LLP • Bala Cynwyd (PA)

On-site
USD 110,000 - 150,000
Senior ML Performance Engineer: LLM Benchmarking & GPU
Senior ML Performance Engineer: LLM Benchmarking & GPU

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
+2
ML Systems Performance Engineer
ML Systems Performance Engineer

Cerebras • United States

On-site
USD 100,000 - 130,000
Staff ML Performance Engineer - GPU & Inference
Staff ML Performance Engineer - GPU & Inference

Modal • San Francisco (CA)

On-site
USD 180,000 - 260,000
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Selby Jennings • Chicago (IL)

On-site
USD 140,000 - 210,000