Senior ML Engineer, Serving & Optimization

Selby Jennings

New York (NY)

On-site

USD 180,000 - 250,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Selby Jennings is building a new machine learning infrastructure initiative and seeks an experienced engineer to own the design and deployment of large-scale model serving platforms. You will influence architecture decisions and work closely with researchers and engineers to push state-of-the-art models into production on GPU farms.

The role emphasizes ownership from day one, with opportunities to shape the platform roadmap and optimize hardware utilization across growing GPU clusters in a

Qualifications

  • 8+ years in software, systems or ML infra.
  • Deep experience with model serving and optimization at scale.
  • Strong knowledge of CUDA, kernel optimization and GPU acceleration.
  • Expert programming in C++ and Python.

Responsibilities

  • Design and build large-scale serving and inference infra for state-of-the-art ML models.
  • Optimize performance, latency and GPU utilization in production.
  • Deploy and scale models across thousands of GPUs; improve efficiency and reliability.
  • Develop compression, quantization, distillation techniques to maximize hardware performance.
  • Move cutting-edge models from research to production with hardware integration.
  • Collaborate with researchers and engineers to speed experimentation and deployment.
  • Shape architecture and roadmap for a rapidly growing ML platform.

Skills

Large-scale infra
CUDA GPU
C++
Python
PyTorch
JAX
TensorFlow
Model serving
Inference optimization
GPU acceleration

Education

BS/MS/PhD in CS/Engineering

Tools

CUDA
Linux

Job description

Join one of the most sophisticated trading firms in the world as they build a new machine learning initiative from the ground up. The team is currently just three people, including engineers from two competitors and a researcher from a top FAANG AI lab, and they're looking to add a few key hires who will help define the future of ML infrastructure at the firm.

This particular hire will focus on model serving, inference optimization, and deploying large-scale models onto GPU infrastructure. The team is already operating at scale with over 1,000 GPUs today and is rapidly expanding toward 10,000+ GPUs, creating unique engineering challenges around performance, efficiency, and hardware utilization.

Rather than joining a large, established ML organization, you'll have significant ownership over architecture, technical direction, and platform decisions from day one.

What You Will Do
  • Design and build large-scale serving and inference infrastructure for state-of-the-art ML models.
  • Optimize model performance, latency, throughput, and GPU utilization in production environments.
  • Deploy and scale models across thousands of GPUs while improving efficiency and reliability.
  • Develop model compression, quantization, distillation, and other optimization techniques to maximize hardware performance.
  • Work on taking cutting-edge models from research into production by bringing them closer to the hardware.
  • Partner closely with researchers and engineers to accelerate experimentation and deployment.
  • Help define the architecture and roadmap for a rapidly growing ML platform.
What You Bring
  • 8+ years of software engineering, systems engineering, or machine learning infrastructure experience.
  • Deep experience with model serving, inference, and performance optimization at scale.
  • Strong understanding of model compression, quantization, kernel optimization, and GPU acceleration.
  • Strong knowledge of CUDA, GPU programming, and hardware-aware optimization.
  • Expert-level programming skills in C++ and Python.
  • Experience with modern ML frameworks such as PyTorch, JAX, or TensorFlow.
  • BS, MS, or PhD in Computer Science, Engineering, Mathematics, or a related field.
Why Consider It
  • Ground-floor opportunity to build a new ML platform inside one of the world's leading trading firms.
  • Work alongside engineers and researchers from top AI labs and elite quantitative trading firms.
  • Ownership over critical infrastructure supporting next-generation ML workloads.
  • Massive scale, with GPU infrastructure growing from the low thousands to 10s of thousands
  • Solve challenging problems at the intersection of distributed systems, machine learning, and high-performance computing.

This role can sit out of NYC or Chicago

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Machine Learning Engineer (Training & Inference Systems)
Machine Learning Engineer (Training & Inference Systems)

Fintal Partners • New York (NY)

On-site
USD 185,000 - 230,000
Machine Learning Engineer (Training & Inference Systems)
Machine Learning Engineer (Training & Inference Systems)

Jobzhr • New York (NY)

On-site
USD 180,000 - 280,000
Applied ML Systems Engineer - Finance
Applied ML Systems Engineer - Finance

Park Lane Recruitment • New York (NY)

On-site
USD 250,000 - 350,000
Relocation support
Sponsorship available
Elite total compensation
+5
Applied ML Systems Engineer – Finance
Applied ML Systems Engineer – Finance

Park Lane Recruitment • New York (NY)

On-site
USD 250,000 - 350,000
401(k) matching
Medical coverage
Wellness reimbursement
+2
Principal ML Infrastructure Engineer (Relocation Available)
Principal ML Infrastructure Engineer (Relocation Available)

Franklin Fitch • Dallas (TX)

On-site
USD 100,000 - 140,000
Senior ML Performance Engineer
Senior ML Performance Engineer

well-funded deeptech startup • California (MO)

On-site
USD 200,000 - 250,000
ML Architect
ML Architect

Blue Signal Search • United States

On-site
USD 180,000 - 280,000
Health insurance
Dental insurance
Life insurance
+1
Machine Learning Performance Engineer
Machine Learning Performance Engineer

Long Ridge Partners • New York (NY)

On-site
USD 600,000 - 1,500,000
Hybrid working options
Free meals (breakfast, lunch)
Wellness reimbursement
+2
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Machine Learning Hardware Engineer
Machine Learning Hardware Engineer

Fintal Partners • New York (NY)

On-site
USD 140,000 - 190,000