ML Systems Engineer — LLM Inference & Multi-Node Performance

ScOp Venture Capital LLC.

Santa Barbara, Northern (CA, KY)

Hybrid

USD 150,000 - 210,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Competitive compensation
Equity
Professional growth opportunities
Access to GPU compute resources

Job summary

ChipAgents seeks an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering its agentic AI platform. This role centers on low-level systems optimization and building scalable multi-node clusters for training and inference that push LLM throughput and latency limits.

You will implement performance optimizations, build evaluation harnesses, and collaborate with researchers to integrate new model architectures into production infrastructure.

Qualifications

  • 3+ years of experience with large-scale ML systems and GPU inference optimization.
  • Strong Python and C++/CUDA skills; hands-on experience with SGLang, vLLM, PyTorch or similar frameworks.
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
  • Experience deploying and optimizing LLMs in production: model serving, batching, distributed inference, or quantization.
  • Strong systems-level debugging and profiling skills; comfortable across CUDA kernels to application logic.
  • Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.
  • Self-directed problem solver with a passion for ambitious optimization challenges.

Responsibilities

  • Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.
  • Implement and benchmark concrete inference optimizations.
  • Profile and analyze inference bottlenecks from GPU kernel execution to memory bandwidth constraints.
  • Build robust evaluation harnesses and benchmarking frameworks measuring accuracy, throughput, latency, and resource utilization across parallelism strategies.
  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
  • Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance.

Skills

Python
C++/CUDA
GPU architecture
ML systems
Profiling
Distributed computing
Model serving
SGLang
vLLM
PyTorch
CUDA kernels
Ray

Education

Bachelor/Master/PhD in CS/EE or related

Tools

SGLang
vLLM
PyTorch

Job description

ChipAgents seeks an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering its agentic AI platform. This role centers on low-level systems optimization and building scalable multi-node clusters for training and inference that push LLM throughput and latency limits.

You will implement performance optimizations, build evaluation harnesses, and collaborate with researchers to integrate new model architectures into production infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ML Systems Engineer
ML Systems Engineer

ScOp Venture Capital LLC. • Santa Barbara (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Competitive compensation
Equity
Professional growth opportunities
+1
Senior AI Inference Engineer: High-Throughput LLMs
Senior AI Inference Engineer: High-Throughput LLMs

Crusoe Energy Systems • San Francisco (CA)

On-site
USD 180,000 - 240,000
Health benefits
401(k) match
Paid time off
+1
LLM Inference Systems Performance Engineer
LLM Inference Systems Performance Engineer

3M HEALTHCARE • Austin (TX)

On-site
USD 150,000 - 210,000
Medical, dental, and vision coverage
Income protection benefits
Paid family leave
+1
Senior AI Systems Engineer: LLM Inference & Optimization
Senior AI Systems Engineer: LLM Inference & Optimization

Showcify • United States

Remote
USD 180,000 - 240,000
ML Systems Engineer: Distributed LLM Training & Inference
ML Systems Engineer: Distributed LLM Training & Inference

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 200,800 - 251,000
Comprehensive health coverage
Equity-based compensation
Retirement benefits
+3
LLM Inference Systems Engineer
LLM Inference Systems Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior ML Engineer: AI Inference & Performance Optimizer
Senior ML Engineer: AI Inference & Performance Optimizer

Nebius • Palo Alto (CA)

Hybrid
USD 195,000 - 263,000
Health insurance
401(k) plan
Parental leave
+2
LLM Systems Engineer & Research
LLM Systems Engineer & Research

Scale AI, Inc. • New York (NY)

On-site
USD 189,000 - 237,000
Comprehensive health coverage
Dental and vision coverage
Retirement benefits
+3
ML Inference Performance Engineer — Optimize Cost & Latency
ML Inference Performance Engineer — Optimize Cost & Latency

Adaption Labs • San Francisco (CA)

On-site
USD 180,000 - 260,000
Flexible work
Adaption Passport
Lunch stipend
+1
Senior AI Inference Engineer — Production LLM Optimizer
Senior AI Inference Engineer — Production LLM Optimizer

Crusoe Energy Systems LLC • San Francisco (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Competitive compensation
Equity packages
Health/dental/vision insurance
+1