Staff Engineer - LLM Inference & Serving at Scale

Prime Intellect

San Francisco (CA)

Hybrid

USD 150,000 - 300,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Cash compensation range of $150-300k
Flexible work arrangement (remote or San Francisco office)
Full visa sponsorship and relocation support
Professional development budget
Regular team off-sites and conference attendance
Opportunity to shape decentralized AI and RL

Job summary

Prime Intellect is looking for a skilled ML Systems Engineer to build and optimize LLM serving infrastructure and inference systems. This hybrid role involves contributing to the scalability of their reinforcement learning training. Successful candidates will have over 3 years of experience in building ML services, strong knowledge of Python and cloud platforms, and a desire to work on cutting-edge AI infrastructure. They offer a cash compensation range of $150-300k with equity and full relocation support.

Qualifications

  • 3+ years building and running large-scale ML/LLM services with clear latency/availability SLOs.
  • Hands-on with vLLM, SGLang, TensorRT‑LLM.
  • Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
  • Deep understanding of prefill vs. decode, KV-cache behavior, batching, and speculative decoding.
  • Comfortable debugging CUDA/NCCL, drivers/kernels, and storage.

Responsibilities

  • Build infrastructure to serve LLMs efficiently at scale.
  • Optimize and integrate inference systems into our RL training stack.
  • Design placement and scheduling algorithms for heterogeneous accelerators.
  • Implement multi-region/zone failover and traffic shifting.
  • Profile kernels, memory bandwidth and transport; apply quantization techniques.

Skills

Building ML Systems at Scale
Inference Backends
Distributed Serving Infra
Inference Internals
Full-Stack Debugging
Python
PyTorch
Cloud & Automation
Kubernetes
GPU & Networking

Job description

Prime Intellect is looking for a skilled ML Systems Engineer to build and optimize LLM serving infrastructure and inference systems. This hybrid role involves contributing to the scalability of their reinforcement learning training. Successful candidates will have over 3 years of experience in building ML services, strong knowledge of Python and cloud platforms, and a desire to work on cutting-edge AI infrastructure. They offer a cash compensation range of $150-300k with equity and full relocation support.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000
Remote ML Engineering Manager: LLM Serving & Infra
Remote ML Engineering Manager: LLM Serving & Infra

Jobgether • United States

Remote
USD 176,000 - 252,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range of $150-300k
Flexible work arrangement (remote or San Francisco office)
Full visa sponsorship and relocation support
+3
Founding ML Infra Engineer — Production-Grade LLMs
Founding ML Infra Engineer — Production-Grade LLMs

Realmlabs • Sunnyvale (CA)

On-site
USD 210,000 - 350,000
Market aligned compensation
Founding engineer equity
Medical, Dental, Vision, and Life insurance
+2
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Senior Cloud AI LLM Serving Engineer
Senior Cloud AI LLM Serving Engineer

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Competitive annual discretionary bonus program
Potential RSU grants
Comprehensive benefits package
ML Systems Engineer: Distributed LLM Training & Inference
ML Systems Engineer: Distributed LLM Training & Inference

Scale AI • Seattle (WA), New York (NY), San Francisco (CA)

On-site
USD 200,000 - 251,000
Comprehensive health coverage
Equity-based compensation
Retirement benefits
+3
Tech Lead Manager — LLM Training Platform
Tech Lead Manager — LLM Training Platform

Scale AI, Inc. • New York (NY)

On-site
USD 264,000 - 331,000
Health, dental and vision coverage
Retirement benefits
Learning and development stipend
+2
Staff ML Systems Engineer - Diffusion LLM Serving
Staff ML Systems Engineer - Diffusion LLM Serving

Inception • San Francisco (CA)

On-site
USD 180,000 - 240,000