Staff Engineer, LLM Inference & Infra

Prime Intellect

United States

Hybrid

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Flexible work arrangement
Full visa sponsorship
Professional development budget
Regular team off-sites

Job summary

A leading AI infrastructure firm is seeking a talented engineer to build and optimize a multi-tenant LLM serving platform at scale. Candidates should have over 3 years of experience in ML systems, with strong skills in tools like PyTorch and a solid understanding of inference optimization. This role offers flexible work arrangements and professional development opportunities for those passionate about democratizing AI development.

Qualifications

  • 3+ years building and running large-scale ML/LLM services.
  • Hands-on with vLLM, SGLang, or TensorRT-LLM.
  • Deep understanding of inference mechanics and performance optimization.

Responsibilities

  • Build multi-tenant LLM serving platform across cloud GPU fleets.
  • Design scheduling algorithms for heterogeneous accelerators.
  • Profile kernels and optimize memory management for maximum performance.

Skills

Building ML Systems at Scale
Inference Backends
Full-Stack Debugging
Python
Kubernetes

Tools

PyTorch
CUDA
TensorRT

Job description

A leading AI infrastructure firm is seeking a talented engineer to build and optimize a multi-tenant LLM serving platform at scale. Candidates should have over 3 years of experience in ML systems, with strong skills in tools like PyTorch and a solid understanding of inference optimization. This role offers flexible work arrangements and professional development opportunities for those passionate about democratizing AI development.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - LLM Inference & Serving at Scale
Staff Engineer - LLM Inference & Serving at Scale

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Founding ML Infra Engineer — Production-Grade LLMs
Founding ML Infra Engineer — Production-Grade LLMs

Realmlabs • Sunnyvale (CA)

On-site
USD 210,000 - 350,000
Market aligned compensation
Founding engineer equity
Medical, Dental, Vision, and Life insurance
+2
Staff Software Engineer: LLM Inference & GPU Infra (Equity)
Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
Senior ML Engineer, AI Platform & Products (LLM)
Senior ML Engineer, AI Platform & Products (LLM)

BetterUp • New York (NY)

Hybrid
USD 200,000 - 275,000
Senior AI Inference Infrastructure Engineer
Senior AI Inference Infrastructure Engineer

Modular • United States

Hybrid
USD 167,000 - 273,000
LLM Serving Engineer, Cloud AI Platform Architect
LLM Serving Engineer, Cloud AI Platform Architect

Qualcomm • Austin (TX)

On-site
USD 158,000 - 238,000
Remote ML Engineering Manager: LLM Serving & Infra
Remote ML Engineering Manager: LLM Serving & Infra

Jobgether • United States

Remote
USD 176,000 - 252,000
Staff Software Engineer- Foundation Model Inference
Staff Software Engineer- Foundation Model Inference

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity
Tech Lead Manager for Scalable LLM Training Platform
Tech Lead Manager for Scalable LLM Training Platform

Scale AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior Cloud AI LLM Serving Engineer
Senior Cloud AI LLM Serving Engineer

Qualcomm • San Diego (CA)

On-site
USD 158,000 - 238,000
Competitive annual discretionary bonus program
Potential RSU grants
Comprehensive benefits package