Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC

Bellevue (KY)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) plan with company match
Paid holidays

Job summary

Designworks Talent LLC is seeking a Staff Inference Engineer to build and operate the model-serving systems powering a next-generation AI inference platform. You will work on high-throughput, low-latency inference for production-scale APIs in a hybrid Bellevue, WA setting.

The role emphasizes optimizing GPU-backed workloads, collaborating with training and platform teams, and contributing to scalable, reliable production services across diverse model architectures.

Qualifications

  • Experience building and operating production ML inference systems at scale.
  • Understanding latency, throughput, memory usage, and cost trade-offs for large models.
  • Experience designing reliable distributed systems or production infrastructure.
  • Understanding GPU-backed AI workloads and scaling inference.
  • Strong engineering fundamentals and ability to own complex problems.
  • Comfortable in fast-moving environments where systems are built from the ground up.

Responsibilities

  • Build and operate production-grade model-serving and inference systems for high-throughput AI workloads.
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost across models.
  • Design systems to maximize GPU utilization with predictable performance.
  • Improve scalability and operational maturity of inference platforms as demand grows.
  • Collaborate with AI training, GPU performance, orchestration, and infra teams for smooth transitions from development to production serving.
  • Develop monitoring, alerting, and operational practices for reliable inference services.
  • Investigate and resolve performance, reliability, and capacity challenges across inference workloads.
  • Contribute to architecture decisions and engineering standards as the platform evolves.

Skills

Production ML inference
Distributed systems
GPU computing
Performance optimization
LLM inference
Cloud infrastructure

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
Kubernetes
GPU scheduling

Job description

Designworks Talent LLC is seeking a Staff Inference Engineer to build and operate the model-serving systems powering a next-generation AI inference platform. You will work on high-throughput, low-latency inference for production-scale APIs in a hybrid Bellevue, WA setting.

The role emphasizes optimizing GPU-backed workloads, collaborating with training and platform teams, and contributing to scalable, reliable production services across diverse model architectures.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Inference Engineer
Staff Inference Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Staff Engineer, Inference Runtime — High-Performance AI Serving
Staff Engineer, Inference Runtime — High-Performance AI Serving

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Hybrid Staff AI Training Infrastructure Engineer
Hybrid Staff AI Training Infrastructure Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Principal Data Center Ops Engineer for AI Infra
Principal Data Center Ops Engineer for AI Infra

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 120,000 - 170,000
Medical Insurance
401(k) Match
Paid Holidays
+1
High-Performance AI Inference Platform Engineer
High-Performance AI Inference Platform Engineer

DRW Holdings, LLC • Chicago (IL)

On-site
USD 200,000 - 250,000
Medical Insurance
Dental Insurance
Vision Insurance
+5
Senior AI Inference Platform Engineer
Senior AI Inference Platform Engineer

Tradermath • Chicago (IL), Northern (KY)

Hybrid
USD 200,000 - 250,000
Medical insurance
Dental insurance
Vision insurance
+1
Staff AI Platform Engineer — Inference & Scale Leader
Staff AI Platform Engineer — Inference & Scale Leader

Greenhouse Software, Inc. • New York (NY), Northern (KY)

Hybrid
USD 220,000 - 330,000
Applied AI Researcher: Inference & Systems Expert
Applied AI Researcher: Inference & Systems Expert

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 280,000
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
INFERENCE ENGINEER
INFERENCE ENGINEER

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000