Senior Inference Platform Engineer

Designworks Talent

Bellevue (WA)

Hybrid

USD 180,000 - 240,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Hybrid work model
Office Bellevue
Competitive compensation

Job summary

Designworks Talent is seeking an Inference Engineer to build and operate production-grade model-serving and inference systems in a hybrid Bellevue, WA setting. You’ll work across distributed systems, GPU optimization, and production AI infrastructure to deliver high-throughput, low-latency inference.

Join a fast-moving team shaping next-gen AI platforms with scalable tooling, monitoring, and reliability practices, collaborating with AI training and platform teams.

Qualifications

  • Experience building and operating production ML inference or model-serving systems at scale.
  • Understanding of latency, throughput, memory utilization, and cost efficiency in serving large AI models.
  • Experience designing reliable distributed systems or production infrastructure.
  • Familiarity with GPU-backed AI workloads and scaling inference systems.

Responsibilities

  • Build and operate production-grade model-serving and inference systems for high-throughput AI workloads.
  • Optimize inference infrastructure for latency, throughput, and cost efficiency across workloads.
  • Design systems to maximize GPU utilization with predictable performance and reliability.
  • Improve scalability and maturity of inference platforms as demand grows.
  • Collaborate with AI training and platform teams for smooth transitions from development to production.
  • Develop monitoring and alerting to maintain reliable inference services.
  • Investigate performance, reliability, and capacity challenges in inference workloads.

Skills

Production ML inference
Distributed systems
GPU workloads understanding
Problem ownership
Fast-moving environment

Tools

vLLM
TensorRT-LLM
Triton Inference Server
Kubernetes
Cloud infrastructure

Job description

Designworks Talent is seeking an Inference Engineer to build and operate production-grade model-serving and inference systems in a hybrid Bellevue, WA setting. You’ll work across distributed systems, GPU optimization, and production AI infrastructure to deliver high-throughput, low-latency inference.

Join a fast-moving team shaping next-gen AI platforms with scalable tooling, monitoring, and reliability practices, collaborating with AI training and platform teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Inference Engineer
Inference Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Office Bellevue
Competitive compensation
Head of AI Inference Platforms & Optimizations
Head of AI Inference Platforms & Optimizations

DigitalOcean • Seattle (WA)

Hybrid
USD 274,000 - 343,000
Staff Engineer, Inference Runtime — High-Performance AI Serving
Staff Engineer, Inference Runtime — High-Performance AI Serving

Anthropic • Seattle (WA)

Hybrid
USD 405,000 - 485,000
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Senior AI Inference Platform Lead
Senior AI Inference Platform Lead

CoreWeave • Bellevue (WA)

On-site
USD 188,000 - 303,000
Medical, dental, and vision insurance
401(k) with generous employer match
Paid Parental Leave
+3
Inference Platform Backend Engineer (Equity & Benefits)
Inference Platform Backend Engineer (Equity & Benefits)

Together • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Health insurance
Competitive compensation
Senior AI Infra Engineer: High-Performance Inference
Senior AI Infra Engineer: High-Performance Inference

Ddn • Sacramento (CA)

On-site
USD 140,000 - 200,000
Senior Systems Engineer, AI Inference Infra
Senior Systems Engineer, AI Inference Infra

United States Digital Space LLC • United States

Remote
USD 150,000 - 190,000
Senior Inference Platform Engineer - Data Center
Senior Inference Platform Engineer - Data Center

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5