Staff Inference Engineer

Designworks Talent LLC

Bellevue (KY)

Hybrid

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
401(k) plan with company match
Paid holidays

Job summary

Designworks Talent LLC is seeking a Staff Inference Engineer to build and operate the model-serving systems powering a next-generation AI inference platform. You will work on high-throughput, low-latency inference for production-scale APIs in a hybrid Bellevue, WA setting.

The role emphasizes optimizing GPU-backed workloads, collaborating with training and platform teams, and contributing to scalable, reliable production services across diverse model architectures.

Qualifications

  • Experience building and operating production ML inference systems at scale.
  • Understanding latency, throughput, memory usage, and cost trade-offs for large models.
  • Experience designing reliable distributed systems or production infrastructure.
  • Understanding GPU-backed AI workloads and scaling inference.
  • Strong engineering fundamentals and ability to own complex problems.
  • Comfortable in fast-moving environments where systems are built from the ground up.

Responsibilities

  • Build and operate production-grade model-serving and inference systems for high-throughput AI workloads.
  • Optimize inference infrastructure for token throughput, latency, scalability, and cost across models.
  • Design systems to maximize GPU utilization with predictable performance.
  • Improve scalability and operational maturity of inference platforms as demand grows.
  • Collaborate with AI training, GPU performance, orchestration, and infra teams for smooth transitions from development to production serving.
  • Develop monitoring, alerting, and operational practices for reliable inference services.
  • Investigate and resolve performance, reliability, and capacity challenges across inference workloads.
  • Contribute to architecture decisions and engineering standards as the platform evolves.

Skills

Production ML inference
Distributed systems
GPU computing
Performance optimization
LLM inference
Cloud infrastructure

Tools

vLLM
SGLang
TensorRT-LLM
Triton Inference Server
Kubernetes
GPU scheduling

Job description

Staff Inference Engineer

Location: Hybrid | Bellevue, WA (downtown)**Titles:**Senior and Staff (multiple roles available)

Build the Inference Platform Powering Next-Generation AI Applications

About the Opportunity

A well-funded, rapidly growing AI infrastructure company is building a next-generation cloud platform designed to power the full lifecycle of artificial intelligence. The organization is developing a comprehensive AI infrastructure, platform, and services portfolio that supports the full spectrum of AI workloads---including large-scale compute, model training, fine-tuning, inference, and emerging agentic AI applications.

Backed by significant long-term investment, the company combines the speed, ownership, and innovation of a startup with the stability and resources of an established parent organization. Engineering teams are intentionally lean, highly collaborative, and AI-native, leveraging modern tooling and automation to build infrastructure capable of supporting the industry's most demanding AI workloads.

We're seeking Inference Engineers to build and operate the model-serving systems behind a next-generation AI inference platform. This team focuses on delivering high-throughput, low-latency, reliable inference experiences that enable customers to consume advanced AI capabilities through production-scale APIs.

The Opportunity

This is a foundational engineering role focused on building the systems that bring AI models from research environments into reliable production services. You'll work on the infrastructure layer responsible for serving large models efficiently, optimizing performance, and ensuring reliability as usage scales.

You'll collaborate closely with GPU performance, AI training infrastructure, platform engineering, and operations teams to solve complex challenges around model serving, latency optimization, resource efficiency, and production reliability.

This opportunity is ideal for engineers who enjoy working at the intersection of distributed systems, machine learning infrastructure, GPU computing, and large-scale production systems.

What You'll Do
  • Build and operate production-grade model-serving and inference systems supporting high-throughput, low-latency AI workloads.

  • Optimize inference infrastructure for token throughput, latency, scalability, and cost efficiency across different model architectures and workloads.

  • Design systems that maximize GPU utilization while maintaining predictable performance and reliability.

  • Improve the scalability and operational maturity of inference platforms as customer demand grows.

  • Partner with AI training, GPU performance, orchestration, and infrastructure teams to ensure smooth transitions from model development to production serving.

  • Develop monitoring, alerting, and operational practices to maintain reliable inference services.

  • Investigate and resolve performance, reliability, and capacity challenges across inference workloads.

  • Contribute to architecture decisions and engineering standards as the platform evolves.

What We're Looking For
  • Experience building and operating production machine learning inference or model-serving systems at scale.

  • Strong understanding of the performance trade-offs involved in serving large AI models, including latency, throughput, memory utilization, and cost efficiency.

  • Experience designing reliable distributed systems or production infrastructure.

  • Understanding of GPU-backed AI workloads and the challenges of scaling inference systems.

  • Strong engineering fundamentals and the ability to independently own complex technical problems.

  • Comfortable working in a fast-moving environment where systems and processes are being built from the ground up.

Preferred Qualifications
  • Experience with modern inference-serving frameworks such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar technologies.

  • Experience optimizing LLM inference workloads or large-scale AI serving platforms.

  • Background operating API-based AI products or high-volume production services.

  • Experience with GPU scheduling, distributed systems, Kubernetes, or cloud infrastructure platforms.

  • Familiarity with model optimization techniques such as quantization, batching, caching, or performance tuning.

  • Experience working at a hyperscaler, AI lab, GPU cloud provider, or large-scale ML infrastructure organization.

Compensation
  • Competitive base pay for Bellevue market

  • Certain roles are eligible for additional rewards, including merit increases, annual bonus, and long term incentives. These awards are allocated based on individual performance

  • U.S. based employees have access to medical, dental, and vision insurance, a 401(k) plan and company match, employees also receive per calendar year, paid holidays.

Location
  • Hybrid role based in the Bellevue, WA area.

  • Approximately three days per week in the office.

  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

  • U.S. work authorization is required. Visa sponsorship is not currently available.

Why Join?
  • Build the inference platform powering the next generation of AI applications.

  • Work directly on large-scale model serving, GPU optimization, and production AI systems.

  • Solve complex challenges around latency, throughput, reliability, and cost efficiency.

  • Join early enough to influence architecture, tooling, and engineering practices.

  • Collaborate with a highly experienced team building critical AI infrastructure from the ground up.

  • Enjoy the ownership and technical impact of a startup environment backed by significant long-term investment.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Training Infrastructure Engineer
Senior AI Training Infrastructure Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 180,000 - 260,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Staff AI Training Infrastructure Engineer
Staff AI Training Infrastructure Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 150,000 - 210,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Applied Researcher – AI Expert
Applied Researcher – AI Expert

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 280,000
Applied Researcher – AI Expert
Applied Researcher – AI Expert

Designworks Talent • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Staff Principal Data Center Operations and Maintenance Engineer
Staff Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 120,000 - 170,000
Medical Insurance
401(k) Match
Paid Holidays
+1
Staff GPU Performance / Kernel Engineer
Staff GPU Performance / Kernel Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Medical, dental, and vision insurance
401(k) plan with company match
Paid holidays
Staff Principal Data Center Operations and Maintenance Engineer
Staff Principal Data Center Operations and Maintenance Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 130,000 - 200,000
Medical, dental, vision insurance
401(k) with company match
Paid holidays
Senior GPU Performance / Kernel Engineer
Senior GPU Performance / Kernel Engineer

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 170,000 - 250,000
Hybrid work model
Medical, dental, vision insurance
Staff Data Center Infrastructure Software Engineer
Staff Data Center Infrastructure Software Engineer

Designworks Talent LLC • Bellevue (WA)

Hybrid
USD 140,000 - 200,000
Merit-based bonus
Long-term incentives
Medical, dental, vision