Senior ML Infra Engineer - Remote, Scalable Inference

AZX

Seattle (WA)

Remote

USD 180,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
Equity
Flexible paid time off
Bonus eligibility
Fully remote culture with Seattle team

Job summary

AZX is seeking an ML Engineer to own the technical backbone of model serving at scale, shaping inference platforms, GPU scheduling, and autoscaling across cloud and customer-managed clusters for vLLM/SGLang models.

You’ll set standards, guide evaluation systems, and mentor engineers in a high-impact, architecture-focused role that balances cost, performance, and reliability for critical AI workloads.

Qualifications

  • 3+ years of experience with ML infrastructure and inference serving at production scale.
  • Strong background in evaluation and reliability engineering for ML/LLM systems.
  • Solid Kubernetes experience, GPU scheduling constraints (node pools, autoscaling).
  • Track record of technical leadership at staff/senior level.
  • Research fluency (PhD/publications) is a plus but infra-first role.

Responsibilities

  • Own architecture for inference serving and GPU scheduling — Kubernetes operators, autoscaling, and dynamic capacity across vLLM/SGLang.
  • Design and calibrate eval systems for model, prompt, and agent changes, including golden datasets and CI integration.
  • Advise on cost-aware model routing and cascading decisions, balancing latency, cost, and quality.
  • Apply physics-informed ML and enterprise AI expertise to client and platform problems.
  • Set technical standards for ML infrastructure and evaluation practice; mentor engineers.
  • Partner with related teams to keep architecture coherent as the platform grows.

Skills

ML infrastructure
Kubernetes
GPU scheduling
Staff/senior leadership
Evaluation systems
Reliability engineering

Tools

vLLM
SGLang
TensorRT-LLM
PyTorch
CUDA

Job description

AZX is seeking an ML Engineer to own the technical backbone of model serving at scale, shaping inference platforms, GPU scheduling, and autoscaling across cloud and customer-managed clusters for vLLM/SGLang models.

You’ll set standards, guide evaluation systems, and mentor engineers in a high-impact, architecture-focused role that balances cost, performance, and reliability for critical AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infra Engineer — Remote AI Serving
Senior ML Infra Engineer — Remote AI Serving

United States Digital Space LLC • United States

Remote
USD 90,000 - 150,000
Senior ML Infrastructure Engineer — Remote
Senior ML Infrastructure Engineer — Remote

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 105,000 - 143,000
Staff AI Inference Engineer — Remote/Hybrid
Staff AI Inference Engineer — Remote/Hybrid

Zoom • Seattle (WA), Northern (KY)

Hybrid
USD 152,000 - 332,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Remote ML Platform Engineer — Scalable Inference
Remote ML Platform Engineer — Scalable Inference

Bright Vision Technologies • Framingham (MA)

On-site
USD 100,000 - 160,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Remote ML Systems Engineer — Scalable Inference & GPU Ops
Remote ML Systems Engineer — Scalable Inference & GPU Ops

Bright Vision Technologies • Euless (TX), Bedford (TX)

On-site
USD 145,000 - 165,000
Remote ML Systems Engineer: Scalable AI Inference
Remote ML Systems Engineer: Scalable AI Inference

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Staff ML Infra Engineer: Distributed Training & Inference
Staff ML Infra Engineer: Distributed Training & Inference

Jobtailor • Boston (MA)

On-site
USD 120,000 - 160,000
Remote MLOps Engineer - Scalable AI Inference Platform
Remote MLOps Engineer - Scalable AI Inference Platform

Bright Vision Technologies • Sammamish (WA)

On-site
USD 100,000 - 150,000