Senior ML Infra Engineer — Scalable Inference & Platform

Unity Technologies

Seattle (WA)

On-site

USD 187,000 - 243,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Unity Technologies in Seattle is seeking a Senior ML Engineer to design and evolve Unity Vector’s online model inference platform. You will build reliable production infrastructure for serving ML models with low latency and high reliability, and enable safe experimentation at scale.

You will work closely with ML engineers, platform teams, and product stakeholders to ensure models can be deployed, scaled, monitored, and iterated on efficiently.

Qualifications

  • Experience building and operating production-grade online ML inference systems.
  • Experience with model serving frameworks such as Triton, TorchServe, Ray Serve, or TensorFlow Serving.
  • Experience optimizing inference workloads (dynamic batching, quantization, GPU acceleration, kernel tuning).
  • Strong experience with distributed systems, Kubernetes, autoscaling, and production observability.
  • Strong Python programming skills for production ML systems.
  • Experience with PyTorch and model packaging, validation, and serving lifecycle management.
  • Experience designing infrastructure for safe model rollout, canary testing, and automated rollback.
  • Ability to reason about latency, throughput, reliability, and cost tradeoffs in online systems.

Responsibilities

  • Design and operate large-scale online inference infrastructure with low latency and high reliability.
  • Develop infrastructure supporting distributed training workflows (e.g., PyTorch, Ray Data, Ray Train).
  • Integrate ML pipelines with workflow orchestration (Flyte, Airflow) for multi-stage training workflows.
  • Optimize model performance via compilation, GPU/CPU utilization, scheduling, and kernel fusion.
  • Improve observability of ML systems (latency, throughput, error-rate, cost, model-health).
  • Collaborate with ML engineers to accelerate model iteration while ensuring safety and scalability.
  • Improve reliability and reproducibility of serving workflows (packaging, validation, deployment automation).
  • Lead architectural improvements to make the platform robust and cost-efficient.

Skills

Production ML infra
Distributed systems
System design
Observability
Performance optimization
Leadership (influence)

Tools

NVIDIA Triton Inference Server
TorchServe
Ray Serve
TensorFlow Serving
Kubernetes
Python
PyTorch
Flyte/Airflow

Job description

Unity Technologies in Seattle is seeking a Senior ML Engineer to design and evolve Unity Vector’s online model inference platform. You will build reliable production infrastructure for serving ML models with low latency and high reliability, and enable safe experimentation at scale.

You will work closely with ML engineers, platform teams, and product stakeholders to ensure models can be deployed, scaled, monitored, and iterated on efficiently.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer — Online Inference at Scale
Senior ML Infrastructure Engineer — Online Inference at Scale

Socket.dev • Washington

Hybrid
USD 187,000 - 243,000
Equity awards
Health insurance
Retirement plans
+1
Remote Senior ML Infra Engineer — Real-Time Model Serving
Remote Senior ML Infra Engineer — Real-Time Model Serving

3M HEALTHCARE • Mountain View (CA)

On-site
USD 183,000 - 249,000
Health insurance
Stock options
Retirement plan
+1
Senior Real-Time ML Infrastructure Engineer
Senior Real-Time ML Infrastructure Engineer

3M HEALTHCARE • Bellevue (WA)

Remote
USD 183,000 - 249,000
Senior ML Engineer, Data Infrastructure & Pipelines
Senior ML Engineer, Data Infrastructure & Pipelines

Unity Enterprise • Mountain View (CA)

Hybrid
USD 200,000 - 261,000
Comprehensive health insurance
Life and disability insurance
Employee stock ownership
+3
Senior ML Data Infrastructure Engineer — Scalable Pipelines
Senior ML Data Infrastructure Engineer — Scalable Pipelines

Unity • Mountain View (CA)

On-site
USD 200,000 - 261,000
Comprehensive health insurance
Employee stock ownership
Competitive retirement plans
+4
Senior ML Engineer, Data Infrastructure & Pipelines
Senior ML Engineer, Data Infrastructure & Pipelines

Unity Technologies • Mountain View (CA)

On-site
USD 200,000 - 261,000
Health insurance
Stock options
Retirement plan
+3
Offline ML Infrastructure Engineer for Scalable Pipelines
Offline ML Infrastructure Engineer for Scalable Pipelines

3M HEALTHCARE • Mountain View (CA)

On-site
USD 112,000 - 163,000
Employee stock ownership
Commuter subsidy
Generous vacation and personal days
+1
Senior Machine Learning Engineer, ML Infrastructure- Online
Senior Machine Learning Engineer, ML Infrastructure- Online

Socket.dev • Washington

Hybrid
USD 187,000 - 243,000
Equity awards
Health insurance
Retirement plans
+1
Remote ML Systems Engineer: Scalable AI Inference
Remote ML Systems Engineer: Scalable AI Inference

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior ML Inference Engineer – Platform
Senior ML Inference Engineer – Platform

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000