Kubernetes ML Inference Engineer: Model Serving

Abridge

San Francisco (CA)

On-site

USD 221,000 - 260,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

14 paid holidays
Flexible PTO
Health plans
HSA contributions
Parental leave
401(k) matching
Personal device allowance
FSA and commuter benefits
Sabbatical after 5 years
Equity

Job summary

Abridge is seeking an ML Infrastructure Engineer, Model Inference in San Francisco to build and optimize the core inference infrastructure powering our AI‑driven healthcare solutions. You will collaborate across Infrastructure and Research to deploy, optimize, and orchestrate AI models at scale.

Ideal candidates have 2+ years of production ML infrastructure experience, strong Kubernetes know‑how, and a track record of engineering scalable APIs and distributed systems for real-time workloads.

Qualifications

  • 2+ years building and deploying machine learning models in production.
  • Deep understanding of container orchestration and distributed systems architecture.
  • Experience developing APIs and managing distributed systems for both batch and real-time workloads.

Responsibilities

  • Design, deploy and maintain scalable Kubernetes clusters for AI model inference and training.
  • Develop, optimize, and maintain ML model serving infrastructure, ensuring high-performance and low-latency.
  • Collaborate with ML and product teams to scale backend infrastructure for AI-driven products, focusing on model deployment, throughput optimization, and compute efficiency.
  • Optimize compute-heavy workflows and enhance GPU utilization for ML workloads.
  • Build a robust model API orchestration system.
  • Collaborate with leadership to define and implement strategies for scaling infrastructure as the company grows, ensuring long-term efficiency and performance.

Skills

Kubernetes administration
APIs development
Distributed systems
Model inference
Compute optimization
Collaboration with ML/product teams

Tools

NVIDIA Triton Server
VLLM
TRT-LLM
Terraform
Ansible
GitOps

Job description

Abridge is seeking an ML Infrastructure Engineer, Model Inference in San Francisco to build and optimize the core inference infrastructure powering our AI‑driven healthcare solutions. You will collaborate across Infrastructure and Research to deploy, optimize, and orchestrate AI models at scale.

Ideal candidates have 2+ years of production ML infrastructure experience, strong Kubernetes know‑how, and a track record of engineering scalable APIs and distributed systems for real-time workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer - Model Inference & Scale
ML Infrastructure Engineer - Model Inference & Scale

Abridge • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Generous Time Off
Comprehensive Health Plans
401(k) Matching
+2
ML Infra Tech Lead: Scalable Training & Inference
ML Infra Tech Lead: Scalable Training & Inference

Reducto • San Francisco (CA)

On-site
USD 180,000 - 260,000
Unlimited PTO
Daily Lunch
Commuter Reimbursement
+3
ML Inference Infrastructure Engineer
ML Inference Infrastructure Engineer

Baseten • San Francisco (CA)

On-site
USD 90,000 - 130,000
100% coverage of medical, dental, and vision insurance
Generous PTO policy including Winter Break
Company-facilitated 401(k)
Senior ML Inference Engineer: Production Systems
Senior ML Inference Engineer: Production Systems

MakerMaker • San Francisco (CA)

On-site
USD 180,000 - 240,000
Cloud-Scale Backend Engineer for ML Inference
Cloud-Scale Backend Engineer for ML Inference

Praxis, Inc. • San Francisco (CA)

On-site
USD 170,000 - 250,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Inference Platform Engineer (Kubernetes)
ML Inference Platform Engineer (Kubernetes)

Baseten • New York (NY)

On-site
USD 165,000 - 330,000
ML Infra Engineer: Scale Training & Inference (Hybrid)
ML Infra Engineer: Scale Training & Inference (Hybrid)

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
Lead ML Infra Architect for High-Performance NLP
Lead ML Infra Architect for High-Performance NLP

Cohere • New York (NY)

On-site
USD 180,000 - 240,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+5
Scale AI Infra Engineer | Kubernetes & CI/CD
Scale AI Infra Engineer | Kubernetes & CI/CD

BaseTen • New York (NY), San Francisco (CA)

Hybrid
USD 160,000 - 210,000
Meaningful equity
100% medical, dental, vision coverage
Flexible PTO including Winter Break
+4