Scale ML Inference & Model-Serving Engineer

Mixpeek

San Mateo (CA)

On-site

USD 150,000 - 230,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Full benefits (medical, dental, vision
401k
H-1B visa sponsorship and immigration”

Job summary

zaimler is seeking an engineer to own our inference and model-serving infrastructure end to end, enabling production-grade AI agents to run fast and reliably at scale. You’ll report to Sofus and collaborate with ML and infra teams on deployment and scalability.

You will set up and scale inference systems, manage GPU resources for concurrency, and optimize the engine that powers scalable agent orchestration, joining an early engineering team in a cutting-edge context.

Qualifications

  • Proven ability to build scalable ML/AI platforms end-to-end for production.
  • Deep understanding of the inference stack: vLLM, KV cache, and the optimization layers underneath model serving.
  • Experience building distributed AI/ML workloads at scale, connecting them to real product or vertical integrations.
  • 3+ years of relevant experience. We care about capability, not tenure.

Responsibilities

  • Set up and scale inference/Ray Serve for ML and LLM model serving.
  • Scale agent GPU infrastructure for concurrency and efficiency across multiple agent workloads.
  • Optimize and improve the engine builder and model server that power scalable agent orchestration.

Skills

Scalable ML platforms
Distributed systems
Model serving
Ray Serve

Tools

Ray Serve
vLLM
KV cache
Distributed systems tooling

Job description

zaimler is seeking an engineer to own our inference and model-serving infrastructure end to end, enabling production-grade AI agents to run fast and reliably at scale. You’ll report to Sofus and collaborate with ML and infra teams on deployment and scalability.

You will set up and scale inference systems, manage GPU resources for concurrency, and optimize the engine that powers scalable agent orchestration, joining an early engineering team in a cutting-edge context.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Applied ML Engineer, AI Infrastructure
Staff Applied ML Engineer, AI Infrastructure

Mixpeek • San Mateo (CA)

On-site
USD 180,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Mixpeek • San Mateo (CA)

On-site
USD 150,000 - 230,000
Full benefits (medical, dental, vision
401k
H-1B visa sponsorship and immigration”
Staff Applied Machine Learning Engineer
Staff Applied Machine Learning Engineer

Mixpeek • San Mateo (CA)

On-site
USD 180,000 - 240,000
Senior ML Systems Engineer - Scalable Inference
Senior ML Systems Engineer - Scalable Inference

Atlassian • Seattle (WA)

Hybrid
USD 206,000 - 269,000
Health and wellbeing resources
Paid volunteer days
Model Inference Engineer: ML Serving & Hardware
Model Inference Engineer: ML Serving & Hardware

Google Inc. • Mountain View (CA)

On-site
USD 180,000 - 240,000
Systems ML Engineer - High-Scale AI Infra (Equity)
Systems ML Engineer - High-Scale AI Infra (Equity)

Meta • Concord (NH)

On-site
USD 154,000 - 217,000
Semantic ML Engineer for Enterprise Data & LLMs
Semantic ML Engineer for Enterprise Data & LLMs

zaimler • San Mateo (CA)

On-site
USD 140,000 - 190,000
Staff ML Infra / MLOps Engineer: Scale & Automate AI
Staff ML Infra / MLOps Engineer: Scale & Automate AI

Quince • Palo Alto (CA)

On-site
USD 218,000 - 285,000
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO
GenAI ML Systems Engineer: Scalable Training & Inference
GenAI ML Systems Engineer: Scalable Training & Inference

Meta • Menlo Park (CA)

On-site
USD 180,000 - 300,000