ML Inference Engineer — Scale & Reliability

Clera

San Mateo (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Clera is an early-stage enterprise AI startup hiring an hands-on ML Infrastructure Engineer to own end-to-end inference and model-serving infrastructure. You will build scalable systems enabling AI agents to run reliably under high concurrency, collaborate with ML and infra teams for performance, and optimize latency and throughput in production environments.

You will work on Docker/Kubernetes based deployments with TensorFlow Serving, TorchServe, Triton, and KServe, using observability stacks

Qualifications

  • 5+ years of experience building and operating ML inference systems or ML infrastructure in production.
  • Hands-on experience with inference-serving frameworks (TensorFlow Serving, TorchServe, Triton, KServe) and custom systems.
  • Strong track record optimizing latency, throughput, and reliability at scale.

Responsibilities

  • Own inference and model-serving infrastructure end to end, from design through production deployment.
  • Build and scale systems enabling AI agents to run reliably under increasing concurrency.
  • Collaborate with ML and infra teams to ensure seamless integration and performance optimization.
  • Identify infrastructure bottlenecks and drive cross-functional solutions across engineering teams.

Skills

ML inference
Latency optimization
Distributed systems
Observability
Cloud platforms
Backend languages

Tools

Docker
Kubernetes
TensorFlow Serving
TorchServe
Triton
KServe
Prometheus
Grafana
ELK
Neo4j

Job description

Clera is an early-stage enterprise AI startup hiring an hands-on ML Infrastructure Engineer to own end-to-end inference and model-serving infrastructure. You will build scalable systems enabling AI agents to run reliably under high concurrency, collaborate with ML and infra teams for performance, and optimize latency and throughput in production environments.

You will work on Docker/Kubernetes based deployments with TensorFlow Serving, TorchServe, Triton, and KServe, using observability stacks

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production ML Engineer Lead - Pipelines & Inference Scale
Production ML Engineer Lead - Pipelines & Inference Scale

Clera • San Francisco (CA)

On-site
USD 130,000 - 160,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Clera • San Mateo (CA)

On-site
USD 180,000 - 240,000
ML Inference Platform Engineer — Scale Production AI
ML Inference Platform Engineer — Scale Production AI

The Consensus • New York (NY)

On-site
USD 120,000 - 150,000
100% coverage of medical, dental, and vision insurance
Flexible PTO policy
Paid parental leave
+2
LLM Training & Inference Systems Engineer
LLM Training & Inference Systems Engineer

United States Digital Space LLC • New York (NY), San Francisco (CA)

On-site
USD 190,000 - 237,000
ML Infra Engineer: Build Scalable ML Services & MLOps
ML Infra Engineer: Build Scalable ML Services & MLOps

Stripe • United States

Hybrid
CAD 172,000 - 258,000
Equity
Retirement plans
Health benefits
+1
Lead ML Platform Engineer: Training & Inference at Scale
Lead ML Platform Engineer: Training & Inference at Scale

Paramount • Burbank (CA)

On-site
USD 157,000 - 235,000
Benefits package
On-site & virtual events
Generous PTO
ML Platform Engineer: Scale AI & Inference
ML Platform Engineer: Scale AI & Inference

Apply • San Francisco (CA)

Hybrid
USD 245,000 - 345,000
Flexible Time Off
Health Insurance
Work From Home Allowance
+2
Inference Infra Architect for Multimodal ML Systems
Inference Infra Architect for Multimodal ML Systems

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Remote MLOps Engineer: Scale AI Inference & Serving
Remote MLOps Engineer: Scale AI Inference & Serving

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior Staff ML Engineer - Scalable LLM Infra
Senior Staff ML Engineer - Scalable LLM Infra

Moveworks • Mountain View (CA), Northern (KY)

Hybrid
USD 190,000 - 280,000