ML Infrastructure Engineer

Clera

San Mateo (CA)

On-site

USD 180,000 - 240,000

Full time

5 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Clera is an early-stage enterprise AI startup hiring an hands-on ML Infrastructure Engineer to own end-to-end inference and model-serving infrastructure. You will build scalable systems enabling AI agents to run reliably under high concurrency, collaborate with ML and infra teams for performance, and optimize latency and throughput in production environments.

You will work on Docker/Kubernetes based deployments with TensorFlow Serving, TorchServe, Triton, and KServe, using observability stacks

Qualifications

  • 5+ years of experience building and operating ML inference systems or ML infrastructure in production.
  • Hands-on experience with inference-serving frameworks (TensorFlow Serving, TorchServe, Triton, KServe) and custom systems.
  • Strong track record optimizing latency, throughput, and reliability at scale.

Responsibilities

  • Own inference and model-serving infrastructure end to end, from design through production deployment.
  • Build and scale systems enabling AI agents to run reliably under increasing concurrency.
  • Collaborate with ML and infra teams to ensure seamless integration and performance optimization.
  • Identify infrastructure bottlenecks and drive cross-functional solutions across engineering teams.

Skills

ML inference
Latency optimization
Distributed systems
Observability
Cloud platforms
Backend languages

Tools

Docker
Kubernetes
TensorFlow Serving
TorchServe
Triton
KServe
Prometheus
Grafana
ELK
Neo4j

Job description

About the Role

This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI startup, where you'll own the end-to-end inference and model-serving infrastructure that keeps production AI agents running reliably and at scale. You'll sit at the intersection of ML and platform engineering, directly shaping the systems that power real-world, high-stakes deployments in regulated industries like insurance, banking, and healthcare.

What You'll Do
  • Own inference and model-serving infrastructure end to end, from design through production deployment.

  • Build and scale systems that enable AI agents to run reliably and efficiently under increasing concurrency.

  • Collaborate closely with ML and infrastructure teams to ensure seamless integration and performance optimization.

  • Identify infrastructure bottlenecks and drive cross-functional solutions across engineering teams.

What We're Looking For
  • 5+ years of experience building and operating ML inference systems, model-serving platforms, or ML infrastructure in production.

  • Hands-on experience designing and scaling inference-serving infrastructure using frameworks such as TensorFlow Serving, TorchServe, Triton, KServe, or equivalent custom systems.

  • Strong track record optimizing production ML systems for latency, throughput, and reliability at scale.

  • Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads.

  • Experience building distributed systems that handle concurrent requests and manage resource allocation under load.

  • Proficiency with observability and debugging tooling for production systems (e.g., Prometheus, Grafana, ELK, distributed tracing).

  • Cloud platform experience on AWS, GCP, or Azure for deploying and managing ML systems.

  • Proficiency in at least one systems or backend language — Python, Go, Rust, C++, or Java.

  • Nice to have: experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune); real-time or low-latency inference systems; agentic or multi-step reasoning pipelines; enterprise data infrastructure or integration platforms.

Location

On-site in San Mateo, CA. No visa sponsorship is available for this role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

ML Infrastructure Engineer
ML Infrastructure Engineer

Lattice, Inc. • San Francisco (CA)

Hybrid
USD 200,000 - 280,000
Competitive salary
Premium health, dental, and vision insurance
Unlimited PTO
+2
ML Infrastructure Engineer
ML Infrastructure Engineer

Acceler8 Talent • San Francisco (CA)

Hybrid
USD 233,000 - 275,000
Sr. Platform Engineer, ML Infrastructure
Sr. Platform Engineer, ML Infrastructure

Insilico Search Partners • Cambridge (MA)

On-site
USD 140,000 - 210,000
Backend Software Engineer (ML Infra)
Backend Software Engineer (ML Infra)

Rockstar • San Francisco (CA)

On-site
USD 100,000 - 130,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
Senior Machine Learning Engineer
Senior Machine Learning Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Mach9 • San Francisco (CA)

On-site
USD 150,000 - 210,000
Competitive salary
Health insurance
Flexible hours
+1
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Parental leave
Lead Data and ML Infrastructure Engineer
Lead Data and ML Infrastructure Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 180,000 - 260,000
Competitive compensation
Meaningful equity
Autonomy and ownership
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 280,000