MLOps Engineer

Zof AI

San Francisco, Northern (CA, KY)

Hybrid

USD 170,000 - 240,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Zof AI seeks an experienced MLOps Engineer to own the platform for AI systems, spanning model serving, scaling, and reliability. You will push the platform that enables a small team to run large-scale AI workloads efficiently in production, with a focus on GPU throughput, observability, and cost control.

As a senior on-site engineer in San Francisco, you will design and operate the backend infrastructure, partner with software engineers, and drive incident practices, rollback automation, and

Qualifications

  • Experience operating production AI/ML systems.
  • Strong infra and systems engineering fundamentals.
  • Experience with cloud platforms, containers, and orchestration.
  • Experience with observability and reliability tooling.
  • Judgment about cost, performance, and operational trade-offs.
  • Clear written and verbal communication.
  • High ownership of production systems.

Responsibilities

  • Design and operate model serving, scaling, and compute infrastructure.
  • Own GPU and inference efficiency, capacity, and cost.
  • Build the platform for provisioning, running, and scaling agents.
  • Build observability for AI systems: tracing, monitoring, and alerting.
  • Keep production AI systems stable, debuggable, and recoverable.
  • Automate deployment, rollback, and environment management.
  • Set reliability standards and incident practices for AI workloads.
  • Partner with engineers to make the platform fast to build on.

Skills

Production AI
Cloud platforms
Containers
Observability
Reliability tooling
Ownership
Clear communication

Tools

Kubernetes
Terraform
Docker

Job description

Zof AI is seeking a MLOps Engineer to own the platform our AI systems run on. This is a consolidated platform role spanning what the market posts as AI Infrastructure, MLOps, LLMOps, and agent platform engineering: model serving and scaling, GPU and compute efficiency, and the provisioning, observability, and reliability infrastructure that keeps production agents stable. The ideal candidate has operated real AI workloads in production and builds infrastructure that lets a small team run systems well above its weight.

Engineering · Senior · Full-time · On-site · San Francisco, CA

Responsibilities
  • Design and operate model serving, scaling, and compute infrastructure.
  • Own GPU and inference efficiency, capacity, and cost.
  • Build the platform for provisioning, running, and scaling agents.
  • Build observability for AI systems: tracing, monitoring, and alerting.
  • Keep production AI systems stable, debuggable, and recoverable.
  • Automate deployment, rollback, and environment management.
  • Set reliability standards and incident practices for AI workloads.
  • Partner with engineers to make the platform fast to build on.
Requirements
  • Experience operating production AI, ML, or high-scale backend systems.
  • Strong infrastructure and systems engineering foundation.
  • Experience with cloud platforms, containers, and orchestration.
  • Experience with observability and reliability tooling.
  • Judgment about cost, performance, and operational trade-offs.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership of systems in production.
Nice to have
  • Experience with GPU clusters, inference servers, or model gateways.
  • Experience running agent workloads or long-lived AI processes.
  • Experience with Kubernetes, Terraform, or similar tooling.
  • Experience in early-stage platform teams.

Experience operating production AI or ML systems at scale is required

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer – MLOps & Infra
Senior AI Platform Engineer – MLOps & Infra

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
AI Infrastructure / MLOps Engineer — NYC
AI Infrastructure / MLOps Engineer — NYC

LaStellar Group • New York (NY)

On-site
USD 140,000 - 180,000
MLOps Engineer
MLOps Engineer

Evlo AI • Boston (MA)

On-site
USD 130,000 - 180,000
MLOps Engineer
MLOps Engineer

Sierracorp • San Francisco (CA)

On-site
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

ACI Infotech • Atlanta (GA)

On-site
USD 100,000 - 120,000
MLOps Engineer
MLOps Engineer

Compunnel, Inc. • San Antonio (TX)

On-site
USD 100,000 - 130,000
AI Infrastructure Engineer MLOps
AI Infrastructure Engineer MLOps

EITACIES Inc. • San Francisco (CA)

On-site
USD 120,000 - 150,000
401(k)
ML Ops Engineer — Agentic AI Lab (Founding Team)
ML Ops Engineer — Agentic AI Lab (Founding Team)

Fabrion • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive salary
Meaningful equity
MLOps Engineer: Scalable ML Pipelines & Infra
MLOps Engineer: Scalable ML Pipelines & Infra

Compunnel, Inc. • San Antonio (TX)

On-site