AI Platform Engineer

Zof AI

San Francisco, Northern (CA, KY)

Hybrid

USD 170,000 - 240,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Zof AI is seeking an AI Platform Engineer to own the platform our AI systems run on, spanning model serving, GPU compute, and reliability infrastructure. This senior role requires building observability, automating deployment, and driving cost-aware scaling in production.

You will work on-site in San Francisco with a small, high-skill team, shaping the platform for rapid experimentation and robust production workloads.

Qualifications

  • Experience operating production AI/ML or high-scale backend systems.
  • Strong infrastructure and systems engineering foundation.
  • Experience with cloud platforms, containers, and orchestration.
  • Experience with observability and reliability tooling.
  • Judgment about cost, performance, and operational trade-offs.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership of systems in production.

Responsibilities

  • Design and operate model serving, scaling, and compute infrastructure.
  • Own GPU and inference efficiency, capacity, and cost.
  • Build the platform for provisioning, running, and scaling agents.
  • Build observability for AI systems: tracing, monitoring, and alerting.
  • Keep production AI systems stable, debuggable, and recoverable.
  • Automate deployment, rollback, and environment management.
  • Set reliability standards and incident practices for AI workloads.
  • Partner with engineers to make the platform fast to build on.

Skills

Production AI experience
Cloud platforms
Observability tooling
Cost-performance trade-offs
Clear communication
Ownership of systems
Fast-moving environment

Tools

Kubernetes
Terraform
Containers

Job description

Zof AI is seeking an AI Platform Engineer to own the platform our AI systems run on. This is a consolidated platform role spanning what the market posts as AI Infrastructure, MLOps, LLMOps, and agent platform engineering: model serving and scaling, GPU and compute efficiency, and the provisioning, observability, and reliability infrastructure that keeps production agents stable. The ideal candidate has operated real AI workloads in production and builds infrastructure that lets a small team run systems well above its weight.

Engineering · Senior · Full-time · On-site · San Francisco, CA

Responsibilities
  • Design and operate model serving, scaling, and compute infrastructure.
  • Own GPU and inference efficiency, capacity, and cost.
  • Build the platform for provisioning, running, and scaling agents.
  • Build observability for AI systems: tracing, monitoring, and alerting.
  • Keep production AI systems stable, debuggable, and recoverable.
  • Automate deployment, rollback, and environment management.
  • Set reliability standards and incident practices for AI workloads.
  • Partner with engineers to make the platform fast to build on.
Requirements
  • Experience operating production AI, ML, or high-scale backend systems.
  • Strong infrastructure and systems engineering foundation.
  • Experience with cloud platforms, containers, and orchestration.
  • Experience with observability and reliability tooling.
  • Judgment about cost, performance, and operational trade-offs.
  • Clear written and verbal communication.
  • Comfort operating in a fast-moving environment.
  • High ownership of systems in production.
Nice to have
  • Experience with GPU clusters, inference servers, or model gateways.
  • Experience running agent workloads or long-lived AI processes.
  • Experience with Kubernetes, Terraform, or similar tooling.
  • Experience in early-stage platform teams.

Experience operating production AI or ML systems at scale is required

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer for Production-Grade Inference
Senior AI Platform Engineer for Production-Grade Inference

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 170,000 - 240,000
Senior AI Applications Engineer
Senior AI Applications Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Applied AI Engineer
Applied AI Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 190,000
Software Engineer, AI Platform
Software Engineer, AI Platform

Fresh Ventures • San Francisco (CA)

On-site
USD 150,000 - 230,000
Health Insurance
Stock Options
401k
Backend Software Engineer
Backend Software Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 120,000 - 190,000
Senior Developer
Senior Developer

ICE • Atlanta (GA)

On-site
USD 180,000 - 240,000
Senior Platform Engineer – AI/ML Infrastructure & Reliability
Senior Platform Engineer – AI/ML Infrastructure & Reliability

StratITech • San Francisco (CA)

On-site
USD 210,000 - 260,000
Equity
AI Infrastructure Engineer MLOps
AI Infrastructure Engineer MLOps

EITACIES Inc. • San Francisco (CA)

On-site
USD 120,000 - 150,000
401(k)
AI Platform Engineer
AI Platform Engineer

Worky • San Francisco (CA)

On-site
USD 180,000 - 240,000
Senior DevOps Engineer
Senior DevOps Engineer

Zof AI • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 230,000