Staff Software Engineer AI Inference Infra Orchestrator

Together AI

San Francisco (CA)

On-site

USD 240,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Health insurance
Competitive benefits

Job summary

Together AI in San Francisco is hiring a Software Engineer to build the systems that treat infrastructure as software, turning racks of GPUs into running inference clusters through a manifest-driven platform. You will design and implement the engines that manifest schema and execute against it, while producing production-grade code and CI/CD pipelines.

You will own the provisioning lifecycle, create declarative APIs for the inference team, and automate repairs and reliability measures, ensuring

Qualifications

  • Strong software engineering background in Go, Python, or Rust.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time.
  • Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub.
  • A product mindset. You've built internal platforms or APIs consumed by other engineering teams.
  • Nice to have: exposure to bare-metal provisioning and networking fundamentals, or GPU infrastructure.
  • Nice to have: experience with GPU cluster software stacks and hyperscaler environments.
  • Nice to have: systems programming in Rust or Go.

Responsibilities

  • Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery to decommission/RMA.
  • Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call.
  • Automate self-healing: detect degraded nodes, drain safely, trigger repair or replacement, and reintroduce healthy capacity.
  • Own reliability of the pipeline: idempotency, retries, rollback, and drift detection.
  • Partner with the inference/ML platform team to encode cluster shapes as abstractions in the platform.
  • Engineer it like software: strong typing, automated tests, code reviews, and CI/CD for infrastructure code.

Skills

Go
Python
Rust
Software engineering
CI/CD
Event-driven systems

Tools

Temporal
Cadence
Kafka
NATS
SQS
Kubernetes
Workflow engines

Job description

Together AI in San Francisco is hiring a Software Engineer to build the systems that treat infrastructure as software, turning racks of GPUs into running inference clusters through a manifest-driven platform. You will design and implement the engines that manifest schema and execute against it, while producing production-grade code and CI/CD pipelines.

You will own the provisioning lifecycle, create declarative APIs for the inference team, and automate repairs and reliability measures, ensuring

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
Senior AI Infrastructure Engineer — Scale, Reliability & Automation
Senior AI Infrastructure Engineer — Scale, Reliability & Automation

AI Chopping Block • San Francisco (CA)

On-site
USD 190,000 - 270,000
Equity
Health insurance
Startup benefits
Senior Systems Engineer, AI Inference Platform
Senior Systems Engineer, AI Inference Platform

Slope • San Francisco (CA)

On-site
USD 180,000 - 260,000
Staff Software Engineer, Inference Systems at Scale
Staff Software Engineer, Inference Systems at Scale

Anthropic • San Francisco (CA)

On-site
USD 300,000 - 485,000
Competitive salary
Flexible working hours
Generous vacation and parental leave
Senior AI Inference Deployment Engineer
Senior AI Inference Deployment Engineer

Anthropic • Seattle (WA)

Hybrid
USD 320,000 - 485,000
Competitive compensation
Flexible working hours
Generous vacation and parental leave
Inference Platform Backend Engineer (Equity & Benefits)
Inference Platform Backend Engineer (Equity & Benefits)

Together • San Francisco (CA)

On-site
USD 160,000 - 250,000
Equity
Health insurance
Competitive compensation
Staff AI Systems Engineer — Inference & RL
Staff AI Systems Engineer — Inference & RL

Together • San Francisco (CA)

On-site
USD 200,000 - 280,000
Health insurance
Startup equity
Competitive benefits
AI Inference Orchestration - Distributed Systems Engineer
AI Inference Orchestration - Distributed Systems Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 350,000
Staff Software Engineer, Scalable AI Inference Systems
Staff Software Engineer, Scalable AI Inference Systems

Anthropic • San Francisco (CA)

On-site
USD 320,000 - 485,000