Senior Platform Engineer, Inference & GPU Compute Infra

Together

San Francisco (CA)

On-site

USD 240,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Startup equity
Health insurance

Job summary

Together AI is seeking a Staff Software Engineer to build systems that treat infrastructure as software, turning racks of GPUs into running inference clusters via a declarative manifest-driven platform. You’ll design engines to materialize clusters, with a focus on automation, reliability, and a product mindset.

You will own provisioning state machines, self-service APIs, and end-to-end lifecycle management, from discovery to decommission, while ensuring production-grade quality and CI/CD

Qualifications

  • Strong software engineering background in Go, Python, Rust, or similar — you write and test real software for a living.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems — designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
  • A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.

Responsibilities

  • Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA — as explicit, versioned states and transitions.
  • Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call — no ticket, no human in the loop.
  • Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically.
  • Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service.
  • Partner with the inference/ML platform team: understand the cluster shapes they need — topology, interconnect, scheduling constraints — and encode them as first-class abstractions in the platform.
  • Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code — this is a product, not a collection of Ansible playbooks.

Skills

Go
Python
Rust
Temporal
Cadence
Kubernetes
Event-driven design
Product mindset

Tools

Temporal
Cadence
Kafka

Job description

Together AI is seeking a Staff Software Engineer to build systems that treat infrastructure as software, turning racks of GPUs into running inference clusters via a declarative manifest-driven platform. You’ll design engines to materialize clusters, with a focus on automation, reliability, and a product mindset.

You will own provisioning state machines, self-service APIs, and end-to-end lifecycle management, from discovery to decommission, while ensuring production-grade quality and CI/CD

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer AI Inference Infra Orchestrator
Staff Software Engineer AI Inference Infra Orchestrator

Together AI • San Francisco (CA)

On-site
USD 240,000 - 280,000
Equity
Health insurance
Competitive benefits
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Staff Engineer, GPU Inference & Training Platform
Staff Engineer, GPU Inference & Training Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior Platform Engineer – Inference API & Streaming
Senior Platform Engineer – Inference API & Streaming

General Compute Inc. • New York (NY)

On-site
USD 130,000 - 160,000
Senior Platform Engineer – AI/ML Infra (Equity)
Senior Platform Engineer – AI/ML Infra (Equity)

NVIDIA Corporation • Santa Clara (CA), Northern (KY)

Hybrid
USD 200,000 - 322,000
Senior Platform Engineer — GPU Inference & Training
Senior Platform Engineer — GPU Inference & Training

Neura Market • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior Platform Engineer, OpenAI-compatible Inference API
Senior Platform Engineer, OpenAI-compatible Inference API

General Compute Inc. • San Francisco (CA)

On-site
USD 200,000 - 260,000
Senior Inference Platform Engineer (Kubernetes & GPU)
Senior Inference Platform Engineer (Kubernetes & GPU)

NVIDIA • Town of Texas (WI)

On-site
USD 152,000 - 242,000
Senior Inference Platform Engineer
Senior Inference Platform Engineer

Designworks Talent • Bellevue (WA)

Hybrid
USD 180,000 - 240,000
Hybrid work model
Office Bellevue
Competitive compensation
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000