Kubernetes Platform Engineer for GPU Inference

Togetherai

San Francisco (CA)

On-site

USD 160,000 - 280,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Health insurance

Job summary

Together AI is hiring a software engineer to build a Kubernetes-native control plane that provisions and runs a GPU inference fleet. You’ll design a manifest-driven API and controllers that reconcile, select backends, and manage lifecycle, plus defragmentation and bin-packing to improve utilization.

You’ll own the platform end-to-end, including deployment, testing, and production support, with a focus on decoupling developers from underlying complexity.

Qualifications

  • Strong software engineering background in Go, Python, Rust, or similar.
  • Experience with durable workflow orchestration tools to run long-lived, manifest-driven workflows.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time.
  • Experience with event-driven systems (message queues, event streams, pub/sub).
  • A product mindset: built internal platforms or APIs consumed by other teams.

Responsibilities

  • Build the provisioning state machine for full host lifecycle from discovery to decommission.
  • Build a self-service API so the inference team can request and scale clusters with one API call.
  • Automate self-healing: drain, repair or replace, reintroduce healthy capacity.
  • Own reliability: idempotency, retries, rollback and drift detection.
  • Partner with ML platform teams to encode cluster shapes as abstractions.
  • Engineer with strong typing, tests, code review, CI/CD for infrastructure code.

Skills

Go
Python
Rust
Software engineering
Product mindset

Tools

Temporal
Cadence

Job description

Together AI is hiring a software engineer to build a Kubernetes-native control plane that provisions and runs a GPU inference fleet. You’ll design a manifest-driven API and controllers that reconcile, select backends, and manage lifecycle, plus defragmentation and bin-packing to improve utilization.

You’ll own the platform end-to-end, including deployment, testing, and production support, with a focus on decoupling developers from underlying complexity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Platform Engineer - GPU Infra & Kubernetes
Platform Engineer - GPU Infra & Kubernetes

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
Kubernetes Platform Engineer - GPU & AI Infra
Kubernetes Platform Engineer - GPU & AI Infra

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 165,000 - 210,000
Relocation available
Hybrid work model
Senior Platform Engineer, Inference & GPU Compute Infra
Senior Platform Engineer, Inference & GPU Compute Infra

Together • San Francisco (CA)

On-site
USD 240,000 - 280,000
Startup equity
Health insurance
Kubernetes-Native GPU AI Platform Engineer
Kubernetes-Native GPU AI Platform Engineer

GTN Technical Staffing • Town of Texas (WI), Fort Worth (TX)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work arrangement
Remote work flexibility
GPU Compute Infrastructure Engineer
GPU Compute Infrastructure Engineer

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000
[Junior / Senior / Staff] Software Engineer, Inference / Compute Infrastructure Engineering
[Junior / Senior / Staff] Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • San Francisco (CA)

On-site
USD 160,000 - 280,000
Equity
Health insurance
Competitive benefits
[Junior / Senior / Staff] Software Engineer, Inference / Compute Infrastructure Engineering
[Junior / Senior / Staff] Software Engineer, Inference / Compute Infrastructure Engineering

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 280,000
Health insurance
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes

FriendliAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
+1
GPU AI Platform Engineer - Kubernetes & DevOps
GPU AI Platform Engineer - Kubernetes & DevOps

AMD • San Jose (CA)

On-site
USD 140,000 - 170,000
AMD benefits
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)
Staff Engineer, AI Cloud Infra (Kubernetes + GPUs)

Lambda • San Francisco (CA)

Hybrid
USD 314,000 - 465,000
Health, dental, and vision coverage
401k with 2% company match
Wellness stipend
+1