Platform Engineer - GPU Infra & Kubernetes

Together AI

San Francisco (CA)

On-site

USD 160,000 - 280,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Equity
Health insurance
Competitive benefits

Job summary

Together AI, the AI Native Cloud, is hiring a software engineer to build a Kubernetes-native control plane for provisioning and running a GPU inference fleet. You will design a manifest-driven API where inference teams request clusters or models, with controllers handling reconciliation, provider selection, and lifecycle management.

The role emphasizes decoupling users from runtime complexity, improving utilization via defragmentation, bin-packing, and right-sizing, and owning the end-to-end

Qualifications

  • Strong software engineering background with real software delivery experience.
  • Experience building control planes or orchestration systems that model state over time.
  • Experience with event-driven messaging and modern queues or pub/sub systems.
  • Product-minded developer who ships usable platform APIs for other teams.

Responsibilities

  • Build provisioning state machine covering discovery to decommission and repair.
  • Create self-service API and control plane for request, scale, tear down.
  • Automate self-healing, draining, and capacity reintroduction.
  • Ensure reliability with idempotency, retries, and drift detection.
  • Collaborate with ML/inference teams to encode cluster requirements as abstractions.
  • Develop software-like practices with strong typing, tests, and CI/CD.

Skills

Go/Python/Rust
Kubernetes controllers
Event-driven systems
Product mindset
Testing & CI/CD

Tools

Temporal/Cadence
Kafka/NATS/SQS
Kubernetes

Job description

Together AI, the AI Native Cloud, is hiring a software engineer to build a Kubernetes-native control plane for provisioning and running a GPU inference fleet. You will design a manifest-driven API where inference teams request clusters or models, with controllers handling reconciliation, provider selection, and lifecycle management.

The role emphasizes decoupling users from runtime complexity, improving utilization via defragmentation, bin-packing, and right-sizing, and owning the end-to-end

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Kubernetes Platform Engineer for GPU Inference
Kubernetes Platform Engineer for GPU Inference

Togetherai • San Francisco (CA)

On-site
USD 160,000 - 280,000
Health insurance
GPU Compute Infrastructure Engineer
GPU Compute Infrastructure Engineer

Linuxcareers • San Francisco (CA), Northern (KY)

Hybrid
USD 210,000 - 270,000
AI Infra Engineer: GPU Cloud & Kubernetes POC Leader
AI Infra Engineer: GPU Cloud & Kubernetes POC Leader

Vcluster • Northern (KY)

Hybrid
USD 140,000 - 165,000
Competitive Salary
Equity participation
Health, dental, vision, life Insurance
+2
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes
Lead Cloud Infrastructure Engineer - GPU AI Kubernetes

FriendliAI • San Francisco (CA)

On-site
USD 150,000 - 190,000
Flexible working hours
Lunch and dinner provided
Health check-up support with top-tier硬
+1
Kubernetes-Native GPU AI Platform Engineer
Kubernetes-Native GPU AI Platform Engineer

GTN Technical Staffing • Town of Texas (WI), Fort Worth (TX)

Hybrid
USD 150,000 - 210,000
Relocation assistance
Hybrid work arrangement
Remote work flexibility
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Kubernetes Platform Engineer - GPU & AI Infra
Kubernetes Platform Engineer - GPU & AI Infra

GTN Technical Staffing • Dallas (TX)

Hybrid
USD 165,000 - 210,000
Relocation available
Hybrid work model
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)
AI Infra Engineer: Kubernetes on Bare Metal GPUs (Remote)

vCluster • New York (NY)

Hybrid
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
AI Infrastructure Engineer — GPU Kubernetes for Production
AI Infrastructure Engineer — GPU Kubernetes for Production

vCluster • Germany (OH)

On-site
USD 150,000 - 200,000
Competitive Salary
Platinum-Level Insurance
Flexible Working Schedule
+1
GPU AI Platform Engineer - Kubernetes & DevOps
GPU AI Platform Engineer - Kubernetes & DevOps

AMD • San Jose (CA)

On-site
USD 140,000 - 170,000
AMD benefits