Staff Platform Engineer - GPU Inference Control Plane

Togetherai

Greater London

On-site

GBP 120,000 - 180,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Together AI is seeking a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You will design a manifest-driven API where the inference team declares needs such as a cluster or a model deployment, and our controllers reconcile and manage lifecycle behind the scenes.

You will own reliability, performance optimizations, defragmentation, and scheduling improvements, shaping a platform that decouples developers from runtime complexity.

Qualifications

  • Strong software engineering background in Go, Python, Rust.
  • Experience with durable workflow orchestration tools such as Temporal or Cadence.
  • Experience building control planes or orchestration systems (Kubernetes controllers/operators).
  • Experience with event-driven architectures (message queues).
  • Product mindset, internal platforms or APIs for other teams.

Responsibilities

  • Build the provisioning state machine modelling the full lifecycle of a host and GPU stack.
  • Build a self-service API and control plane for cluster provisioning and teardown.
  • Automate self-healing: drain, repair or replace, reintegrate capacity.
  • Own reliability: idempotency, retries, rollback, drift detection.
  • Partner with ML platform teams to encode cluster constraints as abstractions.
  • Engineer it like software: strong typing, tests, CI/CD for infrastructure code.

Skills

Kubernetes controllers
Go
Python
Rust
Event-driven
API design
Product mindset

Tools

Temporal
Cadence
Kafka

Job description

Together AI is seeking a software engineer to build the Kubernetes-native control plane that provisions and runs our GPU inference fleet. You will design a manifest-driven API where the inference team declares needs such as a cluster or a model deployment, and our controllers reconcile and manage lifecycle behind the scenes.

You will own reliability, performance optimizations, defragmentation, and scheduling improvements, shaping a platform that decouples developers from runtime complexity.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Kubernetes-native GPU Inference
Staff Software Engineer, Kubernetes-native GPU Inference

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • Greater London

On-site
GBP 100,000 - 160,000
Staff Software Engineer, Inference / Compute Infrastructure Engineering London or Amsterdam
Staff Software Engineer, Inference / Compute Infrastructure Engineering London or Amsterdam

Togetherai • Greater London

On-site
GBP 120,000 - 180,000
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

CoreWeave • Greater London

On-site
GBP 120,000 - 180,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+6
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Platform Engineer – Scale AI Infra & GPU Orchestration
Platform Engineer – Scale AI Infra & GPU Orchestration

Ineffable Intelligence • Greater London

On-site
GBP 70,000 - 110,000
Staff Cloud Native Engineer: Kubernetes AI Infra Leader
Staff Cloud Native Engineer: Kubernetes AI Infra Leader

Greenhouse Software, Inc. • United Kingdom

Remote
GBP 90,000 - 120,000
Staff Cloud Native Engineer — AI GPU Infra Architect
Staff Cloud Native Engineer — AI GPU Infra Architect

Nscale • Greater London

On-site
GBP 110,000 - 170,000
AI Infrastructure DevOps Engineer (Kubernetes + GPU)
AI Infrastructure DevOps Engineer (Kubernetes + GPU)

Carbon3.ai • Greater London

On-site
GBP 70,000 - 110,000
Senior HPC Systems Engineer - GPU & Bare-Metal Kubernetes
Senior HPC Systems Engineer - GPU & Bare-Metal Kubernetes

remotestar-team • Cambourne

Hybrid
GBP 90,000 - 150,000
Indefinite contract
Equal pay guaranteed
Variable performance bonus
+7