Staff Engineer, Kubernetes GPU Inference Platform

Togetherai

Greater London

On-site

GBP 110,000 - 150,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Together AI, the AI Native Cloud, is seeking a software engineer to build a Kubernetes-native control plane that provisions and runs our GPU inference fleet.

You will design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving

Qualifications

  • Strong software engineering background and hands-on experience delivering production software.
  • Experience with durable workflow orchestration tools to run long-lived, manifest-driven workflows.

Responsibilities

  • Build the provisioning state machine: model full lifecycle of a host from discovery to GPU driver stack and decommission.
  • Build the self-service API to let inference teams request, scale, and teardown clusters with one API call.
  • Automate self-healing: detect degraded nodes, drain, repair or replace, and reintroduce capacity automatically.

Skills

Go
Python
Rust
Kubernetes controllers/operators
Event-driven systems
Product mindset
Temporal/Cadence

Tools

Temporal
Cadence

Job description

Together AI, the AI Native Cloud, is seeking a software engineer to build a Kubernetes-native control plane that provisions and runs our GPU inference fleet.

You will design a manifest-driven API where the inference team declares what they need, whether that's a cluster, a model deployment, or a capacity change, and our controllers handle the reconciliation, provider/runtime selection, and lifecycle management underneath, so the inference team never has to know or care which specific serving

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Kubernetes-native GPU Inference
Staff Software Engineer, Kubernetes-native GPU Inference

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Staff Software Engineer, Inference / Compute Infrastructure Engineering

Togetherai • Greater London

On-site
GBP 110,000 - 150,000
Staff Software Engineer, Inference / Compute Infrastructure Engineering
Staff Software Engineer, Inference / Compute Infrastructure Engineering

Together AI • Greater London

Hybrid
GBP 100,000 - 160,000
Staff Cloud Native Engineer — AI GPU Infra Architect
Staff Cloud Native Engineer — AI GPU Infra Architect

Nscale • Greater London

On-site
GBP 110,000 - 170,000
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

CoreWeave • Greater London

On-site
GBP 120,000 - 180,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+6
AI Infrastructure DevOps Engineer (Kubernetes + GPU)
AI Infrastructure DevOps Engineer (Kubernetes + GPU)

Carbon3.ai • Greater London

On-site
GBP 70,000 - 110,000
AI Platform Architect: Kubernetes & GPU Infrastructure
AI Platform Architect: Kubernetes & GPU Infrastructure

Carbon3.ai • United Kingdom

On-site
GBP 110,000 - 160,000
Platform Engineer – Scale GPU Infra for AI Platform
Platform Engineer – Scale GPU Infra for AI Platform

Ineffable Intelligence LTD • Greater London

Hybrid
GBP 85,000 - 120,000
Kubernetes & GPU Infra DevOps Engineer (Remote)
Kubernetes & GPU Infra DevOps Engineer (Remote)

Carbon3ai Limited. • United Kingdom

Remote
GBP 70,000 - 110,000
Occasional office visits
Senior GPU & AI Infrastructure Architect
Senior GPU & AI Infrastructure Architect

Referment • Greater London

On-site
GBP 90,000 - 120,000