Staff Software Engineer, Runtime Systems

CoreWeave

Greater London

On-site

GBP 120,000 - 180,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Living Wage Accredited Employer

Job summary

CoreWeave in London seeks a Staff Software Engineer, Runtime Systems to design and build the platform layer enabling AI workloads to run reliably at scale. You will define abstractions between workloads and execution environments, and implement production-grade runtimes across Kubernetes, Argo, and related systems.

You will collaborate with Go, Kubernetes, and infrastructure engineers to drive architectural decisions, mentor engineers, and shape the future of distributed AI compute.

Qualifications

  • Significant experience building complex systems software or distributed infrastructure.
  • Deep expertise in runtime systems, workflow/execution engines, or programming languages.

Responsibilities

  • Design and build runtime components for complex AI/workloads.
  • Define abstractions and contracts between systems and environments.
  • Write production-quality software and build adapters for heterogeneous environments.
  • Mentor engineers and influence architectural decisions across teams.

Skills

Distributed systems
Runtime systems
APIs & contracts
Go/Python/C/C++

Tools

Kubernetes
Argo
OSMO

Job description

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.

We're proud to be a Living Wage accredited Employer.

What You’ll Do

The Physical AI Engineering Team at CoreWeave is building the software and infrastructure that enables demanding AI, simulation, robotics, and engineering workloads to run reliably at scale.

As these workloads become more complex, the challenge is no longer simply providing compute. We need to make heterogeneous workloads easier to execute, observe, reproduce, and move across different systems without hiding the capabilities or semantics of the infrastructure underneath them.

About The Role

We’re seeking a Staff Software Engineer, Runtime Systems to help design and build this layer.

This is a hands-on systems engineering role at the intersection of distributed systems, runtimes, workflow execution, programming language concepts, and large-scale compute infrastructure.

You’ll work on the foundations that allow complex workloads to move from an abstract description into reliable execution across systems such as Kubernetes, Argo, OSMO, and future execution environments.

A major part of the role is deciding where abstraction is useful — and where it creates more complexity. Rather than building another universal workflow engine, you’ll help establish clear contracts between our platform and the systems that execute work, while preserving native capabilities.

You’ll operate across architecture and implementation: defining contracts, writing production software, validating assumptions against real workloads, and working closely with platform and infrastructure teams.

In This Role, You Will
  • Runtime Systems & Architecture
  • Design and build runtime components for complex AI, simulation, and engineering workloads.
  • Define abstractions for workloads, execution environments, dependencies, state, capabilities, and failure.
  • Design interfaces between higher-level services and systems such as Kubernetes, Argo, and OSMO.
  • Establish clear boundaries around which systems own state, decisions, and side effects.
  • Make architectural decisions balancing simplicity, extensibility, performance, and operational reality.
  • Execution Models & Contracts
  • Define durable contracts between workload definitions, control-plane services, and execution backends.
  • Develop typed representations and schemas that allow workloads to be transformed safely across systems.
  • Design compatibility and evolution mechanisms for those contracts.
  • Build conformance and validation mechanisms that make guarantees executable rather than dependent on documentation.
  • Reason deeply about retries, partial failure, idempotency, cancellation, dependencies, and uncertain outcomes.
  • Hands-On Systems Engineering
  • Write production-quality software for critical runtime and control-plane components.
  • Build adapters and integrations for heterogeneous execution environments.
  • Diagnose behaviour across application, orchestration, cluster, and infrastructure boundaries.
  • Improve the reliability, observability, and debuggability of distributed workload execution.
  • Work closely with Go, Kubernetes, and infrastructure engineers to turn architecture into production systems.
  • Performance & Experimentation
  • Develop rigorous ways to understand workload performance across large-scale GPU infrastructure.
  • Design experiments that separate real performance gains from noise, warm-up effects, scheduling behaviour, and stragglers.
  • Build repeatable workload and benchmark environments.
  • Use evidence from real execution to challenge assumptions and guide platform development.
  • Technical Leadership
  • Lead ambiguous systems problems where the correct architecture is not yet known.
  • Reduce complex problems into smaller contracts and mechanisms that can actually be implemented.
  • Challenge unnecessary abstraction and simplify designs where complexity has outgrown its value.
  • Influence technical direction across teams without requiring direct authority.
  • Mentor engineers and contribute to technical hiring and engineering standards.
Who You Are
  • Significant experience building complex systems software, distributed infrastructure, runtimes, workflow systems, or adjacent technology.
  • Deep expertise in at least one of:
    • Distributed systems
    • Runtime systems
    • Workflow or execution engines
    • Programming languages, compilers, or interpreters
    • Cluster scheduling and orchestration
    • High-performance or systems software
  • Strong software engineering fundamentals and production coding ability.
  • Experience designing APIs, protocols, schemas, or contracts between independently evolving systems.
  • Strong understanding of distributed-system failure modes, state, authority, retries, concurrency, and side effects.
  • Strong technical judgement around when abstraction helps and when it simply moves complexity elsewhere.
  • Comfortable entering unfamiliar technical domains and building depth quickly.
  • Strong communication skills and experience influencing architectural decisions across teams.
Experience with some of the following would be valuable, but is not required:
  • Go, Rust, C/C++, or Python.
  • Kubernetes and containerised infrastructure.
  • Argo, OSMO, Temporal, Ray, Kubeflow, or similar systems.
  • GPU clusters or large-scale AI infrastructure.
  • High-performance computing.
  • Simulation, robotics, autonomous systems, or Physical AI.
  • Programming language or compiler research.
  • Performance engineering.
  • Cloud infrastructure at scale.
  • Platforms designed to be operated by autonomous software or AI agents.

Wondering If You’re a Good Fit? We Believe In Investing In Our People, And Value Candidates Who Can Bring Their Own Diversified Experiences To Our Teams – Even If You Aren’t a 100% Skill Or Experience Match.

Here Are a Few Qualities We’ve Found Compatible With Our Team
  • Systems Thinker: You ask where authority lives, what guarantees actually exist, and what happens when systems fail.
  • Technically Deep: You want to understand how systems really behave, not just how they are supposed to behave.
  • Pragmatic: You value elegant engineering, but care more about whether it works for real workloads.
  • Evidence-Driven: You test assumptions and change direction when the evidence says you should.
  • Strong Technical Leader: You can form a view, challenge senior stakeholders, and bring others with you.
  • Commercially Aware: You understand that technical decisions need to improve customer outcomes, engineering velocity, reliability, or economics.
  • Ownership Mentality: You take responsibility for getting difficult systems into production.
Why CoreWeave?
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Runtime Systems
Staff Software Engineer, Runtime Systems

Coreweaveu • Greater London

On-site
GBP 120,000 - 180,000
Living Wage
Staff Software Engineer - Physical AI
Staff Software Engineer - Physical AI

CoreWeave • Greater London

On-site
GBP 120,000 - 180,000
Senior Software Engineer - Physical AI
Senior Software Engineer - Physical AI

CoreWeave • Greater London

On-site
GBP 100,000 - 160,000
Technical Solutions Manager
Technical Solutions Manager

CoreWeave • Greater London

On-site
GBP 116,000 - 155,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+6
Senior Specialist Field Engineer - HPC/AI/ML
Senior Specialist Field Engineer - HPC/AI/ML

CoreWeave • Greater London

On-site
GBP 98,000 - 130,000
Medical Insurance
Dental Insurance
Pension Plan
+5
MLOps Engineer
MLOps Engineer

CoreWeave Europe • Greater London

On-site
GBP 90,000 - 140,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+4
Staff Product Manager
Staff Product Manager

CoreWeave Europe • Greater London

On-site
GBP 120,000 - 170,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Senior Specialist Field Engineer - Security
Senior Specialist Field Engineer - Security

Coreweave • Greater London

On-site
GBP 90,000 - 130,000
Family-level Medical Insurance
Dental Insurance
Pension Contribution
+2
Staff Product Manager - Physical AI/Robotics
Staff Product Manager - Physical AI/Robotics

United States Digital Space LLC • Greater London

On-site
GBP 120,000 - 180,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+4
Senior Specialist Field Engineer - HPC/AI/ML
Senior Specialist Field Engineer - HPC/AI/ML

CoreWeave Europe • Greater London

On-site
GBP 98,000 - 130,000
Living Wage accredited Employer
Discretionary bonus
Equity awards
+1