Staff Software Engineer (MetalDev)

CoreWeave

York and North Yorkshire

On-site

GBP 120,000 - 180,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

CoreWeave is seeking a Staff Software Engineer to join the Compute Architecture group. You will help build software systems that operate the backbone of large-scale GPU data centers, focusing on Go-based distributed services and infrastructure automation.

You will design, build, and operate services that manage lifecycle, health, and firmware state across fleets of GPU servers, with emphasis on reliability, scalability, and observability. Strong leadership and mentoring are expected.

Qualifications

  • 8+ years of software engineering with infra, cloud, and distributed databases.
  • Strong Go expertise and REST/gRPC API experience for mission-critical platforms.
  • Excellent communicator with both technical and non-technical stakeholders.
  • Track record of leading incident response, postmortems, and reliability initiatives.
  • Mentor engineers, lead technical projects, and influence engineering strategy.
  • Hands-on with observability stacks (Prometheus, Grafana, PromQL), CI/CD, and large GPU fleets.
  • Background in architecting and scaling cloud-native Kubernetes infrastructure.
  • Experience contributing to/open-source collaboration and GPU server management.

Responsibilities

  • Design, build, and operate Go-based services for large-scale GPU data centers.
  • Build automation for hardware bring-up, health monitoring, and remediation.
  • Develop reliable REST/gRPC APIs for BMCs, firmware state, and health data.
  • Improve observability and tooling to detect and resolve production issues quickly.
  • Translate incidents into software improvements for platform resilience.
  • Partner with hardware-adjacent and software teams to scale fleet deployments.
  • Provide technical leadership through design and code reviews and mentorship.
  • Make pragmatic architecture decisions balancing reliability, simplicity, and scale.

Skills

Go
Distributed systems
Cloud engineering
Kubernetes
Observability
CI/CD
REST APIs
gRPC
Mentoring

Education

CS degree or equivalent

Tools

Prometheus
Grafana
PromQL
Kafka
ClickHouse
Redfish
CRDB
DMTF

Job description

  • As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers
  • The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack-scale systems
  • This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, where your work directly improves the reliability, safety, and scalability of real-world infrastructure.\
  • Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure
  • Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations
  • Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure
  • Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly
  • Translate incidents and hardware failure modes into software improvements that make the platform more resilient
  • Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale
  • Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship
  • Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden

Skilled in applying a data-driven approach to reliability, optimization, and continuous improvementB.S., M.S., or PhD in Computer Science or related field, or equivalent experience8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large-scale datacenter and cloud environmentsExpertise in Go and proven experience building REST/gRPC APIs for mission-critical platformsExcellent communicator able to work effectively with both technical and non-technical stakeholdersTrack record of leading incident response, postmortems, and driving robust service reliabilityProven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teamsHands-on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU serversStrong background in architecting and scaling cloud-native Kubernetes infrastructure and distributed servicesExperience contributing to and collaborating with open source communitiesWorking knowledge of Kafka, ClickHouse and CRDBDMTF, RedFish APIs, and GPU servers

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer (Infrastructure Engineering)
Senior Software Engineer (Infrastructure Engineering)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 130,000
Staff Software Engineer, GPU Data Center Automation
Staff Software Engineer, GPU Data Center Automation

CoreWeave • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Operations Engineering Manager (m/f/d)
Operations Engineering Manager (m/f/d)

Northern Data Group • Greater London

Hybrid
GBP 90,000 - 130,000
Staff Software Engineer (Data Infrastructure)
Staff Software Engineer (Data Infrastructure)

Peregrine • York and North Yorkshire

On-site
GBP 90,000 - 120,000
Health insurance
Unlimited PTO
401k
+5
Senior Software Infrastructure Engineer
Senior Software Infrastructure Engineer

EngineersOfAI • Greater London

On-site
GBP 40,000 - 60,000
Senior Software Infrastructure Engineer
Senior Software Infrastructure Engineer

EngineersOfAI • Cambridge

On-site
GBP 40,000 - 60,000
Software Engineer (Kubernetes)
Software Engineer (Kubernetes)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 150,000
Staff Software Engineer - Backend
Staff Software Engineer - Backend

ESL • Greater London

On-site
GBP 110,000 - 180,000
Engineering Manager (Fleet Engineering)
Engineering Manager (Fleet Engineering)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 120,000
Staff Software Engineer (Backend)
Staff Software Engineer (Backend)

FACEIT • Greater London

On-site
GBP 120,000 - 170,000
Flexible working hours
Gaming room
Daily lunches provided
+9