Senior Software Engineer (Infrastructure Engineering)

CoreWeave

York and North Yorkshire

On-site

GBP 90,000 - 130,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

CoreWeave is seeking a highly skilled Sr. Infrastructure Engineer to join our Hardware Engineering Dev team. You will own incidents, reliability, and platform observability across production services and bare-metal infrastructure.

Collaborate with cross-functional teams, external vendors, and other stakeholders to deliver scalable, high-performance infrastructure. Build automation, CI/CD pipelines, and proactive monitoring.

Qualifications

  • 7+ years in cloud operations, SRE, or related roles.
  • Experience deploying containerized apps with Kubernetes.
  • Familiarity with incident management frameworks (ITIL, SRE).
  • Proficiency with Prometheus and Grafana.
  • Strong Go or Python coding skills.

Responsibilities

  • Lead incident response and post-incident reviews.
  • Own observability strategy and health metrics.
  • Design scalable, reliable infrastructure for bare-metal services.
  • Develop automation and CI/CD pipelines.
  • Collaborate with cross-functional teams and vendors.
  • Participate in on-call rotation.

Skills

SRE experience
Kubernetes
Incident management
Prometheus/Grafana
Go/Python
On-call
Documentation

Tools

Prometheus
Grafana
AWS
GCP
CI/CD

Job description

Responsibilities
  • CoreWeave is seeking a highly skilled and motivated Sr. Infrastructure Engineer to join our Hardware Engineering Dev team ( Metal Dev)
  • Reporting to the Engineering Manager for Hardware Engineering Dev, you will play a crucial part in the development, deployment, and monitoring of services that manage our bare-metal infrastructure
  • You will collaborate closely with cross-functional teams, external vendors, and other stakeholders to ensure the successful delivery of highly performant and reliable infrastructure solutions
  • Incident Management & Support:
  • Lead incident response efforts by identifying and resolving service disruptions quickly, while coaching other junior team members through resolution
  • Lead the documentation of incidents, conduct in-depth root cause analysis (RCA), and drive post-incident reviews (PIRs) to identify systemic issues. Implement long term improvements that would prevent service degradation
  • Own the development and continuous improvement of incident response playbooks ensuring preparedness for a wide range of failure scenarios
  • Clearly communicate efforts during incidents to the management, stakeholders and the cross functional teams, during an incident. Keep clear records of incident activities
  • Master clear understanding of various services on how they work in production as well as build through knowledge of the internals of these services and how they interact with the entire stack
  • Operational Support & Reliability
  • Build a strategy around making our core services perform at its best at scale. This includes improvements to the services for robustness as well as supportability in production
  • Own system observability and health leveraging tools like Prometheus and Grafana, to proactively detect performance bottlenecks and prevent incidents
  • Lead automation efforts to streamline incident detection and recovery, minimizing manual intervention
  • Define and drive KPIs and SLAs for incident management and ensuring alignment with the organizational reliability objectives
  • Collaborate with engineers across teams to improve platform reliability, resilience improvements, and disaster recovery
  • Collaborate with upstream communities, including Go and Redfish-based services
  • Design and implement solutions to build operational efficiency and stability
  • Document hardware automation workflows and processes
  • Create CI/CD pipelines
  • Ensure smooth operation of all aspects of the server hardware lifecycle, from provisioning to end-of-life, by troubleshooting bugs, automating common tasks, and documenting processes
  • Partner with the Fleet Operations Team to design scalable tooling and processes that enables self-service and reduction in escalation overhead
  • Build out dashboards and alerts to make efficient operational troubleshooting
  • Participate in on-call rotation as well as triage issues that are posted in support channels on an ongoing basis
  • This role is responsible for reducing on-call queries and incidents over time
Requirements
  • Excellent documentation skills and attention to detail
  • 7+ years of experience in cloud operations, site reliability engineering (SRE), or related technical roles
  • Previous experience deploying containerized applications using Kubernetes
  • Familiarity with incident management practices and frameworks (e.g., ITIL, SRE best practices)
  • Prior experience with Prometheus / Grafana
  • Understanding of cloud platforms (e.g., Kubernetes, AWS, GCP) and basic knowledge of cloud infrastructure
  • Proficiency with Go or Python
  • Strong analytical and problem-solving abilities
  • Served on an on-call rotation supporting production services
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer (Kubernetes)
Software Engineer (Kubernetes)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 150,000
Senior Software Engineer (Server Fleet Infrastructure)
Senior Software Engineer (Server Fleet Infrastructure)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 130,000
Engineering Manager (Fleet Engineering)
Engineering Manager (Fleet Engineering)

CoreWeave • York and North Yorkshire

On-site
GBP 90,000 - 120,000
Infrastructure Engineer
Infrastructure Engineer

Jobtailor • Greater London

Hybrid
GBP 90,000 - 130,000
Staff Software Engineer (MetalDev)
Staff Software Engineer (MetalDev)

CoreWeave • York and North Yorkshire

On-site
GBP 120,000 - 180,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Flowcode • York and North Yorkshire

On-site
GBP 90,000 - 120,000
Unlimited Vacation
Health benefits
Stock options
+6
Senior Site Reliability Engineer
Senior Site Reliability Engineer

P2P • Greater London

On-site
GBP 90,000 - 130,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Brevan Howard • Greater London

On-site
GBP 90,000 - 130,000
Senior Solutions Engineer
Senior Solutions Engineer

Kroll • United Kingdom

On-site
GBP 80,000 - 110,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Selby Jennings • Greater London

On-site
GBP 70,000 - 90,000