Senior Site Reliability Engineer, Infrastructure Engineering

coreweaveu

Warszawa

On-site

PLN 262,000 - 350,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

CoreWeave is seeking a Senior Site Reliability Engineer on the MetalDev team to balance production operations (60%) and engineering automation (40%). You will lead incident response, troubleshoot issues, perform root-cause analyses, and participate in on-call rotations, while writing resilient Go code and building dashboards.

You will define SLOs, improve CI/CD pipelines, and create self-service tooling for Fleet Operations and Hardware engineering teams, helping scale our data centre

Qualifications

  • 5+ years of experience in Site Reliability Engineering, production engineering, cloud infrastructure, or software engineering.
  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
  • Working proficiency in Go with production-quality software.
  • Hands-on production experience with Kubernetes and containerised microservices.
  • Experience with Prometheus and Grafana for observability.

Responsibilities

  • Lead incident response, troubleshooting, root-cause analyses, and post-incident reviews.
  • Write resilient Go code and build Prometheus and Grafana dashboards.
  • Develop automated remediation workflows to reduce manual overhead.
  • Define SLOs and KPIs and improve CI/CD deployment pipelines.
  • Create self-service tooling for Fleet Operations and Hardware engineering teams.

Skills

Go
Kubernetes
Prometheus
Grafana
Incident response
On-call experience
Observability
Production engineering
Cloud infrastructure

Education

Bachelor's degree in Computer Science or related field

Tools

Redfish/BMCs
Self-healing automation tools
CI/CD tooling

Job description

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com .

We're proud to be a Living Wage accredited Employer.

What You'll Do

The MetalDev team within CoreWeave's Hardware Compute organisation develops software automation tooling and services used to bring up data centre rack systems and manage bare-metal infrastructure. We provide core reliability, availability, and operational stability functions across regional data centres to ensure seamless infrastructure provisioning.

About the role

As a Senior Site Reliability Engineer on the MetalDev team, you will split your focus between production operations and reliability (60%) and engineering automation (40%). You will lead incident response, troubleshooting, root-cause analyses, and post-incident reviews while participating in an on-call rotation. In this senior role, you will write resilient Go code, build Prometheus and Grafana dashboards, and develop automated remediation workflows to reduce manual overhead across our fleet. Additionally, you will define SLOs and KPIs, improve CI/CD deployment pipelines, and create self-service tooling for Fleet Operations and Hardware engineering teams.

Who You Are

Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).

5+ years of experience in Site Reliability Engineering, production engineering, cloud infrastructure, or software engineering.

Working proficiency in Go with experience developing production-quality software.

Hands-on production experience with Kubernetes and containerised microservices architectures.

Experience with observability and telemetry stacks, specifically Prometheus and Grafana.

Demonstrated track record supporting production services, leading incident management, and participating in on-call rotations.

Excellent troubleshooting, analytical, and technical documentation skills.

Preferred:

Experience managing or automating bare-metal infrastructure.

Familiarity with BMCs, Redfish, or server-management technologies.

Experience building automated remediation or self-healing systems.

Familiarity with public cloud platforms such as AWS or GCP.

Wondering if you're a good fit?

We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams-even if you aren't a 100% skill or experience match.

You love to: Build automated remediation workflows and eliminate operational toil across bare-metal infrastructure.

You're curious about: Developing low-latency telemetry pipelines and scaling self-healing systems across massive data centre clusters.

You're an expert in: Production incident management, Go programming, and designing robust Kubernetes-native observability tools.

Why CoreWeave?

At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

Be Curious at Your Core

Act Like an Owner

Empower Employees

Deliver Best-in-Class Client Experiences

Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for take-off, the organisation's growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

We're hiring across multiple levels. Typical cash compensation ranges from ~262,000-350,000 PLN , with additional performance based bonus & equity that can significantly increase total compensation. The starting salary will be determined by job-related knowledge, skills, experience, and the market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

To fulfill our obligation to protect client data, successful applicants offered employment with CoreWeave will be required to complete a basic criminal record check, conducted in compliance with GDPR. Employment offers are conditional upon receiving satisfactory check results.

What We Offer

In addition to a com

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, Infrastructure Engineering
Site Reliability Engineer, Infrastructure Engineering

coreweaveu • Warszawa

On-site
PLN 223,000 - 298,000
Living Wage accreditation
Site Reliability Engineer, Infrastructure Engineering
Site Reliability Engineer, Infrastructure Engineering

CoreWeave • Warszawa

On-site
PLN 223,000 - 298,000
Living Wage Accredited Employer
Site Reliability Engineer, Infrastructure Engineering
Site Reliability Engineer, Infrastructure Engineering

CoreWeave Europe • Warszawa

On-site
PLN 223,000 - 298,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+2
Software Engineer, OS Automation Platform
Software Engineer, OS Automation Platform

coreweaveu • Warszawa

On-site
PLN 262,000 - 350,000
Hardware Engineer, Server Infrastructure
Hardware Engineer, Server Infrastructure

CoreWeave • Warszawa

On-site
PLN 262,000 - 350,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+4
Senior Software Engineer, Network Development
Senior Software Engineer, Network Development

Coreweaveu • Warszawa

On-site
PLN 98,000 - 130,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Hardware Engineer, Server Infrastructure
Hardware Engineer, Server Infrastructure

coreweaveu • Warszawa

On-site
PLN 180,000 - 300,000
Fleet Engineering Project Manager, Data Center
Fleet Engineering Project Manager, Data Center

CoreWeave • Warszawa

On-site
PLN 189,000 - 252,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Senior Software Engineer, Server Fleet Infrastructure
Senior Software Engineer, Server Fleet Infrastructure

CoreWeave • Warszawa

On-site
PLN 321,000 - 428,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5
Senior Software Engineer, Box Office Platform
Senior Software Engineer, Box Office Platform

CoreWeave • Warszawa

On-site
PLN 321,000 - 428,000
Family-level Medical Insurance
Family-level Dental Insurance
Generous Pension Contribution
+5