Get more replies from employers
Send a job-specific resume in minutes.
CoreWeave is seeking a Senior Site Reliability Engineer on the MetalDev team to balance production operations (60%) and engineering automation (40%). You will lead incident response, troubleshoot issues, perform root-cause analyses, and participate in on-call rotations, while writing resilient Go code and building dashboards.
You will define SLOs, improve CI/CD pipelines, and create self-service tooling for Fleet Operations and Hardware engineering teams, helping scale our data centre
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com .
We're proud to be a Living Wage accredited Employer.
The MetalDev team within CoreWeave's Hardware Compute organisation develops software automation tooling and services used to bring up data centre rack systems and manage bare-metal infrastructure. We provide core reliability, availability, and operational stability functions across regional data centres to ensure seamless infrastructure provisioning.
As a Senior Site Reliability Engineer on the MetalDev team, you will split your focus between production operations and reliability (60%) and engineering automation (40%). You will lead incident response, troubleshooting, root-cause analyses, and post-incident reviews while participating in an on-call rotation. In this senior role, you will write resilient Go code, build Prometheus and Grafana dashboards, and develop automated remediation workflows to reduce manual overhead across our fleet. Additionally, you will define SLOs and KPIs, improve CI/CD deployment pipelines, and create self-service tooling for Fleet Operations and Hardware engineering teams.
Bachelor's degree in Computer Science, Engineering, or a related technical field (or equivalent practical experience).
5+ years of experience in Site Reliability Engineering, production engineering, cloud infrastructure, or software engineering.
Working proficiency in Go with experience developing production-quality software.
Hands-on production experience with Kubernetes and containerised microservices architectures.
Experience with observability and telemetry stacks, specifically Prometheus and Grafana.
Demonstrated track record supporting production services, leading incident management, and participating in on-call rotations.
Excellent troubleshooting, analytical, and technical documentation skills.
Experience managing or automating bare-metal infrastructure.
Familiarity with BMCs, Redfish, or server-management technologies.
Experience building automated remediation or self-healing systems.
Familiarity with public cloud platforms such as AWS or GCP.
We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams-even if you aren't a 100% skill or experience match.
You love to: Build automated remediation workflows and eliminate operational toil across bare-metal infrastructure.
You're curious about: Developing low-latency telemetry pipelines and scaling self-healing systems across massive data centre clusters.
You're an expert in: Production incident management, Go programming, and designing robust Kubernetes-native observability tools.
At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:
Be Curious at Your Core
Act Like an Owner
Empower Employees
Deliver Best-in-Class Client Experiences
Achieve More Together
We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for take-off, the organisation's growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!
We're hiring across multiple levels. Typical cash compensation ranges from ~262,000-350,000 PLN , with additional performance based bonus & equity that can significantly increase total compensation. The starting salary will be determined by job-related knowledge, skills, experience, and the market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).
To fulfill our obligation to protect client data, successful applicants offered employment with CoreWeave will be required to complete a basic criminal record check, conducted in compliance with GDPR. Employment offers are conditional upon receiving satisfactory check results.
In addition to a com