Production Engineering Leader: SRE & Reliability

CoreWeave

Bellevue (WA)

On-site

USD 207,000 - 275,000

Full time

33 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Life Insurance
Disability insurance
Flexible Spending Account
Health Savings Account
Tuition Reimbursement
Employee Stock Purchase Program (ESPP)
Mental Wellness benefits

Job summary

CoreWeave is seeking a Senior Manager of Production Engineering to lead and expand the SRE team for its cloud platform. You will shape resilience, reliability, and efficiency while partnering with engineering and product teams to build secure, scalable systems.

You will champion automation-first practices, drive postmortems, and evolve on-call strategies to support a global 24x7 environment, shaping a culture of ownership and continuous learning.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related fields.
  • 10+ years in a leadership or senior management role at a cloud provider, hyperscaler, or high-growth tech company.
  • Experience hiring, developing, and managing geographically distributed 24x7 engineering teams.
  • Experience in designing and implementing incident management processes including on-call rotations, escalation paths, postmortems, and SLO/SLA framework.
  • Solid foundation in systems engineering, with a deep understanding of distributed systems, networking, and storage architecture.
  • Strong cross-functional collaboration skills, with the ability to influence product, platform, hardware, and security teams.

Responsibilities

  • Contribute and execute the SRE vision, strategy, and roadmap for a large-scale, distributed cloud infrastructure.
  • Lead and mentor a high-performing team of SREs, promoting a culture of ownership, collaboration, and continuous learning.
  • Champion automation-first practices, leveraging AI, tools like Terraform, Kubernetes, and Infrastructure-as-Code to minimize toil and manual interventions.
  • Establish and evolve Operational Excellence best practices ensuring the platform is proactive, and propagates organizational learning.
  • Drive initiatives for incident management, postmortem culture, root cause analysis, and system hardening.
  • Collaborate with engineering, product, and customer support teams to build scalable, resilient, and self-healing systems.
  • Evolve our on-call strategy and processes to support a 24x7, globally distributed platform with minimal disruptions.

Skills

SRE leadership
Incident management
Hybrid cloud
Automation
Mentorship
Cross-functional collaboration
Terraform
Kubernetes
Infrastructure as Code

Education

Bachelor’s degree in Computer Science, Engineering, or related fields

Tools

Terraform
Kubernetes
Infrastructure as Code

Job description

CoreWeave is seeking a Senior Manager of Production Engineering to lead and expand the SRE team for its cloud platform. You will shape resilience, reliability, and efficiency while partnering with engineering and product teams to build secure, scalable systems.

You will champion automation-first practices, drive postmortems, and evolve on-call strategies to support a global 24x7 environment, shaping a culture of ownership and continuous learning.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud Production Engineering Lead (SRE)
Cloud Production Engineering Lead (SRE)

CoreWeave • New York (NY)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
+9
Global Production & Reliability Leader
Global Production & Reliability Leader

Everbridge • Northern (KY)

Hybrid
USD 195,000 - 270,000
Production Reliability Engineer (SRE & Automation)
Production Reliability Engineer (SRE & Automation)

NRnP Technology • Northern (KY)

Hybrid
USD 90,000 - 140,000
Senior Production SRE: Cloud & On-Prem Reliability
Senior Production SRE: Cloud & On-Prem Reliability

Weights & Biases • New York (NY)

On-site
USD 140,000 - 180,000
Medical Insurance
Dental Insurance
Vision Insurance
+15
Senior SRE & Cloud Reliability Architect
Senior SRE & Cloud Reliability Architect

United States Digital Space LLC • United States

Remote
USD 180,000 - 250,000
Operations Manager, Fleet Reliability — 24/7 & Automation
Operations Manager, Fleet Reliability — 24/7 & Automation

CoreWeave • New York (NY)

On-site
USD 143,000 - 191,000
Medical, dental, and vision insurance
Equity awards
401(k) with generous match
+2
VP, Global Production Ops & Reliability Scale & Resilience
VP, Global Production Ops & Reliability Scale & Resilience

Everbridge • United States

Remote
USD 195,000 - 270,000
Health insurance
Dental coverage
Parental planning
+6
Senior SRE Lead: Reliability, Automation & Incidents
Senior SRE Lead: Reliability, Automation & Incidents

Shield AI • San Diego (CA)

On-site
USD 183,000 - 275,000
Equity
Bonus
Benefits
Senior Infrastructure Engineer — AI Cloud Reliability
Senior Infrastructure Engineer — AI Cloud Reliability

Socket.dev • New York (NY)

On-site
USD 182,000 - 242,000
Medical/Dental/Vision
Life Insurance
Tuition Reimbursement
+5
Senior Cloud SRE & 24x7 Incident Resilience Engineer
Senior Cloud SRE & 24x7 Incident Resilience Engineer

Encora • United States

On-site
USD 90,000 - 130,000