Senior Production Engineer for Reliability & Observability

Weights & Biases

Bellevue (WA)

Hybrid

USD 139,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
Tuition Reimbursement
401(k) with employer match
Paid Parental Leave
Catered lunch daily
Casual work environment

Job summary

CoreWeave is hiring a Senior Production Engineer to design, build, and operate the reliability platform underpinning global cloud and on‑prem deployments. You will own observability, incident response, and automation across AWS, GCP, Azure, and self‑hosted environments, with leadership to guide cross‑functional teams.

You will evolve systems for scale and reliability, drive canary deployments, and champion on‑call improvements.

Qualifications

  • Extensive experience designing, building, deploying, and operating production infrastructure.
  • Public cloud expertise (AWS, GCP, or Azure) with on‑premises or hybrid experience.
  • Proficient in infrastructure‑as‑code and automation.
  • Hands‑on Kubernetes and production workloads.
  • Strong scripting skills (Go, Python, Bash).
  • CI/CD and observability tooling experience.
  • Ability to lead technically and influence direction.
  • Comfortable on‑call and reducing toil.

Responsibilities

  • Design, build, deploy, and operate reliability/infrastructure services across cloud and on‑prem.
  • Improve error attribution and alert routing.
  • Own observability patterns and SLI/SLOs, dashboards, and service catalogs.
  • Build release‑safety systems for fast and safe deployments.
  • Advance incident and on‑call program and runbooks.
  • Reduce on‑call burden via architecture and automation.
  • Provide technical leadership across the team.

Skills

Cloud & on-prem deployments
Kubernetes & containers
CI/CD & observability
Multi-cloud (AWS/GCP/Azure)
SRE / DevOps discipline

Tools

Terraform
CloudFormation/CDK
Ansible
Prometheus
Grafana
Datadog

Job description

CoreWeave is hiring a Senior Production Engineer to design, build, and operate the reliability platform underpinning global cloud and on‑prem deployments. You will own observability, incident response, and automation across AWS, GCP, Azure, and self‑hosted environments, with leadership to guide cross‑functional teams.

You will evolve systems for scale and reliability, drive canary deployments, and champion on‑call improvements.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Production Engineer: Reliability Platform & SRE
Senior Production Engineer: Reliability Platform & SRE

Weights & Biases • Livingston (NJ)

On-site
USD 139,000 - 185,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+1
Senior Production Engineer: Reliability & Multi-Cloud Platform
Senior Production Engineer: Reliability & Multi-Cloud Platform

Weights & Biases • New York (NY)

On-site
USD 139,000 - 185,000
Health insurance
401(k) match
Paid parental leave
+2
Senior Infrastructure Engineer — Reliability & Scale
Senior Infrastructure Engineer — Reliability & Scale

CoreWeave • Livingston (NJ)

On-site
USD 153,000 - 242,000
Medical insurance
401(k) with employer match
Paid Parental Leave
+3
Senior Observability Platform Lead
Senior Observability Platform Lead

Coreweave • New York (NY)

On-site
USD 188,000 - 275,000
Health Insurance
Life Insurance
Tuition Reimbursement
+5
Cloud Production Engineering Lead (SRE)
Cloud Production Engineering Lead (SRE)

CoreWeave • New York (NY)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
+9
Senior Production Engineering Leader - SRE & Platform
Senior Production Engineering Leader - SRE & Platform

Coreweave • Livingston (NJ)

On-site
USD 207,000 - 275,000
100% paid Medical, dental, and vision insurance
Flexible PTO
401(k) with employer match
Senior Infrastructure Engineer — AI Cloud Reliability
Senior Infrastructure Engineer — AI Cloud Reliability

Socket.dev • New York (NY)

On-site
USD 182,000 - 242,000
Medical/Dental/Vision
Life Insurance
Tuition Reimbursement
+5
Senior Site Reliability Engineer — Platform & Observability
Senior Site Reliability Engineer — Platform & Observability

Jobtailor • North Carolina

On-site
USD 180,000 - 240,000
Senior Platform Engineer - Kubernetes & Data Infra
Senior Platform Engineer - Kubernetes & Data Infra

CoreWeave • San Francisco (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance 
Company‑paid Life Insurance
Voluntary life insurance
+14
Senior Platform Engineer: Scalable Data & Kubernetes
Senior Platform Engineer: Scalable Data & Kubernetes

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance
401(k) with employer match
Paid parental leave
+1