Senior Production Engineer: Reliability & Multi-Cloud Platform

Weights & Biases

New York (NY)

On-site

USD 139,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) match
Paid parental leave
Flexible PTO
Catered lunch

Job summary

CoreWeave is seeking a Senior Production Engineer to design, build, and operate the reliability platform spanning cloud and on‑prem environments. This hands‑on DevOps/SRE role focuses on automation, observability, incident management, and scalable infrastructure across AWS, GCP, Azure, and on‑premises.

You will lead technical initiatives, mentor teammates, and champion a sustainable on‑call culture. You’ll evolve systems to meet growing scale and performance needs, own canary deployments, and

Qualifications

  • Extensive engineering experience designing, building, deploying, and operating critical production infrastructure and services.
  • Expert in one or more major public clouds (AWS, GCP, or Azure), with on‑premises or hybrid experience.
  • Strong skills in infrastructure‑as‑code, automation, and configuration management (Terraform, CloudFormation, CDK, Ansible).
  • Hands‑on expertise with Kubernetes and containerized workloads in production.
  • Proficient in a systems scripting language (Go, Python, Bash) and comfortable writing tooling.
  • Deep experience with CI/CD and observability/monitoring tooling (Prometheus, Grafana, Datadog).
  • Proven ability to evolve designs for scale, reliability, and performance.
  • Comfortable owning on‑call for services you build and reducing on‑call burden through automation.

Responsibilities

  • Design, build, deploy, and operate reliable infrastructure across multi‑cloud and on‑prem environments.
  • Improve error attribution and alert routing to route issues to the owning team.
  • Own observability patterns, SLI/SLO frameworks, and dashboards for production visibility.
  • Build release‑safety systems including canaries, smoke tests, and safe rollouts.
  • Lead incidents and on‑call programs, driving tooling and runbooks.
  • Reduce on‑call burden via architecture and automation, participate in on‑call rotations.
  • Manage infrastructure with Terraform and push for IaC‑driven operations.
  • Solve large operational problems and translate them into actionable engineering work.
  • Provide technical leadership and influence direction across the team.

Skills

Production engineering
Cloud platforms
Infrastructure as code
Kubernetes
Go/Python scripting
CI/CD
Observability
Incident management

Tools

Terraform
Kubernetes
CI/CD (GitHub Actions)
Prometheus
Grafana
Datadog

Job description

CoreWeave is seeking a Senior Production Engineer to design, build, and operate the reliability platform spanning cloud and on‑prem environments. This hands‑on DevOps/SRE role focuses on automation, observability, incident management, and scalable infrastructure across AWS, GCP, Azure, and on‑premises.

You will lead technical initiatives, mentor teammates, and champion a sustainable on‑call culture. You’ll evolve systems to meet growing scale and performance needs, own canary deployments, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Production Engineer, Reliability Platform
Senior Production Engineer, Reliability Platform

Weights & Biases • San Francisco (CA)

On-site
USD 139,000 - 185,000
Medical, dental, vision insurance
401(k) with employer match
Flexible PTO
Senior Production Engineer: Reliability Platform & SRE
Senior Production Engineer: Reliability Platform & SRE

Weights & Biases • Livingston (NJ)

On-site
USD 139,000 - 185,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+1
Senior Production Engineer for Reliability & Observability
Senior Production Engineer for Reliability & Observability

Weights & Biases • Bellevue (WA)

Hybrid
USD 139,000 - 185,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
+5
Cloud Production Engineering Lead (SRE)
Cloud Production Engineering Lead (SRE)

CoreWeave • New York (NY)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
+9
Senior Production Engineering Leader - SRE & Platform
Senior Production Engineering Leader - SRE & Platform

Coreweave • Livingston (NJ)

On-site
USD 207,000 - 275,000
100% paid Medical, dental, and vision insurance
Flexible PTO
401(k) with employer match
Senior Cloud Reliability Engineer
Senior Cloud Reliability Engineer

Coreweave • New York (NY)

On-site
USD 139,000 - 204,000
Medical, dental, and vision insurance paid by CoreWeave
Company-paid Life Insurance
Flexible Spending Account
+3
Senior Platform Engineer: Scalable Data & Kubernetes
Senior Platform Engineer: Scalable Data & Kubernetes

CoreWeave Europe • San Francisco (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with match
Flexible PTO
+1
Senior Infrastructure Engineer — Reliability & Scale
Senior Infrastructure Engineer — Reliability & Scale

CoreWeave • Livingston (NJ)

On-site
USD 153,000 - 242,000
Medical insurance
401(k) with employer match
Paid Parental Leave
+3
Senior Platform Engineer: Scalable Data & Kubernetes
Senior Platform Engineer: Scalable Data & Kubernetes

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance
401(k) with employer match
Paid parental leave
+1
Senior Platform Engineer: Scalable Kubernetes & Data Infra
Senior Platform Engineer: Scalable Kubernetes & Data Infra

CoreWeave • New York (NY)

On-site
USD 182,000 - 242,000
Medical insurance
Dental & vision insurance
401(k) with employer match
+1