Senior Infrastructure Engineer — Reliability & Scale

CoreWeave

Livingston (NJ)

On-site

USD 153,000 - 242,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical insurance
401(k) with employer match
Paid Parental Leave
Flexible PTO
Tuition Reimbursement
Stock options / equity

Job summary

CoreWeave is seeking a Sr. Infrastructure Engineer to join our Hardware Engineering Dev team.

You will help design, deploy, and monitor services that manage our bare-metal infrastructure, collaborating with cross-functional teams and external partners to ensure high performance and reliability. You will own incident response, observability, and automation efforts, building scalable tooling and dashboards to support production operations while reducing on-call load and ensuring resilience across

Qualifications

  • 7+ years of experience in cloud operations, SRE, or related roles.
  • Understanding of cloud platforms (Kubernetes, AWS, GCP) and basic infra knowledge.
  • Familiarity with incident management practices and ITIL/SRE.
  • Proficiency with Go or Python.
  • Experience with Prometheus / Grafana.
  • Experience deploying containerized apps with Kubernetes.
  • Excellent documentation and attention to detail.
  • Strong analytical and problem‑solving abilities.
  • On-call rotation experience.

Responsibilities

  • Lead incident response, RCA, PIRs, and preventive improvements.
  • Own observability and health using Prometheus and Grafana.
  • Lead automation to detect and recover from incidents.
  • Define KPIs and SLAs aligned with reliability goals.
  • Collaborate to improve platform reliability and disaster recovery.
  • Build dashboards and alerts; participate in on-call rotation.

Skills

Cloud operations
SRE practices
Go or Python
Prometheus/Grafana
Kubernetes
AWS / GCP
On-call rotation
Documentation

Tools

Kubernetes
Go
Python
Prometheus
Grafana
AWS
GCP
CI/CD

Job description

CoreWeave is seeking a Sr. Infrastructure Engineer to join our Hardware Engineering Dev team.

You will help design, deploy, and monitor services that manage our bare-metal infrastructure, collaborating with cross-functional teams and external partners to ensure high performance and reliability. You will own incident response, observability, and automation efforts, building scalable tooling and dashboards to support production operations while reducing on-call load and ensuring resilience across

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infra & SRE Manager — Reliability at Scale
Infra & SRE Manager — Reliability at Scale

CoreWeave • New York (NY)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Infrastructure Reliability Engineering Manager
Infrastructure Reliability Engineering Manager

Socket.dev • New York (NY), Sunnyvale (CA), Livingston (NJ)

On-site
USD 182,000 - 242,000
Medical insurance
Dental insurance
Vision insurance
+8
Senior Production Engineer, Reliability Platform
Senior Production Engineer, Reliability Platform

Weights & Biases • San Francisco (CA)

On-site
USD 139,000 - 185,000
Medical, dental, vision insurance
401(k) with employer match
Flexible PTO
Senior Production Engineer for Reliability & Observability
Senior Production Engineer for Reliability & Observability

Weights & Biases • Bellevue (WA)

Hybrid
USD 139,000 - 185,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
+5
Infra Engineering Manager — Reliability & SRE Lead
Infra Engineering Manager — Reliability & SRE Lead

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2
Senior Production Engineer: Reliability & Multi-Cloud Platform
Senior Production Engineer: Reliability & Multi-Cloud Platform

Weights & Biases • New York (NY)

On-site
USD 139,000 - 185,000
Health insurance
401(k) match
Paid parental leave
+2
Infrastructure Engineering Manager, Bare-Metal Reliability
Infrastructure Engineering Manager, Bare-Metal Reliability

PVH (Tommy Hilfiger/Calvin Klein) • Livingston (NJ)

On-site
USD 140,000 - 180,000
Infra Engineering Manager: Scale Reliability for Bare-Metal
Infra Engineering Manager: Scale Reliability for Bare-Metal

Coreweave • West Covina (CA)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance - {
401(k) with employer match
Tuition Reimbursement
+5
Staff Platform Engineer: Scale & Reliability
Staff Platform Engineer: Scale & Reliability

CoreWeave • Livingston (NJ)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance—
Company-paid Life Insurance
Tuition Reimbursement
+4
Infra Support Engineering Lead — 24/7 Ops
Infra Support Engineering Lead — 24/7 Ops

Neura Market • Livingston (NJ), Northern (KY)

Hybrid
USD 198,000 - 264,000