Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B

Weights & Biases

Bellevue (WA)

Hybrid

USD 139,000 - 185,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
Tuition Reimbursement
401(k) with employer match
Paid Parental Leave
Catered lunch daily
Casual work environment

Job summary

CoreWeave is hiring a Senior Production Engineer to design, build, and operate the reliability platform underpinning global cloud and on‑prem deployments. You will own observability, incident response, and automation across AWS, GCP, Azure, and self‑hosted environments, with leadership to guide cross‑functional teams.

You will evolve systems for scale and reliability, drive canary deployments, and champion on‑call improvements.

Qualifications

  • Extensive experience designing, building, deploying, and operating production infrastructure.
  • Public cloud expertise (AWS, GCP, or Azure) with on‑premises or hybrid experience.
  • Proficient in infrastructure‑as‑code and automation.
  • Hands‑on Kubernetes and production workloads.
  • Strong scripting skills (Go, Python, Bash).
  • CI/CD and observability tooling experience.
  • Ability to lead technically and influence direction.
  • Comfortable on‑call and reducing toil.

Responsibilities

  • Design, build, deploy, and operate reliability/infrastructure services across cloud and on‑prem.
  • Improve error attribution and alert routing.
  • Own observability patterns and SLI/SLOs, dashboards, and service catalogs.
  • Build release‑safety systems for fast and safe deployments.
  • Advance incident and on‑call program and runbooks.
  • Reduce on‑call burden via architecture and automation.
  • Provide technical leadership across the team.

Skills

Cloud & on-prem deployments
Kubernetes & containers
CI/CD & observability
Multi-cloud (AWS/GCP/Azure)
SRE / DevOps discipline

Tools

Terraform
CloudFormation/CDK
Ansible
Prometheus
Grafana
Datadog

Job description

CoreWeave, the AI Hyperscaler™, acquired Weights & Biases to create the most powerful end-to-end platform to develop, deploy, and iterate AI faster. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe, and was ranked as one of the TIME100 most influential companies of 2024. By bringing together CoreWeave’s industry-leading cloud infrastructure with the best-in-class tools AI practitioners know and love from Weights & Biases, we’re setting a new standard for how AI is built, trained, and scaled.

The integration of our teams and technologies is accelerating our shared mission: to empower developers with the tools and infrastructure they need to push the boundaries of what AI can do. From experiment tracking and model optimization to high-performance training clusters, agent building, and inference at scale, we’re combining forces to serve the full AI lifecycle — all in one seamless platform.

Weights & Biases has long been trusted by over 1,500 organizations — including AstraZeneca, Canva, Cohere, OpenAI, Meta, Snowflake, Square, Toyota, and Wayve — to build better models, AI agents and applications. Now, as part of CoreWeave, that impact is amplified across a broader ecosystem of AI innovators, researchers, and enterprises.

As we unite under one vision, we’re looking for bold thinkers and agile builders who are excited to shape the future of AI alongside us. If you're passionate about solving complex problems at the intersection of software, hardware, and AI, there's never been a more exciting time to join our team.

What You’ll Do

The Production Engineering team builds and operates the platform that lets CoreWeave's engineers ship software quickly, reliably, and safely. We own the observability systems, reliability tooling, incident management systems, and the infrastructure-as-code that underpins our engineering organization's ability to deliver software to our enterprise customers across GCP, AWS, Azure, and self-hosted, on-premises environments. Our mission is simple: you build it, you deploy it, you run it – and we build the tools and processes that make owning your service in production straightforward and safe.

About The Role

We are seeking a Senior Production Engineer with deep expertise across both cloud and on-premises environments to design, build, and operate the reliability platform at the core of CoreWeave's engineering organization. This is a hands‑on DevOps/SRE‑flavored role: you'll work across infrastructure-as-code, CI/CD, observability, and incident systems – automating away toil and giving service teams the visibility and guardrails they need to run their own code in production.

You'll evolve our systems to meet increasingly challenging scale, reliability, and performance demands, and own large, ambiguous operational problems end to end – from error attribution and alert routing to release safety and capacity planning. You'll also provide technical leadership across the team, partnering with senior and principal engineers to influence direction and drive outcomes for cross‑functional stakeholders. We care deeply about a sustainable on‑call culture, so you'll participate in on‑call rotations while championing the architecture and automation that reduce on‑call burden over time.

In This Role, You Will
  • Design, build, deploy, and operate critical reliability and infrastructure services across AWS, GCP, Azure, and on‑premises / hybrid environments.
  • Improve error attribution and alert routing, automatically routing errors and pages to the team that owns the affected service, so engineers are only on‑call for what they own.
  • Own observability patterns, SLI/SLO frameworks, service catalog metadata, and dashboards that give teams visibility into their services from code change through production.
  • Build release‑safety systems – canary deployments, smoke tests, staged rollouts, and reliable roll‑back / roll‑forward – so shipping to production is fast and safe.
  • Advance the incident and on‑call program – incident tooling, on‑call rotations, runbooks, and operational readiness reviews; measure incident volume by team to focus reliability investment.
  • Reduce on‑call burden through better architecture, automation, and observability – and participate in on‑call rotations yourself.
  • Provision and manage infrastructure with Terraform and drive manual, incident‑time operations toward automated, repeatable infrastructure‑as‑code.
  • Break down and solve large, ambiguous operational problems, turning them into well‑scoped, shippable engineering work.
  • Provide technical leadership and mentorship, work alongside senior and principal engineers, and influence the broader technical direction of the platform.
Who You Are
  • Extensive engineering experience designing, building, deploying, and operating critical production infrastructure and services across a large enterprise.
  • Expert in one or more major public clouds (AWS, GCP, or Azure), with real experience operating on‑premises or hybrid environments, and in multi‑cloud environments with multi‑account strategies
  • Strong skills in infrastructure‑as‑code, automation, and configuration management (Terraform, CloudFormation, CDK, Ansible, or similar).
  • Hands‑on expertise with Kubernetes and containerized workloads in production.
  • Proficient in a systems/scripting language (Go, Python, Bash, or similar) and comfortable writing tooling and automation, not just operating it.
  • Deep experience with CI/CD (GitHub Actions) and observability / monitoring tooling (Prometheus, Grafana, Datadog, etc.).
  • Proven ability to evolve designs to meet increasingly challenging scale, reliability, and performance requirements.
  • Comfortable owning on‑call for services you build, and passionate about reducing on‑call burden through architecture and automation rather than heroics.
  • Proven ability to provide technical leadership, work effectively with senior and principal engineers, influence technical direction, and contribute to the success of cross‑functional stakeholders.
Preferred
  • Experience building and operating SaaS products.
  • Experience building and operating reliability, incident, or developer‑productivity platforms (service catalogs / Backstage, PagerDuty tooling, SLO frameworks).
  • Experience operating stateful systems and databases in production (PostgreSQL, MySQL, ClickHouse, etc.).
  • Cloud certifications (AWS Solutions Architect, Azure Architect, etc.).
  • Familiarity with security and access‑management best practices in hybrid environments.
Wondering if you’re a good fit?
  • You love making production systems reliable, observable, and boring – and giving other engineers the tools to run their own services confidently.
  • You're curious about reducing on‑call burden through better architecture and automation rather than more heroics.
  • You're comfortable across the DevOps stack – cloud and on‑prem, Kubernetes, Terraform, CI/CD, and observability – and you like automating away toil.
Why CoreWeave?

At CoreWeave, we work hard, have fun, and move fast! We’re in an exciting stage of hyper‑growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values:

About
  • Be Curious at Your Core Act Like an Owner
  • Empower Employees
  • Deliver Best‑in‑Class Client Experiences
  • Achieve More Together

We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

California Applicants
California Consumer Privacy Act
Equal Opportunity & Accommodations

CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.

As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: careers@coreweave.com.

Export Control Compliance

This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. 1157, or (iv) asylee under 8 U.S.C. 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.

  • 1157, or (iv) asylee under 8 U.S.C. 1158.
  • 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency.
Base Salary and Benefits

The base salary range for this role is $139,000 to $185,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

What We Offer

The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location.

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US‑based offerings; for roles in other locations, benefits vary and are shared during the hiring process. These include:

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company‑paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long‑term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family‑Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full‑service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption
"
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B
Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B

Weights & Biases • New York (NY)

On-site
USD 139,000 - 185,000
Health insurance
401(k) match
Paid parental leave
+2
Senior Software Engineer, W&B Agent - W&B
Senior Software Engineer, W&B Agent - W&B

Weights & Biases • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Senior Software Engineer, Infrastructure Operations - Weights & Biases
Senior Software Engineer, Infrastructure Operations - Weights & Biases

Weights & Biases • Sunnyvale (CA)

On-site
USD 153,000 - 204,000
401(k) with generous match
Medical, dental, vision insurance
Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B
Senior Software Engineer, Production Engineering (Cloud & On-Prem) - W&B

Weights & Biases • San Francisco (CA)

On-site
USD 139,000 - 185,000
Medical, dental, vision insurance
401(k) with employer match
Flexible PTO
Senior Software Engineer, Infrastructure Operations - Weights & Biases
Senior Software Engineer, Infrastructure Operations - Weights & Biases

Weights & Biases • New York (NY)

On-site
USD 153,000 - 204,000
Medical, dental, and vision insurance
Life Insurance
Equity awards
+3
Senior Software Engineer, Infrastructure Operations - Weights & Biases
Senior Software Engineer, Infrastructure Operations - Weights & Biases

Weights & Biases • Livingston (NJ)

On-site
USD 153,000 - 204,000
Medical, dental, and vision insurance
401(k) with employer match
Paid Parental Leave
+2
Staff Software Engineer, Infrastructure Engineering
Staff Software Engineer, Infrastructure Engineering

CoreWeave • New York (NY)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Staff Software Engineer
Staff Software Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+3
Staff Software Engineer
Staff Software Engineer

CoreWeave • New York (NY)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with employer match
+2
Staff Software Engineer
Staff Software Engineer

CoreWeave • San Francisco (CA)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2