Staff Software Engineer (Compute Architecture)

Coreweave

New York (NY)

On-site

USD 188,000 - 275,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Tuition reimbursement
Flexible PTO
401(k) with employer match

Job summary

Coreweave is seeking a Staff Software Engineer to join the Compute Architecture team in New York. In this role, you will design, build, and operate Go-based services that manage GPU data center infrastructure and ensure production reliability.

The ideal candidate will have 8+ years of experience, strong expertise in Go, and a solid background in cloud engineering. We offer a competitive salary range, comprehensive benefits, and a supportive workplace culture.

Qualifications

  • 8+ years of software engineering experience with a focus on infrastructure.
  • Expertise in Go and experience building REST/gRPC APIs.
  • Strong background in cloud-native Kubernetes infrastructure.

Responsibilities

  • Design and operate Go-based services for GPU data centers.
  • Build automation for data center operations and hardware discovery.
  • Translate incidents into software improvements for reliability.

Skills

Go
Cloud engineering
Distributed databases
Incident response
Mentoring
Observability stacks

Education

B.S., M.S., or PhD in Computer Science or related field

Tools

Kubernetes
Grafana
Prometheus

Job description

About the Role

As a Staff Software Engineer within our Compute Architecture organization, you will help build the software systems that operate the backbone of our large-scale GPU data centers. The METALDEV team builds Go-based distributed services that bring new infrastructure online, manage hardware lifecycle workflows, monitor production health, and automate safe operations across fleets of GPU servers and rack‑scale systems. This is a software-first role at the intersection of distributed systems, production reliability, and hardware‑aware automation, where your work directly improves the reliability, safety, and scalability of real‑world infrastructure.

What You’ll Do
  • Design, build, and operate Go-based services that manage the lifecycle of large‑scale GPU data center infrastructure.
  • Build automation for data center bring‑up, hardware discovery, health monitoring, remediation, and production operations.
  • Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack‑level infrastructure.
  • Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly.
  • Translate incidents and hardware failure modes into software improvements that make the platform more resilient.
  • Partner with hardware‑adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale.
  • Provide technical leadership through design reviews, code reviews, architectural guidance, and mentorship.
  • Make pragmatic architecture decisions that balance reliability, simplicity, scalability, and operational burden.
Who You Are
  • B.S., M.S., or PhD in Computer Science or related field, or equivalent experience.
  • 8+ years of software engineering experience with a strong focus on infrastructure, cloud engineering, and distributed databases—particularly within large‑scale datacenter and cloud environments.
  • Expertise in Go and proven experience building REST/gRPC APIs for mission‑critical platforms.
  • Strong background in architecting and scaling cloud‑native Kubernetes infrastructure and distributed services.
  • Proven success in mentoring engineers, leading technical projects, and influencing engineering strategy across teams.
  • Experience contributing to and collaborating with open source communities.
  • Skilled in applying a data‑driven approach to reliability, optimization, and continuous improvement.
  • Excellent communicator able to work effectively with both technical and non‑technical stakeholders.
  • Hands‑on experience with observability stacks (Prometheus, Grafana, PromQL), CI/CD pipelines, and operating large fleets of GPU servers.
  • Track record of leading incident response, postmortems, and driving robust service reliability.
Nice To Have Skills
  • Working knowledge of Kafka, ClickHouse, and CRDB.
  • DMTF, RedFish APIs, and GPU servers.
What We Offer

The base salary range for this role is $188,000 to $275,000. The starting salary will be determined based on job‑related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility).

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US‑based offerings; for roles in other locations, benefits vary and are shared during the hiring process.

  • Medical, dental, and vision insurance—100% paid for by CoreWeave.
  • Company‑paid Life Insurance.
  • Voluntary supplemental life insurance.
  • Short and long‑term disability insurance.
  • Flexible Spending Account.
  • Health Savings Account.
  • Tuition Reimbursement.
  • Ability to Participate in Employee Stock Purchase Program (ESPP).
  • Mental Wellness Benefits through Spring Health.
  • Family‑Forming support provided by Carrot.
  • Paid Parental Leave.
  • Flexible, full‑service childcare support with Kinside.
  • 401(k) with a generous employer match.
  • Flexible PTO.
  • Catered lunch each day in our office and data center locations.
  • A casual work environment.
  • A work culture focused on innovative disruption.
Equal Opportunity & Accommodations

CoreWeave is an equal opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information.

As part of this commitment and consistent with the Americans with Disabilities Act (ADA), CoreWeave will ensure that qualified applicants and candidates with disabilities are provided reasonable accommodations for the hiring process, unless such accommodation would cause an undue hardship. If reasonable accommodation is needed, please contact: careers@coreweave.com.

Export Control Compliance

This position requires access to export controlled information. To conform to U.S. Government export regulations applicable to that information, applicant must either be (A) a U.S. person, defined as a (i) U.S. citizen or national, (ii) U.S. lawful permanent resident (green card holder), (iii) refugee under 8 U.S.C. § 1157, or (iv) asylee under 8 U.S.C. § 1158, (B) eligible to access the export controlled information without a required export authorization, or (C) eligible and reasonably likely to obtain the required export authorization from the applicable U.S. government agency. CoreWeave may, for legitimate business reasons, decline to pursue any export licensing process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Software Engineer, Compute Architecture
Senior Software Engineer, Compute Architecture

CoreWeave • Sunnyvale (CA)

Hybrid
USD 180,000 - 260,000
Medical, dental, vision insurance (100
ESPP
401(k) with match
Staff Software Engineer, Compute Architecture
Staff Software Engineer, Compute Architecture

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance – 100% paid
401(k) with a generous employer match
Flexible PTO
+2
Senior Go Engineer: GPU Data Center Orchestration
Senior Go Engineer: GPU Data Center Orchestration

CoreWeave • Sunnyvale (CA)

On-site
USD 180,000 - 260,000
Software Engineer, Kubernetes
Software Engineer, Kubernetes

Coreweave • Livingston (NJ)

On-site
USD 153,000 - 204,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2
Senior Software Engineer, Server Fleet Infrastructure
Senior Software Engineer, Server Fleet Infrastructure

Coreweave • Livingston (NJ)

On-site
USD 110,000 - 150,000
Medical, dental, and vision insurance, fully paid
401(k) with employer match
Flexible Spending Account
Senior Software Engineer, Infrastructure Engineering
Senior Software Engineer, Infrastructure Engineering

CoreWeave • Livingston (NJ)

On-site
USD 153,000 - 242,000
Medical insurance
401(k) with employer match
Paid Parental Leave
+3
Engineering Manager, Infrastructure Engineering
Engineering Manager, Infrastructure Engineering

Socket.dev • New York (NY), Sunnyvale (CA), Livingston (NJ)

On-site
USD 182,000 - 242,000
Medical insurance
Dental insurance
Vision insurance
+8
Staff Software Engineer, Network Development
Staff Software Engineer, Network Development

CoreWeave • Livingston (NJ)

On-site
USD 207,000 - 275,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+1
Senior Platform Engineer, Metal Dev
Senior Platform Engineer, Metal Dev

CoreWeave • New York (NY)

Hybrid
USD 153,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Flexible Spending Account
+3
Engineering Manager, Infrastructure Engineering
Engineering Manager, Infrastructure Engineering

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Paid parental leave
+2