Operations Engineer, Fleet Reliability

CoreWeave

Washington (District of Columbia)

On-site

USD 83,000 - 110,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
HSA/Health Savings Account
Tuition Reimbursement
Employee Stock Purchase Program
Mental Wellness Benefit
Parental Leave and childcare support
401(k) with employer match
Flexible PTO

Job summary

CoreWeave is seeking a Fleet Reliability Operations engineer to manage provisioning, maintenance, and uptime of its expanding fleet of server nodes. You’ll work on configuration, updates, and remote troubleshooting of top-tier HPC clusters and their networking, delivery platforms and tool dependencies.

You will join a focused team to deploy nodes quickly, rack them, and power them on while collaborating with data center, hardware, and platform teams to resolve issues and improve processes.

Qualifications

  • Strong understanding of Linux system administration and internals.
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably.
  • Experience with scripting languages (bash, python, powershell, etc.)

Responsibilities

  • Configure and maintain large‑scale high‑performance supercomputing clusters running state‑of‑the‑art GPUs.
  • Troubleshoot hardware and software issues; escalation and coordination with data center, network, hardware and platform teams to drive resolution.
  • Monitor and analyze system performance and take appropriate remediation actions for cloud health.
  • Approach your work with flexibility and optimism anticipating shifting business and technical priorities.
  • Create and maintain documentation of team processes, knowledge and best practices for system management.
  • Think critically about your day‑to‑day work and collaborate to improve team processes and efficiency.
  • Participate in on‑call rotations, including after‑hours and weekend work.

Skills

Linux administration
Troubleshooting
Scripting (bash, Python, PowerShell)

Education

Bachelor’s degree in a related field or equivalent experience

Tools

Grafana
Prometheus
promsql
Kubernetes

Job description

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com.

What You'll Do

The Fleet Reliability Operations team is responsible for the day‑to‑day provisioning, management and uptime of CoreWeave’s ever‑expanding fleet of server nodes. Playing a central role in CoreWeave’s growth strategy, this team is on the front line for configuration, updates and remote troubleshooting of our highest tier of supercomputing clusters and their networking, delivery platforms and tool dependencies. You will be in a daily battle with entropy to maximize the number of nodes CoreWeave can deliver to customers.

We are seeking curious, creative and persistent problem solvers to join our Fleet Reliability Operations team to help us drive batches of server nodes through our provisioning and validation processes while efficiently and effectively troubleshooting node or cluster problems as they arise. This individual will join a team of committed engineers working to deploy nodes as fast as they can be racked and turned on.

  • Configure and maintain large‑scale high‑performance supercomputing clusters running state‑of‑the‑art GPUs
  • Troubleshoot hardware and software issues; escalation and coordination with data center, network, hardware and platform teams to drive resolution
  • Monitor and analyze system performance and take appropriate remediation actions for cloud health
  • Approach your work with flexibility and optimism anticipating shifting business and technical priorities
  • Create and maintain documentation of team processes, knowledge and best practices for system management
  • Think critically about your day‑to‑day work and collaborate to improve team processes and efficiency
  • Participate in on‑call rotations, including after‑hours and weekend work
Minimum Qualifications
  • Strong understanding of Linux system administration and internals
  • Ability to troubleshoot hardware and software issues and perform system maintenance tasks consistently and reliably
  • Experience with scripting languages (bash, python, powershell, etc.)
Preferred Qualifications
  • 2+ years of experience troubleshooting or administering data center or on‑prem infrastructure (servers, storage, network or a mix)
  • Experience with Grafana, Prometheus, promsql queries or similar observability platforms
  • Experience with data center environments, including server racks, HVAC systems, fiber trays
  • Experience administering Kubernetes
  • Experience with HPC and GPU‑related workloads
  • Bachelor’s degree in a related field or equivalent experience
Compensation & Benefits

The base salary range for this role is $83,000 to $110,000. The starting salary will be determined based on job‑related knowledge, skills, experience, and market location. In addition to base salary, the total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program.

Benefits include:

  • Medical, dental, and vision insurance – 100% paid by CoreWeave
  • Company‑paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long‑term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Employee Stock Purchase Program (ESPP) participation
  • Mental Wellness Benefits through Spring Health
  • Family‑forming support provided by Carrot
  • Paid Parental Leave
  • Full‑service childcare support with Kinside
  • 401(k) with generous employer match
  • Flexible PTO
Equal Opportunity & Accommodations

CoreWeave is an equal‑opportunity employer, committed to fostering an inclusive and supportive workplace. All qualified applicants and candidates will receive consideration for employment without regard to race, color, religion, sex, disability, age, sexual orientation, gender identity, national origin, veteran status, or genetic information. CoreWeave will provide reasonable accommodations for applicants with disabilities in compliance with the Americans with Disabilities Act.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Plano (TX)

Hybrid
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with a generous employer match
Flexible PTO
+2
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Bellevue (WA)

Hybrid
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with employer match
Paid Parental Leave
+2
Operations Engineer, Fleet Reliability
Operations Engineer, Fleet Reliability

CoreWeave • Sunnyvale (CA)

Hybrid
USD 83,000 - 110,000
100% paid medical, dental, and vision insurance
401(k) with generous employer match
Flexible PTO
+2
Operations Engineering Manager, Fleet Reliability
Operations Engineering Manager, Fleet Reliability

CoreWeave • Livingston (NJ)

On-site
USD 143,000 - 191,000
Medical, dental, and vision insurance
Equity awards
401(k) with employer match
+4
Senior Software Engineer, Server Fleet Infrastructure
Senior Software Engineer, Server Fleet Infrastructure

Coreweave • Livingston (NJ)

On-site
USD 110,000 - 150,000
Medical, dental, and vision insurance, fully paid
401(k) with employer match
Flexible Spending Account
Staff Software Engineer, Network Development
Staff Software Engineer, Network Development

CoreWeave • Bellevue (WA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+3
Staff Software Engineer, Network Development
Staff Software Engineer, Network Development

CoreWeave • Sunnyvale (CA)

On-site
USD 190,000 - 260,000
Medical, dental, and vision insurance—
Company‑paid Life Insurance
401(k)with employer match
+3
Sr GPU Infrastructure Software Engineer
Sr GPU Infrastructure Software Engineer

CoreWeave • Sunnyvale (CA)

Hybrid
USD 165,000 - 242,000
100% paid medical, dental, and vision insurance
Company-paid life insurance
401(k) with generous employer match
+2
Staff Software Engineer, Compute Architecture
Staff Software Engineer, Compute Architecture

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance – 100% paid
401(k) with a generous employer match
Flexible PTO
+2
Sr GPU Infrastructure Software Engineer
Sr GPU Infrastructure Software Engineer

CoreWeave • Bellevue (WA)

Hybrid
USD 165,000 - 242,000
Medical, dental, and vision insurance - 100% paid
Company-paid Life Insurance
Flexible Spending Account
+6