HPC Performance Engineer: Kernel & Systems

CoreWeave

New York (NY)

On-site

USD 165,000 - 242,000

Full time

48 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with employer match
Tuition Reimbursement
Flexible PTO
Catered lunch

Job summary

CoreWeave is seeking a highly skilled HPC Performance Engineer to join our HAVOCK Team, reporting to the Manager of Systems Engineering. You will design, develop, and optimize bare-metal systems from POST through joining a Kubernetes cluster, maintaining a custom Linux kernel, Ubuntu-based images, and the container/runtime stack.

You will collaborate with cross-functional teams and stakeholders to ensure low-latency, high-throughput performance, with telemetry and automation driving decisions

Qualifications

  • 5+ years of professional experience in Systems/HPC Performance Engineering, Benchmarking, and/or Validation.
  • Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, or a related field.
  • Strong experience with MPI workloads and distributed system performance analysis.
  • Familiarity with RoCE, InfiniBand, and GPUDirect/Data Direct I/O, NUMA, etc in HPC workloads.
  • Hands-on use of public HPC benchmarks (HPCC, HPL, OSU, MLPerf-HPC, STREAM, IO500).

Responsibilities

  • Develop and maintain tools for establishing systems performance baselines.
  • Develop and maintain performance regression analysis testing automation.
  • Design and maintain performance regression test pipelines for HPC workloads.
  • Debug and tune fabric-level performance to ensure low-latency high-throughput configurations.
  • Development of telemetry for performance analysis across distributed clusters of servers.
  • Triage and fix performance issues in Linux.
  • Define Linux and OS requirements, specifications, and system architecture in relation to systems performance, in collaboration with cross-functional teams.

Skills

Python
Go
bash/sh
C
Prometheus
Victoria Metrics
Grafana
Linux Kernel
Ubuntu
Kubernetes
Docker
KubeVirt
containerd
kubelet
NVIDIA GPUs
Infiniband

Education

Bachelor’s degree in Computer Engineering, Electrical Engineering, Computer Science, or a related field

Tools

MPI workloads
OS benchmarks (HPCC, HPL, OSU, MLPerf-HPC, STREAM, IO500)

Job description

CoreWeave is seeking a highly skilled HPC Performance Engineer to join our HAVOCK Team, reporting to the Manager of Systems Engineering. You will design, develop, and optimize bare-metal systems from POST through joining a Kubernetes cluster, maintaining a custom Linux kernel, Ubuntu-based images, and the container/runtime stack.

You will collaborate with cross-functional teams and stakeholders to ensure low-latency, high-throughput performance, with telemetry and automation driving decisions

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

HPC Performance Engineer for AI Cloud Systems
HPC Performance Engineer for AI Cloud Systems

CoreWeave • Bellevue (WA)

On-site
USD 165,000 - 242,000
Medical, dental, and vision insurance
401(k) with company match
Flexible PTO
+4
Senior Kubernetes HPC Field Engineer
Senior Kubernetes HPC Field Engineer

Neura Market • Sunnyvale (CA), Northern (KY)

Hybrid
USD 182,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
Senior Kubernetes HPC Field Engineer
Senior Kubernetes HPC Field Engineer

Socket.dev • San Francisco (CA)

On-site
USD 182,000 - 242,000
Medical, dental, vision insurance
401(k) with match
Paid parental leave
+1
Senior Linux Systems Engineer - GPU Virtualization
Senior Linux Systems Engineer - GPU Virtualization

Coreweave • Livingston (NJ)

On-site
USD 178,000 - 242,000
Medical, dental, vision insurance
Company-paid Life Insurance
401(k) with employer match
+2
Senior Systems Engineer — Kubernetes Test & HPC Infra
Senior Systems Engineer — Kubernetes Test & HPC Infra

CoreWeave • Bellevue (WA)

On-site
USD 140,000 - 190,000
Medical Insurance
401(k) with Employer Match
Flexible PTO
+1
HPC Performance Engineer
HPC Performance Engineer

CoreWeave • New York (NY)

On-site
USD 165,000 - 242,000
Medical, dental, and vision insurance
401(k) with employer match
Tuition Reimbursement
+2
HPC Performance Engineer
HPC Performance Engineer

CoreWeave • Bellevue (WA)

On-site
USD 165,000 - 242,000
Medical, dental, and vision insurance
401(k) with company match
Flexible PTO
+4
HPC Fleet Reliability Engineer
HPC Fleet Reliability Engineer

CoreWeave • Plano (TX)

On-site
USD 83,000 - 110,000
Medical, dental, and vision insurance
401(k) with generous match
Tuition Reimbursement
+4
Senior Cloud Support Engineer, Kubernetes & HPC
Senior Cloud Support Engineer, Kubernetes & HPC

CoreWeave • Bellevue (WA)

Hybrid
USD 122,000 - 163,000
Medical Insurance
Dental Insurance
Vision Insurance
+5
Senior GPU & Runtime Systems Engineer (Kubernetes & Linux)
Senior GPU & Runtime Systems Engineer (Kubernetes & Linux)

CoreWeave • Bellevue (WA)

On-site
USD 153,000 - 204,000
Medical, dental, and vision insurance
Equity awards
Discretionary bonus
+6