Senior AI Infra Performance & Observability Engineer

Coreweave

United States

On-site

USD 182,000 - 242,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
Tuition Reimbursement
ESPP
Parental Leave
Lunch provided in office

Job summary

CoreWeave is seeking a Senior Engineer for the Benchmarking & Performance team to build observability and insight tooling across GPU fleets and fabrics. You will transform billions of telemetry events into real-time performance insights for engineers, product teams, and executives.

Lead the design of the data lake and time-series stack, ensuring sub-second responses during live incidents. Collaborate across hardware and workloads to drive performance intelligence at scale.

Qualifications

  • 5+ years of experience building distributed systems, observability platforms, or performance engineering tooling.
  • Strong coding in Python or Go; familiarity with GPU infrastructure is a plus.
  • Hands‑on experience with Kubernetes at production scale, CI/CD, and observability stacks.
  • Fluency in PromQL/MetricsQL for real-time alerting and anomaly detection.
  • Knowledge of time-series databases and modern table formats (Iceberg, Parquet, Avro).

Responsibilities

  • Design and build systems that assess AI infrastructure health and performance.
  • Own time-series data layer; optimize PromQL/MetricsQL queries for alerting and analysis.
  • Develop telemetry pipelines for GPU fabric and distributed training workloads.
  • Design data lake components (Iceberg, Parquet, Avro) to support insights.
  • Profile and tune query engines for low latency and SLAs.
  • Build and, if needed, support BI/Reporting views (Grafana/Looker).

Skills

Distributed systems
Observability tooling
Python/Go
Kubernetes
PromQL
Time-series databases

Education

Bachelor’s degree in CS/CE or related field

Tools

Prometheus
Grafana
OpenTelemetry
Iceberg
Parquet/Avro

Job description

CoreWeave is seeking a Senior Engineer for the Benchmarking & Performance team to build observability and insight tooling across GPU fleets and fabrics. You will transform billions of telemetry events into real-time performance insights for engineers, product teams, and executives.

Lead the design of the data lake and time-series stack, ensuring sub-second responses during live incidents. Collaborate across hardware and workloads to drive performance intelligence at scale.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Benchmarking & Performance Architect
AI Benchmarking & Performance Architect

CoreWeave • Sunnyvale (CA)

On-site
USD 206,000 - 333,000
Medical, dental, and vision insurance
Equity awards
401(k) with match
+2
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

TheDataJob • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical, dental, vision insurance
ESPP – Employee Stock Purchase Program
Tuition Reimbursement
+3
Applied AI Inference Engineer: Benchmark & Optimize
Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • United States

On-site
USD 182,000 - 242,000
Medical Insurance
401(k) Matching
Paid Parental Leave
+3
Senior GPU Kernel Engineer - High-Performance Inference
Senior GPU Kernel Engineer - High-Performance Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 182,000 - 242,000
Healthcare coverage
Equity awards
401(k) match
+5
Inference Performance Engineer — Applied AI
Inference Performance Engineer — Applied AI

CoreWeave • Seattle (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
Discretionary bonus
+6
Senior Network Observability Engineer for GPU Cloud
Senior Network Observability Engineer for GPU Cloud

Socket.dev • Sunnyvale (CA), New York (NY)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+1
Senior Performance Engineer - AI Fabric & GPU Scale
Senior Performance Engineer - AI Fabric & GPU Scale

Astera Labs • San Jose (CA)

On-site
USD 135,000 - 170,000
Senior CUDA Kernel Engineer for High-Performance Inference
Senior CUDA Kernel Engineer for High-Performance Inference

CoreWeave • Bellevue (WA)

On-site
USD 182,000 - 242,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Disability insurance
+3
Senior Data Engineer: Data Lake & Observability
Senior Data Engineer: Data Lake & Observability

CoreWeave • Livingston (NJ)

On-site
USD 153,000 - 204,000
Medical, dental, and vision insurance
Company-paid Life Insurance
ESPP
+2