Staff AI Inference Systems Engineer

CoreWeave

Sunnyvale (CA)

On-site

USD 188,000 - 275,000

Full time

3 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Company-paid Life Insurance
Voluntary supplemental life insurance
Health Savings Account
Tuition Reimbursement
Employee Stock Purchase Program (ESPP)

Job summary

CoreWeave is seeking a Staff Software Engineer (IC5) on the Inference team in Sunnyvale, CA to lead architecture and performance across services. You will optimize latency, throughput, and GPU utilization while guiding cross-team design and reliability at scale.

You’ll work with distributed systems, Kubernetes infrastructure, and inference workloads, driving memory optimization and batching strategies. A strong background in Go, Python, or C++ is expected.

Qualifications

  • 8–12+ years of experience building and operating large-scale distributed systems or cloud platforms
  • Proven experience leading cross-team technical initiatives impacting multiple services or organizations
  • Strong programming skills in Go, Python, or C++
  • Deep expertise in Kubernetes at production scale, including orchestration, scheduling, and service design
  • Strong understanding of distributed systems, networking, and performance optimization
  • Experience designing and operating low-latency, high-throughput systems with strict P95/P99 latency requirements
  • Hands-on experience with inference systems, including batching or micro-batching strategies, caching, and memory optimization
  • Experience improving system performance using metrics-driven approaches (e.g., latency, throughput, utilization)
  • Familiarity with mixed precision (BF16, FP8) and streaming inference workloads

Responsibilities

  • Lead architecture, performance, and reliability across multiple services and teams within the Inference platform
  • Optimize inference performance (latency, throughput, and GPU utilization) and memory usage
  • Drive cross-team design initiatives and influence engineering direction
  • Work deeply in distributed systems and Kubernetes-based infrastructure, focusing on scheduling, batching, and memory optimization

Skills

Go
Python
C++
Kubernetes
Distributed systems
Performance optimization
Networking
Inference systems
Latency and throughput

Tools

CUDA
NCCL
RDMA
NUMA
GPU interconnects
vLLM
Triton
TensorRT-LLM
Ray Serve
TorchServe

Job description

CoreWeave is seeking a Staff Software Engineer (IC5) on the Inference team in Sunnyvale, CA to lead architecture and performance across services. You will optimize latency, throughput, and GPU utilization while guiding cross-team design and reliability at scale.

You’ll work with distributed systems, Kubernetes infrastructure, and inference workloads, driving memory optimization and batching strategies. A strong background in Go, Python, or C++ is expected.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer, Inference
Staff Software Engineer, Inference

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Voluntary supplemental life insurance
+3
Senior Engineering Manager, AI Inference Platform
Senior Engineering Manager, AI Inference Platform

CoreWeave • Bellevue (WA)

On-site
USD 162,000 - 198,000
Medical, dental, and vision insurance
401(k) with employer match
Flexible PTO
+2
Senior Engineering Manager – AI Inference Platform
Senior Engineering Manager – AI Inference Platform

Neura Market • Bellevue (WA), Northern (KY)

Hybrid
USD 188,000 - 303,000
Medical, dental, vision
Equity awards
ESPP
+2
Inference AI/ML Engineer - GPU Cloud Platform
Inference AI/ML Engineer - GPU Cloud Platform

CoreWeave • Bellevue (WA)

On-site
USD 92,000 - 135,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous employer match
+2
Applied AI Inference Performance Engineer
Applied AI Inference Performance Engineer

TheDataJob • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical, dental, vision insurance
ESPP – Employee Stock Purchase Program
Tuition Reimbursement
+3
Applied AI Inference Engineer: Benchmark & Optimize
Applied AI Inference Engineer: Benchmark & Optimize

CoreWeave • San Francisco (CA)

On-site
USD 188,000 - 275,000
Medical insurance
Life Insurance
Disability insurance
+5
Inference Performance Engineer — Applied AI
Inference Performance Engineer — Applied AI

CoreWeave • Seattle (WA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Equity awards
Discretionary bonus
+6
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
Staff Software Engineer, Inference Cloud
Staff Software Engineer, Inference Cloud

Cerebras Systems • Sunnyvale (CA)

On-site
USD 130,000 - 160,000
Opportunity to publish and open source AI research
Work on one of the fastest AI supercomputers
Non-corporate work culture
Staff Software Engineer - Real-Time AI Inference Infra
Staff Software Engineer - Real-Time AI Inference Infra

Cerebras Systems, Inc. • Sunnyvale (CA)

On-site
USD 110,000 - 140,000