Inference AI/ML Engineer for GPU Model Serving

re-zoo-me

Sunnyvale (CA)

On-site

USD 92,000 - 135,000

Full time

9 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
Life Insurance
ESPP
Tuition Reimbursement
401(k) with employer match
Flexible PTO
Catered lunch
Mental wellness benefits

Job summary

CoreWeave is seeking an IC1 engineer to join the Inference team and ship production features for model serving on our GPU platform. You will implement well-scoped changes, learn our practices, and grow quickly with mentorship from experienced engineers.

You will work on Python/Go/C++ services such as Triton, vLLM, and TensorRT-LLM, write tests and docs, and contribute to metrics, dashboards, and runbooks. This is an opportunity to advance in a fast-paced AI infrastructure company in Sunnyvale,

Qualifications

  • BS/MS in CS, EE, or related field, or equivalent practical experience.
  • Foundations in data structures, algorithms, and networked services.
  • Experience with Python or Go; C++ is a plus; Linux fundamentals; Git/CI basics.
  • Exposure to containers and Kubernetes (coursework or projects welcome).
  • Curiosity about GPU inference concepts (micro-batching, KV cache, streaming).

Responsibilities

  • Implement features and fixes in Python/Go/C++ for model-serving services (e.g., Triton, vLLM, TensorRT-LLM, Ray Serve).
  • Write tests, code comments, and short design docs; participate in code reviews.
  • Add basic metrics and dashboards; assist with alarms and runbooks.
  • Follow on-call runbooks and learn incident response in a guided rotation.
  • Contribute to performance experiments (e.g., request batching, concurrency, caching) with guidance.

Skills

Python/Go
C++
Linux
Git/CI
Containers/Kubernetes
GPU inference concepts

Education

BS/MS in CS/EE or related

Tools

Triton
vLLM
TensorRT-LLM
Ray Serve
Grafana/Prometheus/OpenTelemetry

Job description

CoreWeave is seeking an IC1 engineer to join the Inference team and ship production features for model serving on our GPU platform. You will implement well-scoped changes, learn our practices, and grow quickly with mentorship from experienced engineers.

You will work on Python/Go/C++ services such as Triton, vLLM, and TensorRT-LLM, write tests and docs, and contribute to metrics, dashboards, and runbooks. This is an opportunity to advance in a fast-paced AI infrastructure company in Sunnyvale,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Inference AI/ML Engineer - GPU Cloud Platform
Inference AI/ML Engineer - GPU Cloud Platform

CoreWeave • Bellevue (WA)

On-site
USD 92,000 - 135,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous employer match
+2
Inference AI Engineer — Scalable GPU ML Services
Inference AI Engineer — Scalable GPU ML Services

CoreWeave • Sunnyvale (CA)

On-site
USD 92,000 - 135,000
Medical, dental, and vision insurance
401(k) match
Equity awards
+5
Staff AI Inference Systems Engineer
Staff AI Inference Systems Engineer

CoreWeave • Sunnyvale (CA)

On-site
USD 188,000 - 275,000
Medical, dental, and vision insurance
Company-paid Life Insurance
Voluntary supplemental life insurance
+3
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays
Software Engineer, Inference AI/ML
Software Engineer, Inference AI/ML

CoreWeave • Sunnyvale (CA)

On-site
USD 92,000 - 135,000
Medical, dental, and vision insurance
401(k) match
Equity awards
+5
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Software Engineer, Inference AI/ML
Software Engineer, Inference AI/ML

CoreWeave • Bellevue (WA)

On-site
USD 92,000 - 135,000
Medical, dental, and vision insurance
Company-paid Life Insurance
401(k) with generous employer match
+2
Machine Learning Engineer (Inference)
Machine Learning Engineer (Inference)

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 240,000
Graduate Backend Inference Engine Engineer
Graduate Backend Inference Engine Engineer

ByteDance • San Jose (CA)

On-site
USD 128,000 - 256,000
Medical insurance
Dental insurance
Vision insurance
+8