Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt

New York (NY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Lunch stipend
Health and dental benefits
RRSP matching
Parental leave top-up
Education stipend
Vacation and offsite
Home office stipend

Job summary

Cohere is hiring for a senior engineer focused on building reliable ML systems and optimizing LLM inference performance. You will work across the inference stack to reduce latency and increase throughput while collaborating with modeling and systems teams.

We offer remote-friendly roles with distributed teams, offices in multiple cities, and extensive perks including health benefits, parental leave, and learning stipends.

Qualifications

  • 5+ years of experience writing high-performance, production-quality code.
  • Strong programming skills in C++ or Python (Rust/Go also welcome).
  • Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.).
  • Ability to diagnose and resolve performance bottlenecks across the model execution stack.
  • A strong bias for action — you ship fast, measure impact, and iterate.

Responsibilities

  • Improve core performance metrics by optimizing model execution.
  • Collaborate with modeling and systems teams to measure and ship improvements.
  • Build expertise in GPU/CUDA optimizations and MoE strategies.

Skills

High-performance coding
C++
Python
LLM inference
Performance optimization

Tools

C++
Python
vLLM
SGLang

Job description

Cohere is hiring for a senior engineer focused on building reliable ML systems and optimizing LLM inference performance. You will work across the inference stack to reduce latency and increase throughput while collaborating with modeling and systems teams.

We offer remote-friendly roles with distributed teams, offices in multiple cities, and extensive perks including health benefits, parental leave, and learning stipends.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Research Engineer, LLM Inference & Efficiency Remote
Staff Research Engineer, LLM Inference & Efficiency Remote

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
Parental leave top‑up
+3
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Staff Engineer, Production AI & Scalable LLMs
Staff Engineer, Production AI & Scalable LLMs

cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend
Health benefits
RRSP/401K
+4
Staff Engineer - AI Systems & Post-Training RL
Staff Engineer - AI Systems & Post-Training RL

Cohere • New York (NY)

On-site
USD 150,000 - 230,000
Lunch stipend and health benefits
Dental benefits
RRSP/401K/Pension
+6
Staff ML Inference Engineer — Model Efficiency (Remote)
Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2
Distributed LLM Inference & Optimization Engineer
Distributed LLM Inference & Optimization Engineer

Together AI • San Francisco (CA)

On-site
USD 160,000 - 230,000
Startup equity
Health insurance
Competitive benefits
Staff Engineer - LLM Inference & Serving at Scale
Staff Engineer - LLM Inference & Serving at Scale

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Cohere • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+5
Software Engineer, Inference Stack (LLM Infra)
Software Engineer, Inference Stack (LLM Infra)

Baseten • United States

Remote
USD 180,000 - 240,000
Senior LLM Inference Engineer: Performance & Optimization
Senior LLM Inference Engineer: Performance & Optimization

Confidential • United States

On-site
USD 180,000 - 240,000