Staff Engineer - ML Inference & Model Efficiency

Cohere

San Francisco (CA)

Remote

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
Mental health budget
Parental leave top-up
Personal enrichment benefits
6 weeks of vacation

Job summary

A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.

Qualifications

  • 5+ years of experience writing high-performance, production-quality code.
  • Strong programming skills in languages such as C++ or Python.
  • Ability to diagnose and resolve performance bottlenecks.

Responsibilities

  • Work across the inference stack to improve performance metrics.
  • Identify bottlenecks and develop optimizations.
  • Collaborate closely with modeling and systems teams.

Skills

High-performance coding
C++ programming
Python programming
LLM inference knowledge
Diagnosing performance bottlenecks
Fast shipping and iteration

Tools

CUDA programming
Distributed systems

Job description

A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Staff ML Inference Engineer — Model Efficiency (Remote)
Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Staff Engineer, Model Efficiency & LLM Inference
Staff Engineer, Model Efficiency & LLM Inference

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff Research Engineer, LLM Inference & Efficiency Remote
Staff Research Engineer, LLM Inference & Efficiency Remote

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
Parental leave top‑up
+3
ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Staff Engineer - LLM Inference & Serving at Scale
Staff Engineer - LLM Inference & Serving at Scale

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Senior Staff Tech Lead — Inference & ML Performance
Senior Staff Tech Lead — Inference & ML Performance

fal • San Francisco (CA)

On-site
USD 150,000 - 200,000
ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000