Staff Engineer - ML Inference & Model Efficiency

Cohere

San Francisco (CA)

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
Mental health budget
Parental leave top-up
Personal enrichment benefits
6 weeks of vacation

Job summary

A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.

Qualifications

  • 5+ years of experience writing high-performance, production-quality code.
  • Strong programming skills in languages such as C++ or Python.
  • Ability to diagnose and resolve performance bottlenecks.

Responsibilities

  • Work across the inference stack to improve performance metrics.
  • Identify bottlenecks and develop optimizations.
  • Collaborate closely with modeling and systems teams.

Skills

High-performance coding
C++ programming
Python programming
LLM inference knowledge
Diagnosing performance bottlenecks
Fast shipping and iteration

Tools

CUDA programming
Distributed systems

Job description

A leading AI research firm in San Francisco is seeking a Member of Technical Staff specialized in Model Efficiency. In this role, you will enhance LLM inference systems by tackling performance issues and collaborating with cross-functional teams. Ideal candidates have over 5 years of coding experience in C++ or Python and a solid understanding of the LLM inference environment. This position offers a remote-friendly work model, a competitive salary, and extensive benefits including a generous vacation policy.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Systems Engineer - Model Inference & Efficiency
Senior ML Systems Engineer - Model Inference & Efficiency

Cohere • New York (NY)

Hybrid
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Staff ML Inference Engineer — Model Efficiency (Remote)
Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
Competitive compensation with equity
100% medical, dental, and vision insurance
Generous PTO policy
+2
ML Inference Systems Engineer
ML Inference Systems Engineer

Gimlet Labs, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
100% coverage of medical, dental, and vision insurance
Generous PTO policy
Paid parental leave
+2
Staff Engineer, LLM Inference & Infra
Staff Engineer, LLM Inference & Infra

Prime Intellect • United States

Hybrid
USD 120,000 - 150,000
Competitive compensation
Flexible work arrangement
Full visa sponsorship
+2
Senior ML Inference Systems Engineer
Senior ML Inference Systems Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000
ML Performance Engineer – Real-Time Inference
ML Performance Engineer – Real-Time Inference

Odyssey • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Staff Engineer - Customer-Facing AI Inference Infra
Staff Engineer - Customer-Facing AI Inference Infra

Simplify • San Francisco (CA)

On-site
USD 200,000 - 300,000
Housing stipend
Uber/Waymo rides
Senior ML Serving Engineer for LLMs & Inference
Senior ML Serving Engineer for LLMs & Inference

Alldus • San Jose (CA)

On-site
USD 180,000 - 220,000