Staff Research Engineer, Model Efficiency - Fast Inference

Cohere

United States

Remote

USD 180,000 - 240,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Lunch stipend
Health & dental benefits
RRSP matching / 401K
Parental leave
Enrichment benefits
Education stipend
Paid vacation
Travel budget

Job summary

Cohere is seeking a Staff Research Engineer to advance model efficiency for foundation models. You will prototype and deploy techniques to speed up production runs and improve inference efficiency in a remote-friendly, startup-like setting.

You’ll work with a distributed team across EST/PST time zones, contributing to research, engineering, and collaboration with product teams. This role emphasizes leadership, rigorous research, and impactful delivery.

Qualifications

  • PhD in Machine Learning or a related field.
  • Deep understanding of LLM architecture and efficient inference.
  • Experience with techniques that improve model efficiency.
  • Strong software engineering skills.
  • Ability to thrive in a fast-paced startup environment.
  • Publications at top conferences (ICLR, ACL, NeurIPS).
  • Mentor others.

Skills

Software engineering
Mentor others
Publications in ML conferences
LLM/ML understanding

Education

PhD in Machine Learning or related field

Job description

Cohere is seeking a Staff Research Engineer to advance model efficiency for foundation models. You will prototype and deploy techniques to speed up production runs and improve inference efficiency in a remote-friendly, startup-like setting.

You’ll work with a distributed team across EST/PST time zones, contributing to research, engineering, and collaboration with product teams. This role emphasizes leadership, rigorous research, and impactful delivery.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Staff Engineer, Model Efficiency - Remote
Senior Staff Engineer, Model Efficiency - Remote

Cohere • United States

Remote
USD 150,000 - 210,000
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Cohere • United States

Remote
USD 150,000 - 210,000
Staff Engineer - Foundation Model Serving & Inference
Staff Engineer - Foundation Model Serving & Inference

Menlo Ventures • San Francisco (CA)

On-site
USD 120,000 - 160,000
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Cohere • San Francisco (CA)

On-site
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+5
Staff Research Engineer (Model Efficiency)
Staff Research Engineer (Model Efficiency)

Cohere • New York (NY)

On-site
USD 120,000 - 150,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Remote Research Engineer: Real-Time AI Inference
Remote Research Engineer: Real-Time AI Inference

ElevenLabs • Maine

Hybrid
USD 140,000 - 190,000
Annual discretionary stipend
Annual company offsite
Co-working stipend
Staff AI Modeling Engineer
Staff AI Modeling Engineer

Cohere • New York (NY)

Remote
USD 180,000 - 240,000
A weekly lunch stipend for lunch
Full health and dental benefits
RRSP matching / 401K / Pension
+5
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Cohere • New York (NY)

On-site
USD 100,000 - 150,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Staff Inference Engineer — Production-Scale AI Platform
Staff Inference Engineer — Production-Scale AI Platform

Designworks Talent LLC • Bellevue (KY)

Hybrid
USD 180,000 - 240,000
Health insurance
401(k) plan with company match
Paid holidays