Member of Technical Staff, Model Efficiency

Cohere

Montreal

Remote

CAD 100,000 - 130,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Open and inclusive culture
Cutting-edge AI research collaboration
Weekly lunch stipend
Full health and dental benefits
Parental Leave top-up for 6 months
Personal enrichment benefits
Remote-flexible work options
6 weeks of vacation

Job summary

A leading AI technology company is seeking a Member of Technical Staff to enhance model efficiency. This role involves improving performance metrics, optimizing bottlenecks, and collaborating with various teams. The ideal candidate has 5+ years in high-performance coding, strong skills in C++ or Python, and familiarity with large language models. Competitive perks include a flexible work environment, health benefits, and generous vacation time.

Qualifications

  • 5+ years of experience writing high-performance, production-quality code.
  • Strong programming skills in C++ or Python (Rust/Go also welcome).
  • Experience with large language models and the LLM inference ecosystem.
  • Ability to diagnose and resolve performance bottlenecks.
  • A strong bias for action — you ship fast, measure impact, and iterate.

Responsibilities

  • Work across the inference stack to improve core performance metrics.
  • Identify bottlenecks and develop innovative optimizations.
  • Collaborate closely with modeling and systems teams.
  • Experiment, measure, and ship improvements to accelerate inference.

Skills

High-performance code
C++ programming
Python programming
Diagnosing performance bottlenecks
Proactive action and iteration

Job description

Member of Technical Staff, Model Efficiency

1 day ago Be among the first 25 applicants

Get AI-powered advice on this job and more exclusive features.

Who are we?

Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI.

Why this role?

Our team is a fast-growing group of researchers and engineers focused on building reliable ML systems and pushing the boundaries of LLM inference efficiency. We develop techniques that improve how models execute in production, driving lower latency, higher throughput, and consistent quality across diverse workloads.

As an engineer on this team, you’ll work across the inference stack to improve core performance metrics by diving deep into model execution, identifying bottlenecks, and developing innovative optimizations. You’ll collaborate closely with modeling and systems teams to experiment, measure, and ship improvements that meaningfully accelerate inference. As the team evolves, you’ll have opportunities to build expertise in advanced performance techniques, including GPU/CUDA optimizations, kernel-level improvements, and model execution strategies for MoE and large-scale architectures.

We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations.

You may be a good fit for the Model Efficiency team if you have:

  • 5+ years of experience writing high-performance, production-quality code
  • Strong programming skills in C++ or Python (Rust/Go also welcome)
  • Experience working with large language models and familiarity with the LLM inference ecosystem (e.g., vLLM, SGLang, etc.)
  • Ability to diagnose and resolve performance bottlenecks across the model execution stack
  • A strong bias for action — you ship fast, measure impact, and iterate

It’s a big plus if you have experience with:

  • GPU programming, CUDA, or low-level systems optimization
  • Language modeling with transformers (MoE, speculative decoding, KV-cache optimizations)
  • Scaling performance-critical distributed systems (e.g., computation, search, storage)

We value and celebrate diversity and strive to create an inclusive work environment for all. We welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs.

Full-Time Employees At Cohere Enjoy These Perks

  • 🤝 An open and inclusive culture and work environment
  • 🧑💻 Work closely with a team on the cutting edge of AI research
  • 🍽 Weekly lunch stipend, in-office lunches & snacks
  • 🦷 Full health and dental benefits, including a separate budget to take care of your mental health
  • 🐣 100% Parental Leave top-up for up to 6 months
  • 🎨 Personal enrichment benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement
  • 🏙 Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend
  • ✈️ 6 weeks of vacation (30 working days!)

Seniority level: Mid-Senior level

Employment type: Full-time

Job function: Engineering and Information Technology

Industries: Software Development

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Cohere • Toronto

Hybrid
CAD 140,000 - 200,000
Lunch stipend
Health benefits
Retirement plan matching (RRSP/401K)
Staff Research Engineer, Model Efficiency
Staff Research Engineer, Model Efficiency

Cohere • Montreal

On-site
CAD 120,000 - 160,000
Open and inclusive culture
Collaboration on cutting-edge AI research
Weekly lunch stipend, in-office lunches & snacks
+5
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

Cohere • Toronto

On-site
CAD 100,000 - 150,000
An open and inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+4
Senior ML Systems Engineer, Frameworks & Tooling
Senior ML Systems Engineer, Frameworks & Tooling

Cohere • Montreal

On-site
CAD 100,000 - 140,000
Open and inclusive culture
Weekly lunch stipend and snacks
Full health and dental benefits
+2
Member of Technical Staff (Sovereign AI)
Member of Technical Staff (Sovereign AI)

Cohere • Toronto

Hybrid
CAD 100,000 - 140,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Member of Technical Staff, Modeling
Member of Technical Staff, Modeling

Cohere • Toronto

Hybrid
CAD 150,000 - 190,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
Member of Technical Staff, Training Performance Engineer
Member of Technical Staff, Training Performance Engineer

Cohere • Toronto

Hybrid
CAD 150,000 - 190,000
Lunch stipend
Health and dental benefits
RRSP matching / Pension
+5
Senior Research Scientist, Model Evaluation
Senior Research Scientist, Model Evaluation

Visa Hunt • Toronto

Hybrid
CAD 140,000 - 190,000
Lunch stipend
Health & dental
RRSP matching
+5
Senior Member of Technical Staff, Safety and Security for Agents
Senior Member of Technical Staff, Safety and Security for Agents

Cohere • Toronto

On-site
CAD 100,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+4
Senior Member of Technical Staff, Safety and Security for Agents
Senior Member of Technical Staff, Safety and Security for Agents

Cohere • Montreal (administrative region)

Hybrid
CAD 120,000 - 180,000
Open and inclusive culture
Collaborative AI research environment
Weekly lunch stipend
+5