Senior ML Inference Engineer — Model Efficiency

Cohere

Montreal

Remote

CAD 100,000 - 130,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Open and inclusive culture
Cutting-edge AI research collaboration
Weekly lunch stipend
Full health and dental benefits
Parental Leave top-up for 6 months
Personal enrichment benefits
Remote-flexible work options
6 weeks of vacation

Job summary

A leading AI technology company is seeking a Member of Technical Staff to enhance model efficiency. This role involves improving performance metrics, optimizing bottlenecks, and collaborating with various teams. The ideal candidate has 5+ years in high-performance coding, strong skills in C++ or Python, and familiarity with large language models. Competitive perks include a flexible work environment, health benefits, and generous vacation time.

Qualifications

  • 5+ years of experience writing high-performance, production-quality code.
  • Strong programming skills in C++ or Python (Rust/Go also welcome).
  • Experience with large language models and the LLM inference ecosystem.
  • Ability to diagnose and resolve performance bottlenecks.
  • A strong bias for action — you ship fast, measure impact, and iterate.

Responsibilities

  • Work across the inference stack to improve core performance metrics.
  • Identify bottlenecks and develop innovative optimizations.
  • Collaborate closely with modeling and systems teams.
  • Experiment, measure, and ship improvements to accelerate inference.

Skills

High-performance code
C++ programming
Python programming
Diagnosing performance bottlenecks
Proactive action and iteration

Job description

A leading AI technology company is seeking a Member of Technical Staff to enhance model efficiency. This role involves improving performance metrics, optimizing bottlenecks, and collaborating with various teams. The ideal candidate has 5+ years in high-performance coding, strong skills in C++ or Python, and familiarity with large language models. Competitive perks include a flexible work environment, health benefits, and generous vacation time.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff ML Systems Engineer - Model Efficiency & Inference
Staff ML Systems Engineer - Model Efficiency & Inference

Cohere • Toronto

Hybrid
CAD 140,000 - 200,000
Lunch stipend
Health benefits
Retirement plan matching (RRSP/401K)
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Cohere • Montreal

Remote
CAD 100,000 - 130,000
Open and inclusive culture
Cutting-edge AI research collaboration
Weekly lunch stipend
+5
Staff Research Engineer: Accelerate LLM Inference (Remote)
Staff Research Engineer: Accelerate LLM Inference (Remote)

Cohere • Montreal

Hybrid
CAD 120,000 - 160,000
Senior AI/ML Engineer & Technical Lead
Senior AI/ML Engineer & Technical Lead

BrainWave Professionals • Canada

On-site
CAD 120,000 - 160,000
Senior ML Systems Engineer: Frameworks & Tooling (Remote)
Senior ML Systems Engineer: Frameworks & Tooling (Remote)

Cohere • Montreal

On-site
CAD 100,000 - 140,000
Senior ML Engineer - Build Production-Grade AI Systems
Senior ML Engineer - Build Production-Grade AI Systems

TheAppLabb • Toronto

On-site
CAD 110,000 - 170,000
Competitive salary
Opportunities for career growth
Fitness challenge incentives
Staff Research Engineer, LLM Inference & Efficiency
Staff Research Engineer, LLM Inference & Efficiency

Cohere • Montreal (administrative region)

Hybrid
CAD 150,000 - 190,000
Weekly lunch stipend
Health & dental benefits
RRSP matching / Pension
+6
Staff Research Engineer, RL & Model Integration (Remote-friendly)
Staff Research Engineer, RL & Model Integration (Remote-friendly)

Cohere • Montreal

Hybrid
CAD 80,000 - 120,000
Enterprise ML Engineer: Build Scalable AI & Agents
Enterprise ML Engineer: Build Scalable AI & Agents

Boson AI • Toronto

On-site
CAD 150,000 - 400,000
Senior ML Inference & Compiler Architect
Senior ML Inference & Compiler Architect

NVIDIA Corporation • Toronto

Hybrid
CAD 135,000 - 185,000