Senior ML Systems Engineer - Model Inference & Efficiency

Cohere

New York (NY)

Hybrid

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
100% Parental Leave top-up for up to 6 months
Personal enrichment benefits
Remote-flexible work options
6 weeks of vacation

Job summary

A leading AI research company is seeking an engineer for the Model Efficiency team to enhance ML systems. The ideal candidate has over 5 years of experience in high-performance coding, proficient in C++ or Python, and familiar with large language models. This role includes optimizing model execution and resolving performance issues. Offering a flexible remote working environment and generous perks, including health benefits and vacation time.

Qualifications

  • 5+ years of experience writing high-performance, production-quality code.
  • Strong programming skills in C++ or Python.
  • Experience with large language models and familiarity with LLM inference ecosystem.

Responsibilities

  • Work across the inference stack to improve core performance metrics.
  • Identify bottlenecks and develop innovative optimizations.
  • Experiment, measure, and ship improvements to accelerate inference.

Skills

High-performance production-quality code
C++
Python
Large language models
Diagnose performance bottlenecks

Tools

CUDA
GPU programming

Job description

A leading AI research company is seeking an engineer for the Model Efficiency team to enhance ML systems. The ideal candidate has over 5 years of experience in high-performance coding, proficient in C++ or Python, and familiar with large language models. This role includes optimizing model execution and resolving performance issues. Offering a flexible remote working environment and generous perks, including health benefits and vacation time.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 150,000 - 200,000
Senior ML Engineer - Low-Latency Inference & Systems
Senior ML Engineer - Low-Latency Inference & Systems

Inworld • Germany (OH)

Hybrid
USD 120,000 - 160,000
Staff ML Inference Engineer — Model Efficiency (Remote)
Staff ML Inference Engineer — Model Efficiency (Remote)

Jaide Health • San Francisco (CA)

On-site
USD 120,000 - 160,000
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
+2
Software Engineer - ML Model Performance
Software Engineer - ML Model Performance

Baseten • San Francisco (CA)

On-site
USD 150,000 - 250,000
ML Model Performance Engineer - Inference and Acceleration
ML Model Performance Engineer - Inference and Acceleration

Baseten • New York (NY)

On-site
USD 200,000 - 275,000
Senior ML Performance Engineer: LLM Benchmarking & GPU
Senior ML Performance Engineer: LLM Benchmarking & GPU

Amadeus Search • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Competitive salary
Equity and bonus opportunities
Medical, dental, and vision coverage
+2
Senior ML Engineer - Real-Time Inference & Systems
Senior ML Engineer - Real-Time Inference & Systems

Inworld AI • Germany (OH)

On-site
USD 120,000 - 180,000
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)
Staff ML Engineer: Build Ultra-Fast AI at Scale (Relocation)

Inworld AI • Mountain View (CA)

On-site
USD 270,000 - 500,000
Real-Time ML Inference Engineer for Scalable Serving
Real-Time ML Inference Engineer for Scalable Serving

Yobi • New York (NY)

Hybrid
USD 100,000 - 150,000
Competitive Base Salary
Meaningful equity
Annual performance bonus
+3
Senior ML Inference Systems Engineer
Senior ML Inference Systems Engineer

Gimlet Labs • San Francisco (CA)

On-site
USD 120,000 - 160,000