Senior ML Systems Engineer - Model Inference & Efficiency
Cohere
New York (NY)
Hybrid
USD 100,000 - 150,000
Full time
14 days+
Get more replies from employers
Send a job-specific resume in minutes.
Start fresh or import an existing resume
Benefits offered by this job
Inclusive culture and work environment
Weekly lunch stipend, in-office lunches & snacks
Full health and dental benefits
100% Parental Leave top-up for up to 6 months
Personal enrichment benefits
Remote-flexible work options
6 weeks of vacation
Job summary
A leading AI research company is seeking an engineer for the Model Efficiency team to enhance ML systems. The ideal candidate has over 5 years of experience in high-performance coding, proficient in C++ or Python, and familiar with large language models. This role includes optimizing model execution and resolving performance issues. Offering a flexible remote working environment and generous perks, including health benefits and vacation time.
Qualifications
5+ years of experience writing high-performance, production-quality code.
Strong programming skills in C++ or Python.
Experience with large language models and familiarity with LLM inference ecosystem.
Responsibilities
Work across the inference stack to improve core performance metrics.
Identify bottlenecks and develop innovative optimizations.
Experiment, measure, and ship improvements to accelerate inference.
Skills
High-performance production-quality code
C++
Python
Large language models
Diagnose performance bottlenecks
Tools
CUDA
GPU programming
Job description
A leading AI research company is seeking an engineer for the Model Efficiency team to enhance ML systems. The ideal candidate has over 5 years of experience in high-performance coding, proficient in C++ or Python, and familiar with large language models. This role includes optimizing model execution and resolving performance issues. Offering a flexible remote working environment and generous perks, including health benefits and vacation time.