Staff Research Engineer (Model Efficiency)

Cohere

New York (NY)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Deepstreamtech is seeking a Staff Research Engineer to focus on improving large language model (LLM) efficiency. This role will involve developing and deploying innovative techniques aimed at boosting model performance during production. Candidates should have a PhD in Machine Learning and strong software engineering skills for a collaborative startup environment.

The ideal candidate will be passionate about mentorship and have a proven track record of publications in top-tier AI conferences.

Qualifications

  • Must have a PhD in Machine Learning or a related field.
  • Strong understanding of LLM architecture and optimization.
  • Experience with techniques to enhance model efficiency.
  • Proven software engineering skills.

Responsibilities

  • Develop, prototype, and deploy techniques to enhance LLM inference efficiency.
  • Work on model architecture and optimization for performance improvements.
  • Contribute to decoding and inference-time algorithm improvements.

Skills

Machine Learning expertise
Software engineering
Model efficiency techniques
Mentoring

Education

PhD in Machine Learning or related field

Job description

Requirements
  • Have a PhD in Machine Learning or a related field
  • Understand LLM architecture, and how to optimize LLM inference given resource constraints
  • Have significant experience with one or more techniques that enhance model efficiency
  • Strong software engineering skills
  • An appetite to work in a fast-paced high-ambiguity start-up environment
  • Publications at top-tier conferences and venues (ICLR, ACL, NeurIPS)
  • Passion to mentor others
  • If some of the above doesn’t line up perfectly with your experience, we still encourage you to apply!
What the job involves
  • Large Language Models (LLMs) continue to push the boundaries of what AI systems can do - but inference is still the bottleneck
  • The Model Efficiency team is responsible for pushing the limits of LLM inference efficiency across our foundation models. We explore and ship breakthroughs across the model execution stack, including:
  • Model architecture and MoE routing optimization
  • Decoding and inference-time algorithm improvements
  • Software/hardware co-design for GPU acceleration
  • Performance optimization without compromising model quality
  • As a Staff Research Engineer, you will develop, prototype, and deploy techniques that materially improve how fast and efficiently our models run in production
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff Research Engineer: LLM Efficiency Architect
Staff Research Engineer: LLM Efficiency Architect

Deepstreamtech • New York (NY)

On-site
USD 120,000 - 150,000
Staff Engineer - ML Inference & Model Efficiency
Staff Engineer - ML Inference & Model Efficiency

Cohere • San Francisco (CA)

Remote
USD 180,000 - 240,000
Inclusive work culture
Weekly lunch stipend
Full health and dental benefits
+4
Member of ML Technical Staff
Member of ML Technical Staff

Pragmatike • San Francisco (CA)

On-site
USD 200,000 - 350,000
Research Engineer
Research Engineer

Nace.AI • Palo Alto (CA)

On-site
USD 120,000 - 160,000
RESEARCHER, EFFICIENT INFERENCE
RESEARCHER, EFFICIENT INFERENCE

MLSys 2020 • San Francisco (CA)

On-site
USD 140,000 - 180,000
Machine Learning Researcher
Machine Learning Researcher

Multicoin • San Francisco (CA)

On-site
USD 250,000 - 350,000
Competitive compensation
Equity in a high-growth startup
Comprehensive benefits
Research Member of Technical Staff- Efficient Modeling
Research Member of Technical Staff- Efficient Modeling

Rhoda AI • Mountain View (CA)

On-site
USD 100,000 - 150,000
LLM Inference Engineer (Mid, Senior, Staff)
LLM Inference Engineer (Mid, Senior, Staff)

Hippocratic AI Inc. • Menlo Park (CA)

On-site
USD 180,000 - 280,000
Member of Technical Staff [Research]
Member of Technical Staff [Research]

NeoCognition Inc. • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Machine Learning Researcher
Machine Learning Researcher

SOLANA FOUNDATION • San Francisco (CA)

On-site
USD 250,000 - 350,000
Equity in a high-growth startup
Comprehensive benefits