Audio Inference Engineer, Model Efficiency

Visa Hunt

New York (NY)

Hybrid

USD 140,000 - 190,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
Parental Leave
Enrichment benefits
Education stipend
Six weeks vacation
Travel budget to other offices
Home office stipend

Job summary

Cohere in New York is seeking an experienced ML systems engineer to advance high‑performance audio and streaming inference on real-time workloads. You will optimize latency, throughput, and quality while collaborating with training and serving teams to ensure seamless model deployment.

Proficiency in C++, Python, and deep learning frameworks (PyTorch or TensorFlow) is essential, with hands‑on experience in audio, speech, or language models.

Qualifications

  • Significant experience developing high-performance audio or machine learning inference systems.
  • Proficiency with programming languages such as C++ and Python.
  • Hands-on experience with deep learning models for audio, speech, or language applications.
  • A bias for action and a strong results-oriented mindset.

Responsibilities

  • Build reliable audio inference systems and optimize latency, throughput, and quality.
  • Collaborate with both training and serving infrastructure teams to ensure seamless integration.
  • Dive deep into systems, identify bottlenecks, and deliver creative streaming solutions for audio workloads.
  • Contribute to real-time and streaming audio model serving improvements.

Skills

C++
Python
Audio ML
GPU programming
Real-time streaming
PyTorch
TensorFlow
SGLang
vLLM
Distributed inference

Tools

PyTorch
TensorFlow

Job description

Who are we?

Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems.

We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft.

We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us!

Why this role?

Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you will work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads.

You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real-time and streaming audio inference.

Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul and London. We embrace a remote-friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations.

You may be a good fit for the team if you have:
  • Significant experience developing high-performance audio or machine learning inference systems.

  • Proficiency with programming languages such as C++ and Python.

  • Hands-on experience with deep learning models for audio, speech, or language applications.

  • A bias for action and a strong results-oriented mindset.

It is a big plus if you also have considerable experience with:
  • GPU programming, low-level system optimization, model parallelization techniques over multiple GPUs

  • Have experience with duplex real-time streaming architectures.

  • Internals of machine learning frameworks for audio (such as PyTorch, TensorFlow, or specialized audio libraries).

  • Have experience with inference framework like vLLM, SGLang, Tensort-LLM, or custom distributed inference systems

  • Sequence modeling (e.g., transformers for audio/speech) and end-to-end audio pipeline optimization

Full-Time Employees at Cohere enjoy these Perks:
  • A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch.

  • Full health and dental benefits, including a separate budget for mental health.

  • RRSP matching, 401K, Pension Scheme.

  • 100% Parental Leave top-up for up to 6 months, for either parent.

  • Annual enrichment benefits:

    Arts & culture, fitness/wellness, quality time, and a workspace improvement credit.

    Education & learning stipend for conferences, courses, and coaching.

  • 6 weeks of paid vacation (30 working days!)

  • Budget for traveling to other offices if you are remote, plus an annual company offsite.

How and Where We Work:
  • Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon.

  • For those in the office: a daily lunch program, plenty of snacks, and regular community and social events.

  • For those not near an office: a co-working benefit so you can work alongside others in your city.

  • Everyone receives a $500 home office stipend to set up your workspace properly.

We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities.

We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Audio Inference Engineer, Model Efficiency
Audio Inference Engineer, Model Efficiency

Cohere • New York (NY)

Hybrid
USD 120,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Cohere • New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Lead Member of Technical Staff, Inference Infrastructure
Lead Member of Technical Staff, Inference Infrastructure

Visa Hunt • San Francisco (CA)

Hybrid
USD 210,000 - 320,000
Lunch stipend
Health benefits
RRSP matching
+4
Staff Research Engineer, Model Efficiency
Staff Research Engineer, Model Efficiency

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
Parental leave top‑up
+3
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency

Visa Hunt • New York (NY)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health and dental benefits
RRSP matching
+4
Site Reliability Engineer, Inference Infrastructure
Site Reliability Engineer, Inference Infrastructure

Cohere • New York (NY)

Hybrid
USD 140,000 - 200,000
Weekly lunch stipend
Health and dental benefits
RRSP matching
+3
Member of Technical Staff, Data Analysis and Evaluation
Member of Technical Staff, Data Analysis and Evaluation

Cohere • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 210,000
Lunch stipend
Health benefits
RRSP matching
+6
Staff Software Engineer, Inference Infrastructure
Staff Software Engineer, Inference Infrastructure

Jaide Health • San Francisco (CA)

On-site
USD 130,000 - 170,000
Open and inclusive culture
Weekly lunch stipend and snacks
Full health and dental benefits
+3
Software Engineer, Data Infrastructure
Software Engineer, Data Infrastructure

Cohere • San Francisco (CA), New York (NY)

Hybrid
USD 160,000 - 230,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+4