Remote Audio Inference Engineer — High-Performance Serving

Cohere

United States

Remote

USD 110,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

A weekly lunch stipend of $75/£75 or"

Job summary

Cohere is hiring an engineer to advance core audio model serving, latency, throughput, and quality for real-time streaming workloads. You will work across the model development and deployment stack with emphasis on C++/Python, GPU optimization, and ML frameworks in a remote-friendly, globally distributed setup.

You’ll collaborate with research and engineering to push frontier audio capabilities, applying attention to low-level optimizations and scalable inference systems across multiple GPUs.

Qualifications

  • Significant experience developing high-performance audio or machine-learning inference systems.
  • Proficiency with programming languages such as C++ and Python.
  • Hands-on experience with deep learning models for audio, speech, or language applications.

Responsibilities

  • Develop and optimize high-performance audio and ML inference systems for real-time streaming workloads.
  • Collaborate with training and serving teams to integrate model development with deployment.
  • Improve latency, throughput, and quality of audio processing pipelines.

Skills

Audio ML inference
C++
Python
Deep learning for audio
GPU programming
Streaming architectures
PyTorch/TensorFlow

Tools

vLLM
SGLang
Tensort-LLM

Job description

Cohere is hiring an engineer to advance core audio model serving, latency, throughput, and quality for real-time streaming workloads. You will work across the model development and deployment stack with emphasis on C++/Python, GPU optimization, and ML frameworks in a remote-friendly, globally distributed setup.

You’ll collaborate with research and engineering to push frontier audio capabilities, applying attention to low-level optimizations and scalable inference systems across multiple GPUs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Audio Inference Engineer — Model Efficiency (Remote-friendly)
Audio Inference Engineer — Model Efficiency (Remote-friendly)

Cohere • New York (NY)

Hybrid
USD 120,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Audio Inference Engineer, Model Efficiency
Audio Inference Engineer, Model Efficiency

Cohere • New York (NY)

On-site
USD 120,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Remote Audio Inference Engineer, Model Efficiency
Remote Audio Inference Engineer, Model Efficiency

Jaide Health • San Francisco (CA)

On-site
USD 100,000 - 140,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+2
Audio Modeling Engineer — Real-Time Production Inference
Audio Modeling Engineer — Real-Time Production Inference

attention labs • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Multimodal Inference Engineer
Multimodal Inference Engineer

OpenAI • United States

Remote
USD 325,000 - 490,000
Medical insurance
Mental health support
401(k) plan with 50% matching
+3
Staff ML Engineer — Real-Time AI for Immersive Audio/Video
Staff ML Engineer — Real-Time AI for Immersive Audio/Video

Dolby Laboratories • Atlanta (GA)

On-site
USD 187,000 - 215,000
Bonus
Benefits
Profit sharing
+2
Embedded Systems Engineer, On-Device Inference
Embedded Systems Engineer, On-Device Inference

attention labs • Memphis (TN), Northern (KY)

On-site
USD 110,000 - 150,000
Inference Engineer, AGI
Inference Engineer, AGI

Amazon • Sunnyvale (CA)

On-site
USD 165,000 - 224,000
Health insurance
401(k) matching
Paid time off
+1
Senior Staff Engineer, Model Efficiency - Remote
Senior Staff Engineer, Model Efficiency - Remote

Cohere • United States

Remote
USD 150,000 - 210,000
Inference Systems Engineer: Optimize AI Serving & Latency
Inference Systems Engineer: Optimize AI Serving & Latency

adaption • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Flexible work in Bay Area
Adaption Passport travel stipend
Lunch stipend
+1