Remote Audio Inference Engineer — High-Performance Serving

Cohere

United States

Remote

USD 110,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

A weekly lunch stipend of $75/£75 or"

Job summary

Cohere is hiring an engineer to advance core audio model serving, latency, throughput, and quality for real-time streaming workloads. You will work across the model development and deployment stack with emphasis on C++/Python, GPU optimization, and ML frameworks in a remote-friendly, globally distributed setup.

You’ll collaborate with research and engineering to push frontier audio capabilities, applying attention to low-level optimizations and scalable inference systems across multiple GPUs.

Qualifications

  • Significant experience developing high-performance audio or machine-learning inference systems.
  • Proficiency with programming languages such as C++ and Python.
  • Hands-on experience with deep learning models for audio, speech, or language applications.

Responsibilities

  • Develop and optimize high-performance audio and ML inference systems for real-time streaming workloads.
  • Collaborate with training and serving teams to integrate model development with deployment.
  • Improve latency, throughput, and quality of audio processing pipelines.

Skills

Audio ML inference
C++
Python
Deep learning for audio
GPU programming
Streaming architectures
PyTorch/TensorFlow

Tools

vLLM
SGLang
Tensort-LLM

Job description

Cohere is hiring an engineer to advance core audio model serving, latency, throughput, and quality for real-time streaming workloads. You will work across the model development and deployment stack with emphasis on C++/Python, GPU optimization, and ML frameworks in a remote-friendly, globally distributed setup.

You’ll collaborate with research and engineering to push frontier audio capabilities, applying attention to low-level optimizations and scalable inference systems across multiple GPUs.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Audio Inference Engineer — Fast ML Serving
Remote Audio Inference Engineer — Fast ML Serving

Visa Hunt • New York (NY)

Hybrid
USD 140,000 - 190,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+6
Audio Inference Engineer — Model Efficiency (Remote-friendly)
Audio Inference Engineer — Model Efficiency (Remote-friendly)

Cohere • New York (NY)

Hybrid
USD 120,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Audio Inference Engineer, Model Efficiency
Audio Inference Engineer, Model Efficiency

Visa Hunt • New York (NY)

Hybrid
USD 140,000 - 190,000
Weekly lunch stipend
Health and dental benefits
RRSP matching / 401K / Pension
+6
Audio Inference Engineer, Model Efficiency
Audio Inference Engineer, Model Efficiency

Cohere • New York (NY)

Hybrid
USD 120,000 - 150,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+3
Remote Audio Inference Engineer, Model Efficiency
Remote Audio Inference Engineer, Model Efficiency

Jaide Health • San Francisco (CA)

On-site
USD 100,000 - 140,000
Open and inclusive culture
Weekly lunch stipend
Full health and dental benefits
+2
Staff Software Engineer — Scalable Inference Infra
Staff Software Engineer — Scalable Inference Infra

Cohere • New York (NY)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health and dental benefits
RRSP matching / 401K
+5
Staff Software Engineer, AI Inference Platform
Staff Software Engineer, AI Inference Platform

Visa Hunt • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Lunch stipend
Health & dental benefits
Parental leave top-up
+3
Staff Engineer – High-Performance Model Inference
Staff Engineer – High-Performance Model Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 180,000 - 440,000
Equity
Medical coverage
Vision coverage
+5
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal

OpenAI • San Francisco (CA)

On-site
USD 310,000 - 460,000
Inference Infrastructure Engineer, Serving
Inference Infrastructure Engineer, Serving

Jobtailor • Palo Alto (CA)

On-site
USD 180,000 - 240,000