Audio Inference Engineer, Model Efficiency

Cohere

Deutschland

Vor Ort

EUR 60.000 - 85.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Cohere is looking for engineers to advance core audio model serving metrics. You will work on high-performance audio systems and enhance real-time inference capabilities.

The position requires significant experience with machine learning systems, C++, Python, and deep learning models for audio and speech. Flexibility in remote work options is encouraged, with teams distributed across various time zones for collaborative efficiency.

Qualifikationen

  • Significant experience developing high-performance audio or machine learning inference systems.
  • Proficiency with programming languages such as C++ and Python.
  • Hands-on experience with deep learning models for audio, speech, or language applications.

Aufgaben

  • Advance core audio model serving metrics, including latency, throughput, and quality.
  • Collaborate closely with the training and serving infrastructure teams.
  • Identify bottlenecks and deliver creative solutions for audio processing.

Kenntnisse

High-performance audio systems
C++
Python
Deep learning models
Real-time streaming

Tools

PyTorch
TensorFlow

Jobbeschreibung

Who are we?

Our mission is to scale intelligence to serve humanity. We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. We believe that our work is instrumental to the widespread adoption of AI.

We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. We like to work hard and move fast to do what’s best for our customers.

Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft. Each person is one of the best in the world at what they do. We believe that a diverse range of perspectives is a requirement for building great products.

Join us on our mission and shape the future!

Why this role?

Our team is a fast-growing group of committed researchers and engineers. The mission of the team is to build reliable machine learning systems and optimize audio inference serving efficiency using innovative techniques. As an engineer on this team, you’ll work on advancing core audio model serving metrics, including latency, throughput, and quality by diving deep into our systems, identifying bottlenecks, and delivering creative solutions for audio processing and streaming workloads.

You’ll collaborate closely with both the training and serving infrastructure teams to ensure seamless integration between model development and deployment, with a special focus on real‑time and streaming audio inference.

Please Note: We have offices in Toronto, Montreal, San Francisco, New York, Paris, Seoul, and London. We embrace a remote‑friendly environment, and as part of this approach, we strategically distribute teams based on interests, expertise, and time zones to promote collaboration and flexibility. You'll find the Model Efficiency team concentrated in the EST and PST time zones, these are our preferred locations.

YOU MAY BE A GOOD FIT FOR THE TEAM IF YOU HAVE:
  • Significant experience developing high‑performance audio or machine learning inference systems.
  • Proficiency with programming languages such as C++ and Python.
  • Hands‑on experience with deep learning models for audio, speech, or language applications.
  • A bias for action and a strong results‑oriented mindset.
IT IS A BIG PLUS IF YOU ALSO HAVE CONSIDERABLE EXPERIENCE WITH:
  • GPU programming, low‑level system optimization, model parallelization techniques over multiple GPUs.
  • Have experience with duplex real‑time streaming architectures.
  • Internals of machine learning frameworks for audio (such as PyTorch, TensorFlow, or specialized audio libraries).
  • Have experience with inference framework like vLLM, SGLang, Tensort‑LLM, or custom distributed inference systems.
  • Sequence modeling (e.g., transformers for audio/speech) and end‑to‑end audio pipeline optimization.
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Systems Software Engineer (Rust, ML Inference)
Systems Software Engineer (Rust, ML Inference)

ai-coustics • Berlin

Vor Ort
Vertraulich
Machine Learning Engineer
Machine Learning Engineer

ai|coustics • Berlin

Vor Ort
EUR 65.000 - 85.000
Competitive Compensation
Stock Options
Learning Opportunities
+4
Systems Software Engineer (Rust, ML Inference)
Systems Software Engineer (Rust, ML Inference)

ai-coustics • Berlin

Vor Ort
EUR 60.000 - 80.000
Competitive salary package
Additional benefits and stock options
Dynamic startup culture
Staff Research Engineer – Multimodal Generative Modelling
Staff Research Engineer – Multimodal Generative Modelling

Synthesia • Deutschland

Hybrid
EUR 120.000 - 190.000
Deep Learning Engineer
Deep Learning Engineer

Audatic • Berlin

Hybrid
EUR 65.000 - 85.000
30 days of paid vacation
Free daily meals, drinks, and snacks
Regular team events
+1
Inference Optimization — Member of Technical Staff
Inference Optimization — Member of Technical Staff

Construct Labs • Berlin

Vor Ort
EUR 120.000 - 180.000
Member of Technical Staff, Post-Training
Member of Technical Staff, Post-Training

Cohere • Deutschland

Hybrid
EUR 70.000 - 90.000
Spatial Audio Engineer
Spatial Audio Engineer

Audatic • Berlin

Hybrid
EUR 90.000 - 120.000
30 days paid vacation
Conference visits
Free daily meals, drinks and snacks
Senior Deep Learning Engineer
Senior Deep Learning Engineer

Audatic • Berlin

Hybrid
EUR 90.000 - 120.000
30 days of paid vacation
Free daily meals, drinks, and snacks
Conference visits
+2
Software Engineer, Data Infrastructure & Acquisition - Berlin, Germany
Software Engineer, Data Infrastructure & Acquisition - Berlin, Germany

Clutch Canada • Berlin

Remote
EUR 85.000 - 110.000
Competitive salaries
Fast-growing environment
Hands-off management approach
+2