Research Engineer - Inference

Coinscapture

Northern (KY)

Remote

USD 120,000 - 180,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Remote work

Job summary

ElevenLabs is seeking a Research Engineer to join the research team, focusing on deploying and optimizing frontier AI models in production. You will own systems that turn research breakthroughs into real-time products used by millions, ensuring fast, reliable, scalable serving.

You will thrive by deploying state-of-the-art models, optimizing latency and throughput, and building tooling to help researchers ship new models safely and efficiently.

Qualifications

  • Experience deploying and serving ML models in production, latency-sensitive or real-time apps.
  • Strong engineering skills in GPU programming and inference optimization (CUDA, Triton, TensorRT, vLLM, SGLang).
  • Ability to profile, diagnose, and eliminate bottlenecks across the serving stack and build tooling to measure it.

Responsibilities

  • Deploy state-of-the-art models to production and own the path from research checkpoint to serving infrastructure.
  • Optimize inference performance across the stack, including latency, throughput, and cost.
  • Build high-performance serving systems for real-time, streaming workloads where every millisecond matters.
  • Create tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics.

Skills

Production ML deployment
GPU programming
Inference optimization
Profiling & bottleneck diagnosis
Research-to-production handoff

Education

No degree required

Tools

CUDA
TensorRT
Triton
vLLM
SGLang

Job description

# Research Engineer - InferenceElevenlabsRemoteUnited StatesSoftware EngineeringPosted Oct 8, 2026About ElevenLabs ElevenLabs is an AI research and product company transforming how we interact with technology. We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $22B - multiples of 11, always. We have expanded from voice into three main platforms: • ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale. • ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages. • ElevenAP I gives developers access to our leading AI audio foundational models. Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you. How we work • High-velocity Rapid experimentation, lean autonomous teams, and minimal bureaucracy. • Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you. • AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations. • Excellence everywhere: Everything we do should match the quality of our AI models. • Global team: We prioritize your talent, not your location. What we offer • Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible. • Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities. • Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend. • Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose. • Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy. • Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend. About the role We are looking for a Research Engineer to join the research team at ElevenLabs, focused on deploying and optimizing our frontier AI models in production. The quality of our models only matters if they can be served fast, reliably, and at scale. You will own the systems that turn research breakthroughs into real-time products used by millions. You will thrive in this role if you enjoy: • Deploying state-of-the-art models to production and owning the path from research checkpoint to serving infrastructure. • Optimizing inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels. • Building and tuning high-performance serving systems for real-time, streaming workloads where every millisecond matters. • Creating tooling and infrastructure that lets researchers ship new models to production quickly, safely, and with confidence in their performance characteristics. Requirements We do not require any formal certifications or degrees. Instead, we are seeking enthusiastic engineers who can showcase solving impressively hard problems with artifacts such as past projects, designs, or GitHub contributions. Ideally, you bring: • Experience deploying and serving ML models in production, ideally for latency-sensitive or real-time applications. • Strong engineering skills in GPU programming and inference optimization (e.g., CUDA, Triton, TensorRT, or serving frameworks such as vLLM or SGLang). • The capacity to autonomously profile, diagnose, and eliminate bottlenecks across the serving stack, from model architecture to kernels to orchestration, and to build the tooling to measure it. Location This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw. #LI-Remote We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Engineer - Inference
Research Engineer - Inference

Devconnectplatform • United States

Remote
USD 120,000 - 190,000
Annual discretionary stipend
Social travel stipend
Annual company offsite
+1
Research Engineer – Inference
Research Engineer – Inference

Jobicy • United States

Remote
USD 120,000 - 170,000
Innovative culture
Growth paths
Learning & development
+3
Research Engineer - Inference
Research Engineer - Inference

UMAI Consultancy • United States

Remote
USD 140,000 - 210,000
Research Engineer - Inference
Research Engineer - Inference

ElevenLabs • Maine

On-site
USD 140,000 - 190,000
Annual discretionary stipend
Annual company offsite
Co-working stipend
Research Engineer
Research Engineer

Coinscapture • Northern (KY)

Hybrid
USD 120,000 - 180,000
Annual discretionary stipend
Annual company offsite
Co-working stipend
+2
Research Engineer
Research Engineer

UMAI Consultancy • United States

Remote
USD 120,000 - 180,000
Innovative culture
Growth opportunities
Learning stipend
+3
Full-Stack Engineer
Full-Stack Engineer

Coinscapture • Northern (KY)

Hybrid
USD 120,000 - 170,000
Remote work stipend
Annual discretionary stipend
Annual company offsite
+2
Full-Stack Engineer (Back-End Leaning)
Full-Stack Engineer (Back-End Leaning)

Coinscapture • Northern (KY)

Hybrid
USD 120,000 - 180,000
Learning & development stipend
Annual social travel stipend
Annual company offsite
+1
Engineering - Internal AI Transformation
Engineering - Internal AI Transformation

Coinscapture • Northern (KY)

Remote
USD 150,000 - 190,000
Full-Stack Engineer (Front-End Leaning)
Full-Stack Engineer (Front-End Leaning)

Coinscapture • Northern (KY)

Hybrid
USD 120,000 - 170,000
Discretionary stipend
Social travel stipend
Annual offsite
+2