Senior Voice Infrastructure Engineer

AethexAI

Indiana (PA)

On-site

USD 150,000 - 230,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

AethexAI is building a latency‑aware voice platform and seeks an experienced infrastructure engineer to own GPU‑driven model serving at scale. You will optimize fleets for cost per minute while maintaining strict latency budgets across regions.

Based in London and working on an in‑office team of ~10, you will drive reliability, observability, and fast architectural decisions to keep the pipeline responsive and secure.

Qualifications

  • 5+ years running production infrastructure with a focus on latency
  • Hands-on GPU inference serving at low latency and GPU fleet cost management
  • Experience with Kubernetes, cloud infra (AWS preferred) and infra-as-code

Responsibilities

  • Own serving STT, LLM, and TTS model pipelines on GPUs within latency budgets
  • Drive cost-per-minute as a core metric
  • Ensure reliability, observability, and on-call readiness across the platform
  • Maintain security and compliance for a multi-tenant voice system across regions

Skills

Kubernetes
AWS cloud infra
Infra-as-code
GPU inference serving
Cost optimization
Observability
Security & compliance
Multi-tenant systems

Tools

vLLM
Triton
TensorRT-LLM

Job description

The problem

Every voice conversation on our platform is a race against a latency budget: speech in, transcription, a language model, speech back out, all on GPUs, all in the time before a human starts to feel the lag. We run that pipeline across Africa and the Middle East, over noisy lines and real dialects, under infrastructure constraints most companies never have to design around. You'll own the systems that make it fast, reliable, and affordable at scale.

Why it's hard here

GPU is where the difficulty and the money live. You'll be orchestrating model serving across a GPU fleet, keeping it warm enough to hit latency targets and lean enough that cost-per-minute keeps falling. Scale-to-zero pools, warm spares, capacity strategy, and the observability to know what any of it is doing under load. This is the core of the job, not a side quest.

What You'll Own
  • Serving STT, LLM, and TTS models on GPU under hard latency budgets, and the fleet behind them
  • Cost-per-minute as a first-class metric you're accountable for driving down
  • Reliability, observability, and on-call posture across the whole platform
  • Security and compliance for a multi-tenant system handling voice and PII across regions
  • The major architectural calls, made independently and fast
What We're Looking For
  • 5+ years running infrastructure in production, not just building it
  • Hands-on GPU inference serving at low latency, and managing GPU fleets and their cost
  • Fluency with Kubernetes, cloud infra (AWS ideally), and infra-as-code
  • Familiarity with the inference serving ecosystem (vLLM, Triton, TensorRT-LLM or similar)
  • A track record of owning complex systems end to end, and shipping from zero with little
Nice to have
  • Real-time or low-latency background: telephony, streaming, audio, or voice
  • First infra hire or founding-era engineer at an early-stage startup
  • FinOps instincts at scale
  • Genuine interest in emerging markets or speech tech
The setup

You'll work directly with our CTO on a team of ~10 that ships constantly. High ownership, short path to decisions, no one leaving you alone when something breaks at 2am. London-based, in-office.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior GPU Infra Engineer - Low-Latency Voice Platform
Senior GPU Infra Engineer - Low-Latency Voice Platform

AethexAI • Indiana (PA)

On-site
USD 150,000 - 230,000
Staff Software Engineer
Staff Software Engineer

Lupitor • San Francisco (CA)

On-site
USD 150,000 - 190,000
GPU Optimization Engineer
GPU Optimization Engineer

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

AI Chopping Block • Northern (KY)

Hybrid
USD 150,000 - 230,000
HPC Infrastructure Engineer - GPU Clusters
HPC Infrastructure Engineer - GPU Clusters

ElevenLabs • Northern (KY)

Hybrid
USD 140,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Netpreme • Northern (KY)

Hybrid
USD 150,000 - 210,000
Performance bonus
Equity grant
Health, dental, vision fully paid
+5
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding Engineer - AI
Founding Engineer - AI

LeoForce • Long Beach (CA)

On-site
USD 180,000 - 275,000
Milestone bonus
Founding AI Engineer
Founding AI Engineer

PetsApp • New York (NY)

On-site
USD 200,000 - 250,000
100% employer-paid health insurance
Flexible PTO
Seed-stage equity grant
+2
Founding Engineer - ML Infrastructure
Founding Engineer - ML Infrastructure

uRun • San Francisco (CA)

On-site
USD 120,000 - 160,000
Health, dental, and vision
401(k)
Paid time off
+2