Senior GPU Infra Engineer - Low-Latency Voice Platform

AethexAI

Indiana (PA)

On-site

USD 150,000 - 230,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

AethexAI is building a latency‑aware voice platform and seeks an experienced infrastructure engineer to own GPU‑driven model serving at scale. You will optimize fleets for cost per minute while maintaining strict latency budgets across regions.

Based in London and working on an in‑office team of ~10, you will drive reliability, observability, and fast architectural decisions to keep the pipeline responsive and secure.

Qualifications

  • 5+ years running production infrastructure with a focus on latency
  • Hands-on GPU inference serving at low latency and GPU fleet cost management
  • Experience with Kubernetes, cloud infra (AWS preferred) and infra-as-code

Responsibilities

  • Own serving STT, LLM, and TTS model pipelines on GPUs within latency budgets
  • Drive cost-per-minute as a core metric
  • Ensure reliability, observability, and on-call readiness across the platform
  • Maintain security and compliance for a multi-tenant voice system across regions

Skills

Kubernetes
AWS cloud infra
Infra-as-code
GPU inference serving
Cost optimization
Observability
Security & compliance
Multi-tenant systems

Tools

vLLM
Triton
TensorRT-LLM

Job description

AethexAI is building a latency‑aware voice platform and seeks an experienced infrastructure engineer to own GPU‑driven model serving at scale. You will optimize fleets for cost per minute while maintaining strict latency budgets across regions.

Based in London and working on an in‑office team of ~10, you will drive reliability, observability, and fast architectural decisions to keep the pipeline responsive and secure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Voice Infrastructure Engineer
Senior Voice Infrastructure Engineer

AethexAI • Indiana (PA)

On-site
USD 150,000 - 230,000
GPU Optimization Engineer
GPU Optimization Engineer

techire ai • San Francisco (CA)

On-site
USD 230,000 - 300,000
Senior GPU Video Streaming Engineer (Low-Latency Cloud)
Senior GPU Video Streaming Engineer (Low-Latency Cloud)

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Senior GPU Cloud Infrastructure Engineer
Senior GPU Cloud Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 240,000
Platform Engineer
Platform Engineer

Harrison Clarke • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior AI Platform Architect — GPU & Inference
Senior AI Platform Architect — GPU & Inference

Accellor • Mountain View (CA)

On-site
USD 180,000 - 280,000
Voice AI Engineer – Equity, Multilingual, Low-Latency
Voice AI Engineer – Equity, Multilingual, Low-Latency

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 450,000
Equity
Medical, vision, dental coverage
401(k) retirement plan
+3
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity
Senior InfraOps Engineer — GPU Infra, Hybrid, Equity

Lightning AI • New York (NY)

Hybrid
USD 160,000 - 200,000
Health coverage
Equity/RSUs
401(k) matching
+1
Staff AI Infra Engineer: GPU Fleet Reliability Leader
Staff AI Infra Engineer: GPU Fleet Reliability Leader

Luma AI • United States

Remote
USD 210,000 - 320,000
Low-Latency Inference Systems Engineer (Multi-GPU)
Low-Latency Inference Systems Engineer (Multi-GPU)

Digital Waffle • San Francisco (CA)

Hybrid
USD 180,000 - 280,000