Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
AethexAI is building a latency‑aware voice platform and seeks an experienced infrastructure engineer to own GPU‑driven model serving at scale. You will optimize fleets for cost per minute while maintaining strict latency budgets across regions.
Based in London and working on an in‑office team of ~10, you will drive reliability, observability, and fast architectural decisions to keep the pipeline responsive and secure.
Every voice conversation on our platform is a race against a latency budget: speech in, transcription, a language model, speech back out, all on GPUs, all in the time before a human starts to feel the lag. We run that pipeline across Africa and the Middle East, over noisy lines and real dialects, under infrastructure constraints most companies never have to design around. You'll own the systems that make it fast, reliable, and affordable at scale.
GPU is where the difficulty and the money live. You'll be orchestrating model serving across a GPU fleet, keeping it warm enough to hit latency targets and lean enough that cost-per-minute keeps falling. Scale-to-zero pools, warm spares, capacity strategy, and the observability to know what any of it is doing under load. This is the core of the job, not a side quest.
You'll work directly with our CTO on a team of ~10 that ships constantly. High ownership, short path to decisions, no one leaving you alone when something breaks at 2am. London-based, in-office.