Get more replies from employers
Send a job-specific resume in minutes.
techire.® seeks engineers to build the infrastructure that runs next-generation multimodal foundation models at scale. You will design and implement real-time inference pipelines and distributed systems to support low latency, high reliability production workloads.
You’ll collaborate with researchers to productionise new model architectures and drive ownership from day one, focusing on end-to-end reliability and observability across the stack.
Help build the inference stack behind the next generation of multimodal foundation models.
Most inference roles are about making existing models run faster.
This one is about helping define how entirely new model architectures are served at scale.
You’ll be building real-time multimodal AI capable of processing enormous streams of text, audio and video. The research is pushing beyond today’s transformer limitations, and your work will make those models usable in production.
You’ll sit between frontier research and product engineering, designing the infrastructure that allows cutting-edge models to run with low latency, high reliability and at scale. If you enjoy solving systems problems where every millisecond matters, you’ll feel at home here.
Experience with vLLM, SGLang, Continuous Batching, CUDA or Triton would be particularly valuable but isn’t essential.
Alongside compensation, you’ll receive fully covered medical, dental and vision insurance, 401(k), relocation and immigration support, plus daily meals in the office.
If you’re interested in building the infrastructure that enables the next generation of AI models to run in real time, we’d love to tell you more.