Public AI is building the Public AI Inference Utility and seeks its first dedicated infrastructure hire. You will own the platform’s operational backbone, reporting to the CTO, and help productionize public AI services with reliability and transparency. In the first months you will harden the platform for launches, build end-to-end observability, implement routing provenance, and improve performance of chat.publicai.co, while collaborating with open-source stacks and partners.
Qualifications
Significant experience operating production inference or ML serving infrastructure.
Strong distributed systems and SRE instincts: observability, incident response.
Ability to autonomously work in a small team and travel occasionally.
Experience coordinating multiple open-source stacks and projects.
Responsibilities
Technical owner of the platform's operational backbone, reporting to the CTO.
Harden the platform for launches, load testing, and fallback routing.
Build end-to-end observability across OpenWebUI, AWS, CSCS, and Infomaniak.
Implement routing transparency and endpoint provenance for inferences.
Improve reliability and performance of chat.publicai.co.
Skills
Production ML infra
Distributed systems/SRE
Orchestrating stacks
Maintainer/integrator mindset
Autonomy & travel
Tools
OpenWebUI
vLLM
Job description
Public AI is building the Public AI Inference Utility and seeks its first dedicated infrastructure hire. You will own the platform’s operational backbone, reporting to the CTO, and help productionize public AI services with reliability and transparency. In the first months you will harden the platform for launches, build end-to-end observability, implement routing provenance, and improve performance of chat.publicai.co, while collaborating with open-source stacks and partners.