Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.
Compunnel, Inc. is seeking a Real-Time Inference Engineering Lead to design and industrialize low-latency model-serving services for predictive AI use cases.
You will define deployment patterns, capacity controls, monitoring, and high-availability practices across cloud and on-premises environments. The role requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, CI/CD, and technical leadership to mentor inference and platform
The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments. This position requires strong expertise in real-time inference architecture, distributed services, Kubernetes, performance engineering, production operations, CI/CD, and high-availability engineering, along with the ability to provide technical leadership and mentorship to inference and platform engineering teams.