We are looking for a highly experienced Principal Computer Vision & AI Engineer to lead the technical direction of production AI systems operating over live video streams in real-world environments.
This is a hands-on technical leadership role rather than a people-management position. You will architect, build, optimise and deploy advanced computer vision, video understanding and multimodal/Vision-Language Model (VLM) systems while providing technical direction to senior engineers.
Key Responsibilities
- Architect and build production computer vision and video analytics systems.
- Develop solutions across object detection, segmentation, tracking, event understanding and spatial reasoning.
- Design and deploy VLM/multimodal AI systems operating over live and recorded video.
- Own the full ML lifecycle: data strategy, experimentation, fine-tuning, evaluation, deployment and monitoring.
- Optimise models for edge and cloud inference, balancing accuracy, latency, throughput and cost.
- Establish automated model evaluation, regression testing and production monitoring.
- Develop agentic AI systems capable of reasoning over video and interacting safely with tools, APIs, cameras and sensors.
- Set technical architecture and engineering standards while remaining highly hands-on with code.
What We're Looking For
- Deep hands-on experience in Computer Vision / Video AI.
- Strong production experience working with live or streaming video.
- Strong PyTorch and/or TensorFlow experience.
- Expertise across detection, segmentation and tracking using technologies such as YOLO, DETR/RT-DETR, GroundingDINO or equivalent.
- Production experience with Vision-Language Models / multimodal models such as Qwen-VL, InternVL, LLaVA, Gemini or equivalent.
- Experience deploying and optimising AI models in production.
- Exposure to edge AI, ideally TensorRT, ONNX Runtime, NVIDIA Jetson or similar.
- Strong understanding of model evaluation, MLOps, active learning and production monitoring.
- Ability to operate as a Principal/Staff-level technical authority while remaining hands-on.
Nice to Have
- Agentic AI / tool-calling systems such as LangGraph, LangChain or equivalent.
- Real-time CCTV, camera or video-streaming systems.
- Industrial, construction, robotics, autonomous systems or safety-critical AI experience.
- RAG / multimodal retrieval and vector search.
- Self-hosted VLM/LLM serving such as vLLM or NIM.
Critical requirement:
Candidates must have substantial hands-on experience with video-based computer vision systems. Pure GenAI/LLM profiles without strong production video/CV experience will not be suitable.