Principal Computer Vision & AI Engineer
Location: Al Khobar, Saudi Arabia
Role Overview
We are looking for a highly experienced Principal Computer Vision & AI Engineer to lead the technical direction of production AI systems operating over live video streams in real-world environments.
This is a hands‑on technical leadership role rather than a people‑management position. You will architect, build, optimise and deploy advanced computer vision, video understanding and multimodal/Vision-Language Model (VLM) systems while providing technical direction to senior engineers.
Key Responsibilities
- Architect and build production computer vision and video analytics systems.
- Develop solutions across object detection, segmentation, tracking, event understanding and spatial reasoning.
- Design and deploy VLM/multimodal AI systems operating over live and recorded video.
- Own the full ML lifecycle: data strategy, experimentation, fine‑tuning, evaluation, deployment and monitoring.
- Optimise models for edge and cloud inference, balancing accuracy, latency, throughput and cost.
- Establish automated model evaluation, regression testing and production monitoring.
- Develop agentic AI systems capable of reasoning over video and interacting safely with tools, APIs, cameras and sensors.
- Set technical architecture and engineering standards while remaining highly hands‑on with code.
What We're Looking For
- Deep hands‑on experience in Computer Vision / Video AI.
- Strong production experience working with live or streaming video.
- Strong PyTorch and/or TensorFlow experience.
- Expertise across detection, segmentation and tracking using technologies such as YOLO, DETR/RT-DETR, GroundingDINO or equivalent.
- Production experience with Vision‑Language Models / multimodal models such as Qwen‑VL, InternVL, LLaVA, Gemini or equivalent.
- Experience deploying and optimising AI models in production.
- Exposure to edge AI, ideally TensorRT, ONNX Runtime, NVIDIA Jetson or similar.
- Strong understanding of model evaluation, MLOps, active learning and production monitoring.
- Ability to operate as a Principal/Staff‑level technical authority while remaining hands‑on.
Nice to Have
- Agentic AI / tool‑calling systems such as LangGraph, LangChain or equivalent.
- Real‑time CCTV, camera or video‑streaming systems.
- Industrial, construction, robotics, autonomous systems or safety‑critical AI experience.
- RAG / multimodal retrieval and vector search.
- Self‑hosted VLM/LLM serving such as vLLM or NIM.
Critical requirement: Candidates must have substantial hands‑on experience with video‑based computer vision systems. Pure GenAI/LLM profiles without strong production video/CV experience will not be suitable.