An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Infolabs Global in Dubai seeks a Senior Computer Vision & Multimodal AI Engineer to build real-time vision systems and next‑generation vision‑language solutions.
You will design end-to-end pipelines, integrate VLMs like LLaVA, Florence, and Qwen‑VL, and optimize for edge and cloud deployment. The role requires 4+ years of hands‑on CV experience with cameras, OpenCV, C++, Python, and DL tooling.
Role: Senior Computer Vision & Multimodal AI Engineer
Experience Level: 4+ Years
Employment Type: Full-Time (Paid Position)
Location: Dubai
We’re Hiring: Senior Computer Vision & VLM Engineer (4+ Years Experience)
Are you passionate about bridging the gap between classic computer vision and frontier Vision-Language Models (VLMs)? We are looking for an experienced Senior Computer Vision Engineer to join our team and build next-generation, real-world visual perception and multimodal AI systems.
If you have spent the last 4+ years working directly with camera pipelines, real-time vision processing, deep learning models, and multimodal architectures, we want to hear from you.
DomainRequired QualificationsExperience4+ years of hands-on professional experience building and deploying production computer vision systems.Deep Learning & VLMsDemonstrated experience with PyTorch/TensorFlow, fine-tuning VLMs/multimodal models, and prompt engineering/grounding.Camera & Video StreamsDeep understanding of camera hardware, frame capture, OpenCV, RTSP streams, and real-time video stream optimization.Edge & PerformanceProficiency in C++ and Python; experience optimizing models with TensorRT, ONNX, CUDA, or similar frameworks.FundamentalsSolid background in linear algebra, geometry, camera calibration, tracking algorithms, and loss function design.