Senior Computer Vision Engineer - Multimodal AI

Shapoorji Pallonji Finance Private Limited

Mumbai

On-site

INR 4,500,000 - 7,500,000

Full time

4 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Shapoorji Pallonji Finance Private Limited is seeking a Senior Computer Vision Engineer specializing in Multimodal AI and Document Intelligence to design, build, and deploy production‑grade systems that make complex visual information machine‑understandable. Role covers OCR, diagram interpretation, charts, and multimodal retrieval across documents, images, and video frames.

This hands‑on product engineering role requires deploying CV and multimodal ML solutions into enterprise environments with

Qualifications

  • 5+ years of relevant experience in Computer Vision, Multimodal AI, or Document AI.
  • Strong Python programming and hands-on PyTorch experience.
  • Proven track record deploying production‑grade CV or multimodal ML systems.
  • Hands-on with OCR, layout analysis, image processing, or visual information extraction.
  • Strong understanding of CNNs, transformers, and vision‑language models.
  • Experience deploying models on GPU infrastructure and optimizing for production.

Responsibilities

  • Design, build, and deploy production-grade document intelligence systems.
  • Develop OCR, diagram interpretation, and chart extraction capabilities.
  • Implement multimodal retrieval, vision-language models, and RAG workflows.
  • Evaluate models with structured benchmarks and monitor production performance.

Tools

ONNX
TensorRT
vLLM
Neo4j

Job description

We build enterprise intelligence platforms that extract and reason over information contained within documents, images, diagrams, charts, drawings, video, and other visual media.

We are looking for a Senior Computer Vision Engineer specialising in Multimodal AI and Document Intelligence to design, build, and deploy production‑grade systems that make complex visual information machine‑understandable.

The role covers document intelligence, OCR, diagram and chart interpretation, vision‑language models, multimodal retrieval, image restoration, and video understanding. You will work across the complete lifecycle, including data preparation, model selection, training, fine‑tuning, evaluation, optimisation, deployment, and production monitoring.

This is a hands‑on product engineering role, not a research‑only position. The person must have experience deploying computer vision or multimodal machine learning solutions into production environments used by enterprise customers.

Key Responsibilities

1. Document Intelligence and OCR

  • Build systems for document layout analysis, region detection, reading‑order reconstruction, table extraction, and structural understanding across scanned and digitally generated documents.
  • Develop and improve multilingual OCR pipelines, including Indic‑language documents, handwriting, stamps, seals, signatures, watermarks, redactions, and annotations.
  • Improve OCR quality through engine selection, preprocessing, model ensembling, post‑correction, and confidence calibration.

2. Visual and Diagram Understanding

  • Build solutions to interpret diagrams, flowcharts, organisation charts, network diagrams, schematics, technical drawings, maps, timelines, and scientific figures.
  • Develop techniques for symbol recognition, connector tracing, label association, dimension extraction, legend interpretation, and relationship identification.
  • Build visual comparison capabilities to identify meaningful changes across document, drawing, or diagram versions.

3. Charts, Tables and Data Visualisation

  • Extract structured information from charts, dashboards, infographics, and complex tables.
  • Develop methods for axis detection, scale inference, legend association, series separation, value extraction, and borderless table reconstruction.
  • Implement confidence scoring where the source image does not allow precise extraction.

4. Multimodal AI and Visual Retrieval

  • Build vision‑language model applications for grounded question answering across documents, images, diagrams, and video frames.
  • Design multimodal retrieval pipelines using visual embeddings, hybrid retrieval, reranking, and late‑interaction approaches such as ColPali or ColQwen.
  • Fine‑tune open‑weight vision‑language models using LoRA, QLoRA, instruction tuning, or other parameter‑efficient techniques.
  • Build multimodal RAG and GraphRAG solutions that preserve source context, structure, relationships, and provenance.

5. Image Processing and Synthetic Data

  • Build image preprocessing and restoration pipelines covering de‑skewing, de‑noising, de‑warping, shadow removal, moiré reduction, binarisation, and super‑resolution.
  • Use generative methods, including diffusion models and GANs, for restoration, inpainting, augmentation, and synthetic data creation.
  • Generate degraded and long‑tail visual samples to test model robustness and improve performance where labelled training data is limited.

6. Video and Temporal Understanding

  • Develop capabilities for keyframe extraction, scene segmentation, object tracking, text extraction, and temporal event detection.
  • Extract slides, screens, documents, and other relevant visual information from recorded video and screen captures.
  • Build solutions for video summarisation, action recognition, and event‑based retrieval.

7. Evaluation and Production Deployment

  • Define evaluation metrics before model development and build golden datasets, regression tests, hallucination checks, structural fidelity measures, and field‑level accuracy benchmarks.
  • Conduct systematic failure analysis across resolution, language, layout, skew, visual density, and domain variation.
  • Implement confidence scoring and human review workflows for uncertain outputs.
  • Optimise models for accuracy, latency, throughput, memory consumption, infrastructure cost, and production reliability.
  • Deploy and monitor models in cloud, on‑premises, or restricted client environments.
Required Experience and Skills
  • 5+ years of relevant experience in Computer Vision, Machine Learning, Document AI, Multimodal AI, or Applied AI Engineering.
  • Strong Python programming skills and deep hands‑on experience with PyTorch.
  • Proven experience building and deploying production‑grade computer‑vision or multimodal machine‑learning systems.
  • Hands‑on experience with document AI, OCR, layout analysis, image processing, or visual information extraction.
  • Strong working knowledge of:
  • Classical computer vision, including geometric transformations, morphology, contour analysis, feature matching, and image registration
  • CNN architectures such as ResNet, EfficientNet, and U‑Net
  • Detection and segmentation models such as YOLO, DETR, Mask R‑CNN, and SAM
  • Vision transformers and vision‑language models such as ViT, CLIP, Qwen‑VL, InternVL, Donut, or Pix2Struct
  • Generative models including diffusion models, GANs, and VAEs
  • Experience training, fine‑tuning, evaluating, and deploying models using GPU infrastructure.
  • Strong understanding of model throughput, batching, memory optimisation, inference cost, and production scalability.
  • Ability to evaluate when classical computer vision, a specialised model, or a large vision‑language model is the most effective solution.
  • Strong analytical ability with a disciplined approach to experimentation, benchmarking, and measurable improvement.
Preferred Experience
  • Multimodal RAG, GraphRAG, knowledge graphs, or Neo4j‑based retrieval.
  • Visual retrieval using multimodal embeddings, hybrid search, reranking, or late‑interaction models.
  • Diagram, schematic, engineering drawing, CAD, BIM, medical imaging, scientific imaging, or geospatial understanding.
  • Indic‑language or multilingual document processing.
  • Model serving and optimisation using ONNX, TensorRT, vLLM, quantisation, or inference batching.
  • Experience deploying solutions within on‑premises, private‑cloud, or air‑gapped environments.
  • Data labelling strategy, annotation workflows, active learning, and synthetic dataset generation.
  • Open‑source contributions, publications, patents, or demonstrable technical work in computer vision, document AI, or multimodal systems.
What We Are Looking For
  • A hands‑on engineer who has taken models beyond experimentation and deployed them in production.
  • Strong problem‑solving ability across unfamiliar document, image, diagram, and video formats.
  • A quality‑first mindset supported by measurable evaluation and structured failure analysis.
  • Ability to balance model accuracy with latency, infrastructure cost, security, and scalability.
  • Strong ownership across model development, deployment, monitoring, and continuous improvement.
  • Ability to collaborate effectively with Product, Data Science, Engineering, and enterprise implementation teams.
Success in This Role
Success will be measured by:
  • Accuracy and structural fidelity of extracted visual information.
  • Effectiveness of retrieval and question answering across visual content.
  • Improvement against defined evaluation benchmarks and golden datasets.
  • Reliability, cost efficiency, and scalability of deployed models.
  • Reduction in manual review through well‑calibrated confidence scoring.
  • Ability to bring new visual formats into production‑grade processing pipelines.
  • Business adoption and performance of visual intelligence capabilities in enterprise products.

Candidates should have experience delivering production‑grade computer vision, document AI, or multimodal AI solutions. Experience limited to academic research, notebooks, prototypes, hackathons, or proofs of concept will not be sufficient.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Applied Machine Learning / Computer Vision Engineer
Applied Machine Learning / Computer Vision Engineer

TechBiz Global GmbH • Mohali

On-site
INR 1,200,000 - 2,200,000
Senior Applied AI/ML Engineer – Computer Vision & Video
Senior Applied AI/ML Engineer – Computer Vision & Video

Objectways • Chennai District

On-site
INR 1,500,000 - 2,500,000
Computer Vision Engineer
Computer Vision Engineer

Nxtwave Disruptive Technologies(Hiring for a client) • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Computer Vision Engineer
Computer Vision Engineer

Ripik.AI • Dadri

On-site
INR 800,000 - 1,200,000
Computer Vision Engineer
Computer Vision Engineer

Difinity Digital • Ernakulam

On-site
INR 1,800,000 - 3,200,000
Deputy Director - AI Solution and Platforms Observability
Deputy Director - AI Solution and Platforms Observability

PepsiCo • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Deputy Director - AI Solution and Platforms Observability
Deputy Director - AI Solution and Platforms Observability

PepsiCo Inc. • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Senior/Lead AI Engineer
Senior/Lead AI Engineer

Modinity Technologies • Zone 7 Ambattur

On-site
INR 1,200,000 - 1,800,000
Computer Vision Developer Onsite (Coimbatore, India)
Computer Vision Developer Onsite (Coimbatore, India)

S27a • Coimbatore District

On-site
INR 1,500,000 - 2,100,000
Vision AI Solution Architect
Vision AI Solution Architect

Tata Consultancy Services • Bengaluru

On-site
INR 3,000,000 - 5,400,000