Vision AI Solution Architect

Tata Consultancy Services

Bengaluru

On-site

INR 3,000,000 - 5,400,000

Full time

2 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Tata Consultancy Services seeks a Vision Solution Architect to lead end-to-end computer vision and multimodal GenAI projects. You will design, implement and deploy architectures across cloud and edge environments, guiding model selection, fine-tuning and productionization for real-world use cases.

The role requires strong Python, PyTorch, and MLOps collaboration, with experience across hyperscalers and edge devices. Bengaluru-based full-time position.

Qualifications

  • CLIP and contrastive embedding models; VLMs for captioning, visual QA and grounding.
  • Design Vision RAG systems with multimodal embeddings and vector databases.
  • Fine-tune/adapt vision/VLMs using LoRA/QLoRA and transfer learning.
  • Experience with two or more cloud vision services (AWS, GCP, Azure).
  • Edge AI deployment with accelerators and quantization techniques.
  • Proficient in Python; strong PyTorch, HF Transformers, and OpenCV skills.
  • Familiar with ONNX/TensorRT/OpenVINO/CoreML tooling and formats.
  • Knowledge of Docker/Kubernetes for scalable serving and orchestration.

Responsibilities

  • Architect end-to-end CV and multimodal GenAI solutions from discovery to deployment.
  • Lead design of CLIP-style embeddings, VLMs, and traditional CV models.
  • Build Vision RAG pipelines with embedding retrieval and grounding.
  • Guide fine-tuning and domain adaptation for accuracy in production.
  • Evaluate services across AWS, GCP, Azure; balance managed vs self-hosted.
  • Architect edge deployments on Jetson/edge TPU with model compression.
  • Collaborate with data science, MLOps and client teams; translate business needs.
  • Document architectures, accelerators, and reusable patterns.

Skills

CLIP embedding models
Vision-Language Models
Vision RAG systems
LoRA/QLoRA fine-tuning
Hyperscaler AI/vision services
Edge AI deployment
Python
PyTorch
Hugging Face Transformers
OpenCV
ONNX/TensorRT/OpenVINO/CoreML
Docker
Kubernetes
Vision model serving
Architectural communication

Tools

PyTorch
Hugging Face Transformers
OpenCV
ONNX
TensorRT
OpenVINO
CoreML
Docker
Kubernetes

Job description

Computer Vision, VLMs & Multimodal GenAI Architecture

Experience Level: 8-12 years

Employment Type: Full-Time

About the Role

We are looking for a Vision Solution Architect to design and lead end-to-end computer vision and multimodal AI solutions — spanning classical vision models, CLIP-style embedding models, Vision-Language Models (VLMs), and Vision RAG architectures. The ideal candidate can architect solutions across cloud hyperscalers and edge deployments, guiding fine-tuning, integration, and productionization of vision and multimodal GenAI systems for real-world business use cases.

Key Responsibilities
  • Architect end-to-end vision and multimodal GenAI solutions, from use-case discovery and model selection through deployment and monitoring.
  • Design solutions leveraging CLIP-style embedding models, VLMs (e.g., LLaVA, Gemini Vision, GPT-4V/Claude Vision, Qwen-VL), and traditional CV models (detection, segmentation, OCR).
  • Design and implement Vision RAG pipelines — multimodal embedding generation, vector indexing/retrieval, and grounding of VLM outputs against visual and textual knowledge bases.
  • Lead fine-tuning and domain adaptation of vision and vision-language models (LoRA/QLoRA, adapter-based tuning, contrastive fine-tuning) for domain-specific accuracy.
  • Architect solutions across hyperscaler vision/AI services — AWS (Rekognition, Bedrock multimodal models, SageMaker), GCP (Vertex AI Vision, Gemini multimodal APIs), and Azure (AI Vision, Azure AI Foundry) — selecting the right managed vs. self-hosted approach.
  • Design edge deployment architectures for vision models on constrained devices (NVIDIA Jetson, edge TPUs, mobile/embedded accelerators), including model compression and quantization.
  • Select and optimize model formats, runtimes, and accelerators (ONNX, TensorRT, OpenVINO, CoreML) for target hardware and latency/throughput requirements.
  • Define evaluation frameworks for vision/multimodal model quality — embedding retrieval accuracy, grounding fidelity, hallucination, and task-specific benchmarks.
  • Collaborate with data science, MLOps, and client-facing teams to translate business requirements into scalable, cost-efficient vision architectures.
  • Document reference architectures, best practices, and reusable accelerators for vision and multimodal GenAI engagements.
Required Skills & Experience
  • Strong hands-on experience with CLIP and other contrastive embedding models, and Vision-Language Models (VLMs) for tasks such as captioning, visual QA, and grounding.
  • Practical experience designing Vision RAG systems, including multimodal embeddings, vector databases, and retrieval-augmented generation patterns.
  • Experience fine-tuning and adapting vision/VLM models using parameter-efficient techniques (LoRA/QLoRA) and classical CV model training/transfer learning.
  • Working knowledge of hyperscaler AI/vision services across at least two of AWS, GCP, and Azure, and their trade-offs for vision workloads.
  • Familiarity with edge AI deployment — hardware accelerators (NVIDIA Jetson, Coral/edge TPU), model compression, quantization, and runtime optimization.
  • Proficiency in Python and vision/ML frameworks (PyTorch, Hugging Face Transformers, OpenCV) and model format/runtime tooling (ONNX, TensorRT, OpenVINO).
  • Understanding of containerization and orchestration (Docker, Kubernetes) for scalable vision model serving.
  • Strong architectural and communication skills, with the ability to translate business needs into technical vision/GenAI solution designs.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Deputy Director - AI Solution and Platforms Observability
Deputy Director - AI Solution and Platforms Observability

PepsiCo • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Deputy Director - AI Solution and Platforms Observability
Deputy Director - AI Solution and Platforms Observability

PepsiCo Inc. • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Computer Vision Engineer
Computer Vision Engineer

Difinity Digital • Ernakulam

On-site
INR 1,800,000 - 3,200,000
SDE II - ML/AI Engineer
SDE II - ML/AI Engineer

Aurigait • Jaipur

On-site
INR 1,200,000 - 1,800,000
AI Engineer — LLM / VLM
AI Engineer — LLM / VLM

SAI Group Ltd • Kolkata District

On-site
INR 1,500,000 - 3,500,000
Computer Vision Engineer
Computer Vision Engineer

Nxtwave Disruptive Technologies(Hiring for a client) • Hyderabad

On-site
INR 3,000,000 - 5,000,000
AI/ML Architect
AI/ML Architect

Relanto • Bengaluru

On-site
INR 400,000 - 700,000
AI Architect
AI Architect

Ethics Infotech • Vadodara

On-site
INR 1,500,000 - 3,200,000
AI/ML Architect
AI/ML Architect

Indihire Consultants • Bengaluru

On-site
INR 3,500,000 - 7,000,000
AI Architect
AI Architect

HSM Edifice Construction Services • Nagpur District

On-site
INR 1,500,000 - 2,200,000