Ai Ml Engineer

Tata Consultancy Services

Gandhinagar, Indore District, Chennai District

On-site

INR 3,000,000 - 5,000,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Tata Consultancy Services is seeking a senior AI/ML architect to design and govern enterprise-scale data pipelines, models, and platform integrations on Google Cloud. You’ll lead complex deployments, optimize GPUs/TPUs, and drive reliability across CI/CD, monitoring, and governance workflows.

The role requires deep cloud and ML expertise, strong communication with executives, and hands-on work with Vertex AI, Kubeflow, and related tooling in a distributed enterprise environment.

Qualifications

  • Advanced Python/SQL and ML/DL framework expertise.
  • Deep expertise in AI/ML algorithms and DL architectures.
  • Experience with MLOps, CI/CD for ML and model governance.
  • Hands-on with enterprise Vertex AI/Kubeflow/MLflow ecosystem.

Responsibilities

  • Architect complex AI/ML pipelines and data preprocessing.
  • Lead RCA for infrastructure and pipeline issues at scale.
  • Collaborate with C-level stakeholders and architects.
  • Provide architectural guidance and 24x7 advisory support.

Skills

Python/SQL
ML/DL frameworks
PyTorch
TensorFlow
MLOps
CI/CD for ML
Docker
Kubernetes
GCP
Vertex AI

Tools

Kubeflow
MLflow
Triton
vLLM
TensorRT-LLM

Job description

Job Description:
  • Advanced proficiency in Python/SQL, deep expertise in ML/DL frameworks (PyTorch, TensorFlow), and architecting complex data pre-processing/feature engineering pipelines. Deep Subject Matter Expertise in Core AI/ML algorithms, Deep Learning architectures, and Generative AI (LLMs, RAG, Transformer models). Enterprise-level Machine Learning Operations (MLOps), including CI/CD for ML, automated retraining workflows, and model governance. Deep hands-on expertise with the broader AI/ML Ecosystem Tools (Kubeflow, MLflow). Expert-level understanding of GCP compute, storage, networking, IAM, and the entire Vertex AI suite (Workbench, Pipelines, Feature Store, Model Registry, Endpoints, Agent Builder).
  • Extensive experience in provisioning, configuring, and optimizing distributed GPU and Google Cloud TPU (v4/v5) environments for training and inference. Ability to analyze distributed logs, profile CUDA/memory bottlenecks, and diagnose deep-seated infrastructure/pipeline failures. Expert knowledge of designing, integrating, and troubleshooting enterprise RESTful and gRPC AI APIs and microservices. Leading Root Cause Analysis (RCA), diagnosing systemic issues, and implementing permanent architectural resolutions for complex technical problems. Providing high-level technical advisory, architectural consulting, and escalation support to internal engineering teams and enterprise clients. Extensive experience architecting, fine-tuning, and optimizing Conversational Agents, Large Language Models, and Dialogflow CX systems. Exceptional ability to communicate intricate technical information and architectural trade-offs clearly and concisely to C-level stakeholders, architects, and engineering teams. Support with 24x7 operations (Rotational Shifts). English language (verbal and written) proficiency is a must.
Key
  • Troubleshooting Model & Pipeline Issues API & Integration Support: Diagnosing, optimizing, and resolving complex RESTful and gRPC API integrations, latency bottlenecks, and payload serialization issues between client enterprise applications and Vertex AI endpoints. Inference Failures: Leading root-cause investigations for critical inference failures, resolving memory overflows (OOM), optimizing GPU/TPU resource allocation, and fine-tuning serving runtimes (e.g., Triton, vLLM, TensorRT-LLM). Environment Configuration: Architecting and troubleshooting custom Docker containers, GKE/Kubernetes clusters, and cloud-native GCP ML environments with strict enterprise networking and VPC Service Controls (VPC-SC).

  • Data & Performance Monitoring Data Quality Checks: Designing and implementing automated validation frameworks to identify data quality anomalies, schema mismatches, and pipeline corruptions across BigQuery, Dataflow, and Vertex AI Feature Store. Monitoring Drift: Architecting enterprise-grade monitoring solutions using Vertex AI Model Monitoring to detect data drift and concept drift, establishing automated alerting and retraining triggers.

  • Accuracy Inquiries: Providing deep technical analysis on model predictions, bias, and reliability using advanced interpretability and explainability frameworks.

  • Product Education & Technical Documentation Knowledge Base Authoring: Authoring enterprise-grade reference architectures, technical blueprints, and definitive best-practice guides on topics like "Distributed Training on TPUs,", "Production RAG Architecture,", and "LLM Fine-Tuning on Vertex AI." Customer Onboarding: Leading architectural reviews, technical onboarding for enterprise engineering and data science teams. Translating Documentation: Synthesizing complex GCP product roadmaps, cutting-edge AI research, and core engineering release notes into actionable implementation strategies for technical leadership and IT teams. The "Feedback Bridge" to Engineering Bug Reporting: Identifying, reproducing, and isolating complex platform defects, performing core-level debugging, and collaborating directly with Google Cloud / Product Engineering teams to drive fixes. Feature Requests: Aggregating enterprise-level capability gaps, creating detailed technical RFCs, and partnering with Product Management to influence the GCP AI/ML product roadmap. Edge Case Discovery: Documenting unique edge cases and failure modes where AI models/infrastructure fail under load, designing guardrails to improve future model resilience and system stability.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Support Engineer | GCP | Vertex AI | Python | ML Ops (Tech suppo
AI/ML Support Engineer | GCP | Vertex AI | Python | ML Ops (Tech suppo

Tata Consultancy Services • Chennai District

On-site
INR 2,800,000 - 4,200,000
Principal Generative AI & ML Ops Engineer (SME)
Principal Generative AI & ML Ops Engineer (SME)

Tata Consultancy Services • Gandhinagar, Indore District, Chennai District

On-site
INR 1,800,000 - 3,000,000
AI/ML SME - Generative AI | Vertex AI | LLM | RAG | GCP
AI/ML SME - Generative AI | Vertex AI | LLM | RAG | GCP

Tata Consultancy Services • Chennai District

On-site
INR 2,800,000 - 4,800,000
Mlops/AI Engineer - (GCP, Vertex AI, Watsonx, Agentic AI)
Mlops/AI Engineer - (GCP, Vertex AI, Watsonx, Agentic AI)

CIEL HR • Chennai District

On-site
INR 2,500,000 - 5,500,000
Senior Lead AI Engineer
Senior Lead AI Engineer

Wilco Source • Pune District

Hybrid
INR 3,500,000 - 5,500,000
Technical Architect - ML
Technical Architect - ML

Prodapt Solutions Private Limited • Chennai District

On-site
INR 5,000,000 - 7,500,000
Technical Architect - ML
Technical Architect - ML

Prodapt • Chennai District

On-site
INR 4,000,000 - 6,000,000
AI ML Engineer(Technical Support Representative)
AI ML Engineer(Technical Support Representative)

Tata Consultancy Services • Chennai District

On-site
INR 700,000 - 1,100,000
AIML Architect
AIML Architect

Infovision Labs • Bengaluru

On-site
INR 4,500,000 - 6,500,000
AI Infrastructure Architect
AI Infrastructure Architect

Accenture in India • Mumbai

On-site
INR 3,000,000 - 5,400,000