AI SRE/ AI Site Reliability Engineer

Tata Consultancy Services

Bengaluru

On-site

INR 1,800,000 - 2,800,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Tata Consultancy Services in Bengaluru seeks an experienced Senior Site Reliability Engineer to design, implement, and operate scalable, secure platform solutions powering AI/ML applications.

You will automate with Terraform, Helm, Ansible, CloudFormation; manage Kubernetes, Docker, and cloud services; establish SLOs, observability dashboards, and incident response; collaborate with security and engineering teams. 5+ years in SRE/DevOps required.

Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform Engineering
  • Strong programming/scripting skills in Python, Go, Java, or similar languages
  • Hands-on experience with Docker, Kubernetes, and container orchestration
  • Experience with AWS, Azure, or Google Cloud Platform (GCP)
  • Strong knowledge of Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation)
  • Experience with Observability and Monitoring tools such as Grafana, Prometheus, Loki, ELK/EFK, Datadog, OpenTelemetry
  • Knowledge of networking fundamentals: TCP/IP, DNS, Load Balancing, Routing
  • Experience in incident management, RCA, performance tuning, and system scaling
  • Strong understanding of security, compliance, governance, and production support
  • Excellent communication and cross-functional collaboration skills
  • Preferred: Generative AI, LLM, MLOps, ModelOps, or AI Platforms

Responsibilities

  • Manage and support infrastructure powering AI/ML and Generative AI applications
  • Design and implement scalable, highly available, and secure platform solutions
  • Build automation to reduce operational toil and improve platform reliability
  • Develop and maintain Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation)
  • Operate Kubernetes-based environments, container platforms, and cloud services
  • Establish and monitor SLIs, SLOs, SLAs, observability dashboards, and alerting
  • Lead incident response, root cause analysis (RCA), and reliability improvement initiatives
  • Perform capacity planning, performance optimization, and cost management
  • Support GPU-based compute environments, AI model serving, and data pipelines
  • Implement disaster recovery, backup, resiliency, and security controls
  • Collaborate with engineering and security teams to deploy and operate AI services safely
  • Maintain operational documentation, runbooks, and best practices
  • Participate in on-call support and production incident management

Skills

Site Reliability Engineering
Python/Go/Java
Docker & Kubernetes
Cloud Platforms (AWS/Azure/GCP)
Infrastructure as Code (Terraform/Helm
Ansible/CloudFormation
Observability/Monitoring
Networking fundamentals
Incident management & RCA
Security & governance
Communication & collaboration
Generative AI / AI platforms

Tools

Docker
Kubernetes
Terraform
Helm
Ansible
CloudFormation
Grafana
Prometheus
Loki
ELK/EFK
Datadog
OpenTelemetry
Snowflake
Redis
SQL
PostgreSQL
Kafka
Spark
Flink
Slurm

Job description

Key Responsibilities
  • Manage and support infrastructure powering AI/ML and Generative AI applications.
  • Design and implement scalable, highly available, and secure platform solutions.
  • Build automation to reduce operational toil and improve platform reliability.
  • Develop and maintain Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
  • Operate Kubernetes-based environments, container platforms, and cloud services.
  • Establish and monitor SLIs, SLOs, SLAs, observability dashboards, and alerting.
  • Lead incident response, root cause analysis (RCA), and reliability improvement initiatives.
  • Perform capacity planning, performance optimization, and cost management.
  • Support GPU-based compute environments, AI model serving, and data pipelines.
  • Implement disaster recovery, backup, resiliency, and security controls.
  • Collaborate with engineering and security teams to deploy and operate AI services safely.
  • Maintain operational documentation, runbooks, and best practices.
  • Participate in on-call support and production incident management.
Required Skills & Experience
  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform Engineering.
  • Strong programming/scripting skills in Python, Go, Java, or similar languages.
  • Hands‑on experience with Docker, Kubernetes, and container orchestration.
  • Experience with AWS, Azure, or Google Cloud Platform (GCP).
  • Strong knowledge of Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation).
  • Experience with Observability and Monitoring tools such as Grafana, Prometheus, Loki, ELK/EFK, Datadog, OpenTelemetry.
  • Knowledge of networking fundamentals: TCP/IP, DNS, Load Balancing, Routing.
  • Experience in incident management, RCA, performance tuning, and system scaling.
  • Strong understanding of security, compliance, governance, and production support.
  • Excellent communication and cross-functional collaboration skills.
Preferred Skills
  • Experience supporting Generative AI, LLM, MLOps, ModelOps, or AI Platforms.
  • Knowledge of GPU clusters, HPC environments, Slurm, Kubernetes GPU scheduling.
  • Experience with Kafka, Spark, Flink, and distributed data processing frameworks.
  • Familiarity with databases such as Snowflake, Redis, SQL, PostgreSQL.
  • Understanding of Embeddings, Fine-Tuning, RAG, Vector Databases, Model Serving.
  • Experience with Canary Deployments, Blue-Green Deployments, Chaos Engineering.
  • Exposure to financial services or highly regulated environments.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE AI Engineer
SRE AI Engineer

Wissen Technology • Pune District, Mumbai, Bengaluru

Hybrid
INR 900,000 - 1,300,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Maharashtra

On-site
INR 700,000 - 1,800,000
SRE Devops AI Engineer
SRE Devops AI Engineer

Wissen Technology • Pune District, Mumbai, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Dminds Solutions Inc. • Gurugram District

On-site
INR 800,000 - 1,400,000
Remote work stipend
Senior Engineer - AI Platform
Senior Engineer - AI Platform

NetConnectGlobal • Bengaluru

On-site
INR 4,200,000 - 6,000,000
Platform Engineer
Platform Engineer

HCLTech • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Devops Engineer
Devops Engineer

Airtel • Gurugram District

On-site
INR 800,000 - 1,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

SourcingXPress • Hyderabad

On-site
INR 3,000,000 - 5,000,000
AI SRE Engineer
AI SRE Engineer

EDGE Executive Search • Bengaluru

Hybrid
INR 1,800,000 - 3,200,000