Artificial Intelligence / Site Reliability Engineer

Tata Consultancy Services

Bengaluru

On-site

INR 1,500,000 - 2,300,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Tata Consultancy Services seeks an experienced Site Reliability Engineer (SRE) to support its AI Platform in Bengaluru. The role focuses on operating GenAI infrastructure, scaling AI workloads, and ensuring reliability, security, and governance in a regulated financial environment.

You will collaborate with infrastructure, cloud, data, and security teams to deploy, monitor, and optimize training, inference, and data pipelines at scale, balancing cost and performance while driving innovation.

Qualifications

  • Experience in operating and supporting large-scale AI platforms.
  • Hands-on with Kubernetes, cloud services and REST API-based systems.
  • Strong scripting/programming skills in Python/Go/Java.

Responsibilities

  • Operate, monitor, and maintain infrastructure for GenAI apps.
  • Design automation for core platform capabilities and IaC.
  • Establish and enforce SLOs/SLIs/SLAs with dashboards.
  • Lead incident response and postmortems.
  • Perform capacity planning and optimize compute costs.

Skills

SRE infrastructure
IaC
Kubernetes
Cloud platforms
Python/Go/Java
Docker
Monitoring & dashboards
Incident response
Cost optimization

Tools

Terraform
Helm
CloudFormation
Ansible

Job description

Dear Professionals

Greetings from Tata consultancy Services,

Job Title : Artificial Intelligence ( AI) / Site Reliability Engineer ( SRE)

Experiernce:6 to 10 Years

Location: Bengaluru

Mode of Work : Work from Office

Mode of Interview: Virtual

Key Responsibilities

Our mission is to develop a firmwide Artificial Intelligence (AI) Development Platform that aligns with the firms Technology principles and drives efficiency and consistency, controls, security and strong governance and promotes innovation, enabling teams to build applications that leverage AI capabilities and accelerate the adoption of AI across our businesses.

This role is for an experienced and driven Site Reliability Engineer (SRE) to join our AI Platform team to help support, scale and harden the infrastructure that powers our AI/ML systems. You will collaborate closely with infrastrucuture engineering, cloud engineering, data engineering, and security teams to ensure availability, reliability, performance, and security of production AI workloads (training, inference, data pipelines) in a regulated, high-stakes financial environment.

As an SRE on the AI platform, you will bring deep operations, automation, and systems engineering skills to enable our models and pipelines to run reliably at scale, while balancing cost, security, and compliance constraints.

The ideal candidate will have strong hands‑on experience supporting software platforms on any combination of the following platforms - Kubernetes, Cloud (AWS, Azure, and/or Google), API based development, REST framework, data engineering, and large‑scale API Gateway environments etc. Knowledge of AIML and hands‑on experience implementing solutions using Generative AI are also preferable. The candidate will have great communication skills, a team‑based mentality and a strong passion for using AI to increase productivity as well as help generate new ideas for product & technical improvements.

Skills :
  • Operate, monitor, and maintain the infrastructure supporting GenAI applications (training, inference, feature store, data ingestion, model serving)
  • Design and build automation for core platform capabilities, reducing manual toil
  • Develop and maintain infrastructure‑as‑code (IaC) for provisioning and managing compute, storage, network, GPU clusters, Kubernetes / container orchestration, etc.
  • Establish, monitor, and enforce SLOs/SLIs/SLAs, error budgets, alerting, and dashboards
  • Lead incident response, root cause analysis (RCA), postmortems, and systemic remediation
  • Perform capacity planning, scaling strategies, workload scheduling, and resource forecasting
  • Optimize cost vs. performance tradeoffs in large‑scale compute environments
  • Production experience in SRE / Infrastructure / ops for large‑scale systems
  • Strong programming/scripting skills (Python, Go, Java, or equivalent)
  • Deep experience with containerization (Docker), orchestration (Kubernetes, etc.)
  • Infrastructure‑as‑code (Terraform, Helm, CloudFormation, Ansible, etc.)
  • Familiarity with GPU / AI compute clusters, high‑performance data storage, and distributed architectures
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI SRE/ AI Site Reliability Engineer
AI SRE/ AI Site Reliability Engineer

Tata Consultancy Services • Bengaluru

On-site
INR 1,800,000 - 2,800,000
Devops / Site Reliability Engineer (SRE)
Devops / Site Reliability Engineer (SRE)

Tata Consultancy Services • India

On-site
INR 1,800,000 - 2,400,000
SRE Site Reliability Engineering
SRE Site Reliability Engineering

VMC Soft Technologies, Inc • Hyderabad

On-site
INR 3,000,000 - 4,200,000
SRE AI Engineer
SRE AI Engineer

Wissen Technology • Pune District, Mumbai, Bengaluru

Hybrid
INR 900,000 - 1,300,000
AI-MLOps SRE Lead Engineer
AI-MLOps SRE Lead Engineer

Randstad • Hyderabad

Hybrid
INR 2,400,000 - 3,600,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Maharashtra

On-site
INR 700,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior SRE
Senior SRE

Nisum • Hyderabad

On-site
INR 2,800,000 - 4,200,000
SRE Devops AI Engineer
SRE Devops AI Engineer

Wissen Technology • Pune District, Mumbai, Bengaluru

Hybrid
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hilabs • Bengaluru

On-site
INR 2,200,000 - 3,500,000