SRE Operations Lead - AWS, GitLab & AIOps

Tata Consultancy Services

Charlotte (NC)

On-site

USD 70,000 - 120,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Tata Consultancy Services is seeking an experienced SRE/ARE leader to oversee 24x7 production operations, major incident management, and reliability governance. You will drive observability across Dynatrace, Grafana, CloudWatch, and Splunk, while leading automation initiatives with GitLab CI/CD and DevSecOps practices.

The role focuses on optimizing MTTR, SLIs/SLOs, and reliability KPIs, including multi-region AWS and batch processing workloads.

Qualifications

  • SRE/ARE with 24x7 production operations experience.
  • Experience in major incident management and service governance.
  • Strong knowledge of AWS services and observability tooling.
  • Proficient in automation, DevSecOps, and REST APIs.
  • Proven ability to lead incident response and post-incident reviews.

Responsibilities

  • Lead 24x7 SRE operations and incident response coordination.
  • Improve reliability metrics (SLIs, SLOs, SLAs, MTTR).
  • Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, Splunk.
  • Lead automation, AIOps, and AI-driven operations initiatives.
  • Oversee AWS batch processes and Control-M environments.

Skills

SRE/ARE
Major Incident Mgmt
AWS proficiency
Observability tools
Automation & DevSecOps
GitLab CI/CD
Python automation
AIOps
Incident response

Education

Bachelor's degree in Computer Science

Tools

Dynatrace
Grafana
Splunk
CloudWatch
Docker
ECR
Bedrock
ServiceNow

Job description

Job Description
  • SRE / Application Reliability Engineering (ARE) and 24x7 production operations
  • Major Incident Management (P1/P2), ServiceNow, and Incident / Problem / Change Management
  • SLI, SLO, SLA governance; MTTR reduction; service reliability KPIs
  • AWS: CloudWatch, Route 53, S3, CloudFront, Lambda, ECR, and Bedrock
  • Observability: Dynatrace, Grafana, Splunk, and CloudWatch
  • Control-M and enterprise batch operations
  • GitLab CI/CD, REST APIs, Personal Access Tokens, security scanning, and DevSecOps
  • Python GitLab library and API-based automation
  • Automation, AIOps, event correlation, self-healing, and stakeholder management
  • Lead 24x7 SRE operations and coordinate P1/P2 major incident response through service restoration and follow-up.
  • Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.
  • Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.
  • Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.
  • Oversee AWS platform operations, batch processing, and Control-M environments.
  • Integrate Claude AI on AWS Bedrock with GitLab using APIs, PATs, and custom workflows.
  • Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines, and security findings.
  • Design GitLab API automation, custom workflows, DevSecOps controls, and CI/CD pipeline improvements.
  • Build and maintain Python-based GitLab integrations and REST API solutions.
  • Support vulnerability remediation and onboarding/configuration of security scanning tools.
  • Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda configuration, concurrency, and scaling.
  • Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.
Must Have Technical/Functional Skills
  • SRE / Application Reliability Engineering (ARE) and 24x7 production operations
  • Major Incident Management (P1/P2), ServiceNow, and Incident / Problem / Change Management
  • SLI, SLO, SLA governance; MTTR reduction; service reliability KPIs
  • AWS: CloudWatch, Route 53, S3, CloudFront, Lambda, ECR, and Bedrock
  • Observability: Dynatrace, Grafana, Splunk, and CloudWatch
  • Control-M and enterprise batch operations
  • GitLab CI/CD, REST APIs, Personal Access Tokens, security scanning, and DevSecOps
  • Python GitLab library and API-based automation
  • Automation, AIOps, event correlation, self-healing, and stakeholder management
Roles & Responsibilities
  • Lead 24x7 SRE operations and coordinate P1/P2 major incident response through service restoration and follow-up.
  • Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.
  • Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.
  • Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.
  • Oversee AWS platform operations, batch processing, and Control-M environments.
  • Integrate Claude AI on AWS Bedrock with GitLab using APIs, PATs, and custom workflows.
  • Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines, and security findings.
  • Design GitLab API automation, custom workflows, DevSecOps controls, and CI/CD pipeline improvements.
  • Build and maintain Python-based GitLab integrations and REST API solutions.
  • Support vulnerability remediation and onboarding/configuration of security scanning tools.
  • Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda configuration, concurrency, and scaling.
  • Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.
Generic Managerial Skills
  • Lead geographically distributed operations teams and coordinate effectively during critical incidents.
  • Communicate reliability risks, service performance, and remediation plans to technical and business stakeholders.
  • Drive governance, prioritization, continuous improvement, and cross-team collaboration.
  • Mentor engineers and promote automation-first, blameless, and reliability-focused ways of working.
Salary - $70,000 - $120,000 per annum.
Qualifications

BACHELOR OF COMPUTER SCIENCE

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

SRE Operations Lead - AWS, GitLab & AIOps
SRE Operations Lead - AWS, GitLab & AIOps

Envision Technology Solutions • Charlotte (NC)

On-site
USD 140,000 - 190,000
Sr SRE Automation Engineer
Sr SRE Automation Engineer

Compunnel, Inc. • Austin (TX), Northern (KY)

Hybrid
USD 130,000 - 180,000
Senior Lead Site Reliability Engineer
Senior Lead Site Reliability Engineer

JPMorgan Chase & Co. • Jersey City (NJ)

On-site
USD 150,000 - 210,000
SRE Engineer
SRE Engineer

Tata Consultancy Services • Englewood Cliffs (NJ)

On-site
USD 110,000 - 125,000
SRE AWS DevOps with Arize
SRE AWS DevOps with Arize

Tata Consultancy Services • Malvern

On-site
USD 155,000 - 170,000
Discretionary Annual Incentive
Comprehensive Medical Coverage
401K Plan
Senior Site Reliability Engineer
Senior Site Reliability Engineer

GovCIO • Arlington (VA)

On-site
USD 210,000 - 230,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

CardWorks • Pittsburgh

Hybrid
USD 146,000 - 163,000
Competitive Pay
Medical, Dental, and Vision Benefits
401(k) Plan with Company Match
+1
Senior Software Engineer - IP&R Reliability Engineering (Remote)
Senior Software Engineer - IP&R Reliability Engineering (Remote)

The Home Depot • Atlanta (GA)

On-site
USD 130,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Axle • Frederick (MD)

On-site
USD 140,000 - 155,000
Paid Time Off
401K match
Educational Benefits
+5
Site Reliability Engineer
Site Reliability Engineer

TalentDome Staffing • United States

On-site
USD 140,000 - 210,000