Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems)

AT&T

Hyderabad

On-site

INR 4,500,000 - 8,000,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

AT&T, a leading technology and telecom company, seeks a Senior/Lead IC in Hyderabad/Bangalore to own platform reliability, lead SRE and DevOps initiatives, and drive automation across CI/CD, observability, and cloud infrastructure.

The role requires hands-on experience with AKS, Prometheus/Grafana, Python automation, and familiarity with Kafka, Confluent Cloud, and event hubs. Onsite work in India with strong cross-functional influence.

Qualifications

  • 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.
  • Strong hands-on expertise with Kubernetes, especially AKS, and cloud-native platform operations.
  • Advanced experience with CI/CD engineering and GitHub Actions.
  • Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.
  • Strong Python automation scripting skills for reliability engineering, platform tooling, and toil reduction.
  • End-user proficiency with AI-assisted productivity and operations tools for incident analysis and troubleshooting acceleration.
  • Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements.
  • Strong hands-on experience with Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.
  • Experience in governance controls: access management and separation of duties.

Responsibilities

  • Own platform reliability practices for availability, resilience, latency, and efficiency.
  • Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.
  • Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.
  • Lead JFROG Helm chart automation and JFROG images/ACR migration work.
  • Support microservices deployment enablement and platform/tooling upgrades.
  • Own and optimize monitoring, alerting, observability, and logging stack components.
  • Support health-check frameworks including Airflow health-check requirements.
  • Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.
  • Collaborate with architecture and delivery teams on reliability and scalability patterns.
  • Lead cloud infrastructure creation, maintenance, governance, and access controls.
  • Drive capacity planning, DR planning/exercises, and platform best-practice documentation.
  • Support cost management, role enforcement, and license governance.
  • Maintain SOP documentation for established alerts and incident patterns.

Skills

SRE/Platform engineering
Kubernetes (AKS)
CI/CD engineering
GitHub Actions
Observability stack
Python automation
AI-assisted tools (prod incidents)
Java/React/Spring Boot familiarity
Confluent Kafka/Cloud
Azure Event Hub / AWS MSK
Apache Flink
OpenSearch / Grafana
Incident leadership

Tools

Confluent Kafka
Confluent Cloud
Azure Event Hub
AWS MSK
Apache Flink
OpenSearch

Job description

Key Responsibilities


  • Own platform reliability practices for availability, resilience, latency, and operational efficiency.

  • Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.

  • Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.

  • Lead JFROG Helm chart automation and JFROG images/ACR migration work.

  • Support microservices deployment enablement and platform/tooling upgrades.

  • Own and optimize monitoring, alerting, observability, and logging stack components: Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.




  • Support health-check frameworks including Airflow health-check requirements.

  • Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.

  • Collaborate with architecture and delivery teams on reliability and scalability patterns.

  • Lead cloud infrastructure creation, maintenance, governance, and access controls.

  • Drive capacity planning, DR planning/exercises, and platform best-practice documentation.

  • Support cost management, role enforcement, and license management governance.

  • Maintain SOP documentation for established alerts and incident patterns.



Required Qualifications / Must-Have Skills


  • 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.

  • Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations.

  • Advanced experience with CI/CD engineering and GitHub Actions.

  • Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.

  • Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.

  • End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).

  • Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).

  • Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.

  • Experience in governance controls: access management, role enforcement, and separation of duties.

  • Proven high-severity incident leadership and post-incident reliability improvement execution.



Good-to-Have / Nice-to-Have


  • Postgres performance and reliability operations.

  • Telecom-scale high-availability systems experience.



Experience Level

Senior to Lead IC (typically 10 to 17 years)



Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)



What We Offer


  • Opportunity to define and scale platform reliability standards.

  • High technical ownership and strong cross-functional influence.

  • Enterprise-scale impact across observability, automation, and resilience engineering.



Weekly Hours

40



Time Type

Regular



Location

IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City, IND:KA:Bangalore / Intl Tech Park, Navigator Bldg, Whitefield Road: Whitefield Road:Intl Tech Park, Navigator Bldg



It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems)
Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems)

Cricket Wireless LLC. • Bengaluru

On-site
INR 4,500,000 - 7,000,000
Career growth
Tech ownership
Impactful projects
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Event Hub & Azure)
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Event Hub & Azure)

AT&T • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Event Hub & Azure)
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Event Hub & Azure)

AT&T • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Azure/Kubernetes ITSM certifications
Sr Specialist-Tier 2 Application Support Engineer (Confluent Kafka, Azure Event Hub & Azure Kub[...]
Sr Specialist-Tier 2 Application Support Engineer (Confluent Kafka, Azure Event Hub & Azure Kub[...]

AT&T • Hyderabad

On-site
INR 1,800,000 - 3,200,000
Specialist App/Prod Support- Tier 2 Kafka Admin,Confluent, Azure, AKS,Kubernetes
Specialist App/Prod Support- Tier 2 Kafka Admin,Confluent, Azure, AKS,Kubernetes

AT&T • Hyderabad

On-site
INR 1,600,000 - 2,400,000
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Event Hub & Azure)
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Event Hub & Azure)

Cricket Wireless LLC. • Bengaluru

On-site
INR 1,200,000 - 2,000,000
High ownership
Reliability focus
Career growth
Sr Specialist System Engineering (Lead Mobility Core (4G/5G) Test Coordinator)
Sr Specialist System Engineering (Lead Mobility Core (4G/5G) Test Coordinator)

AT&T • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Specialist Tech Service Mgmt- Incident Commander or administrators, Service Desk,24/7 rotational
Specialist Tech Service Mgmt- Incident Commander or administrators, Service Desk,24/7 rotational

AT&T • Hyderabad

On-site
INR 2,000,000 - 3,500,000
Sr Specialist System Engineer – Cloud Automation
Sr Specialist System Engineer – Cloud Automation

Cricket Wireless LLC. • Bengaluru

On-site
INR 2,200,000 - 3,500,000
Sr Specialist System Engineering - Mobile Packet Core Certification Automation Engineer
Sr Specialist System Engineering - Mobile Packet Core Certification Automation Engineer

Cricket Wireless LLC. • Bengaluru

On-site
INR 1,200,000 - 2,400,000