Senior App/Prod Support (SRE, Kafka, K8s, Azure Resource Management)

Cricket Wireless LLC.

Hyderabad

On-site

INR 3,000,000 - 5,000,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Competitive compensation
On-site work opportunities
Professional development

Job summary

AT&T is seeking a Senior/Lead SRE for onsite roles in Hyderabad/Bangalore to own platform reliability, lead cloud infrastructure governance, and drive automation across CI/CD pipelines and observability components. You will collaborate with architecture and delivery teams to improve resilience and scale across enterprise services.

The role requires deep experience with Kubernetes (AKS), Confluent Kafka ecosystems, and extensive scripting in Python to reduce toil and improve incident response.

Qualifications

  • 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.
  • Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations.
  • Advanced experience with CI/CD engineering and GitHub Actions.
  • Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.
  • Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.
  • End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development not required).
  • Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).
  • Strong hands-on experience with Confluent Kafka/Confluent Cloud/Azure Event Hub/AWS-MSK/Aggregated Flink.

Responsibilities

  • Own platform reliability practices for availability, resilience, latency, and operational efficiency.
  • Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.
  • Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.
  • Lead JFROG Helm chart automation and JFROG images/ACR migration work.
  • Support microservices deployment enablement and platform/tooling upgrades.
  • Own and optimize monitoring, alerting, observability, and logging stack components: Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.
  • Support health-check frameworks including Airflow health-check requirements.
  • Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.
  • Collaborate with architecture and delivery teams on reliability and scalability patterns.
  • Lead cloud infrastructure creation, maintenance, governance, and access controls.
  • Drive capacity planning, DR planning/exercises, and platform best-practice documentation.
  • Support cost management, role enforcement, and license management governance.
  • Maintain SOP documentation for established alerts and incident patterns.

Skills

SRE experience
Kubernetes expertise (AKS)
CI/CD engineering
Observability tooling
Python automation
AI-assisted productivity tools
Production troubleshooting
Java/React/Spring Boot familiarity
Confluent Kafka / Cloud
Azure/AWS cloud platforms
Postgres familiarity

Tools

GitHub Actions
Prometheus/Grafana/AlertManager
Azure Monitor
Thanos/OpenSearch/FluentBit
Confluent Kafka/Confluent Cloud/Azure Event Hub/AWS MSK
Apache Flink

Job description

Key Responsibilities
  • Own platform reliability practices for availability, resilience, latency, and operational efficiency.
  • Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.
  • Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.
  • Lead JFROG Helm chart automation and JFROG images/ACR migration work.
  • Support microservices deployment enablement and platform/tooling upgrades.
  • Own and optimize monitoring, alerting, observability, and logging stack components: Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.
  • Support health-check frameworks including Airflow health-check requirements.
  • Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.
  • Collaborate with architecture and delivery teams on reliability and scalability patterns.
  • Lead cloud infrastructure creation, maintenance, governance, and access controls.
  • Drive capacity planning, DR planning/exercises, and platform best-practice documentation.
  • Support cost management, role enforcement, and license management governance.
  • Maintain SOP documentation for established alerts and incident patterns.
Required Qualifications / Must-Have Skills
  • 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.
  • Strong hands‑on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud‑native platform operations.
  • Advanced experience with CI/CD engineering and GitHub Actions.
  • Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.
  • Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.
  • End‑user proficiency with AI‑assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).
  • Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature‑development role).
  • Strong hands‑on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS‑MSK, and Apache Flink.
  • Experience in governance controls: access management, role enforcement, and separation of duties.
  • Proven high‑severity incident leadership and post‑incident reliability improvement execution.
Good‑to‑Have / Nice‑to‑Have
  • Postgres performance and reliability operations.
  • Telecom‑scale high‑availability systems experience.

Experience Level Senior to Lead IC (typically 10 to 17 years)

Location / Work Mode Onsite (Hyderabad / Bangalore or designated AT&T location)

What We Offer
  • Opportunity to define and scale platform reliability standards.
  • High technical ownership and strong cross‑functional influence.
  • Enterprise‑scale impact across observability, automation, and resilience engineering.

Weekly Hours: 40

Time Type: Regular

Location: IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City, IND:KA:Bangalore / Intl Tech Park, Navigator Bldg, Whitefield Road: Intl Tech Park, Navigator Bldg:Intl Tech Park, Navigator Bldg

AT&T and its subsidiaries are committed to equal employment opportunity. All hiring, promotion, and other employment decisions remain merit‑based and free from discrimination on the basis of race, color, religion, religious creed, national origin, ancestry, age, sex, sexual orientation, gender, gender identity, gender expression, physical disability, mental disability, pregnancy, medical condition, genetic information, marital status, citizenship status, military status, veteran status, or any other characteristic protected by federal, state, or local laws. In addition, AT&T will provide reasonable accommodations to qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

We are pioneers of making connections and have been ever since Alexander Graham Bell invented the telephone and founded our company. That was nearly 150 years ago, and we haven’t stopped innovating since. From the widespread and growing availability of 5G and Fiber to working on things we once only dreamed of—at AT&T, we create connections that change the world.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems)
Senior App/Prod Support (Tier 3 SRE for event-driven ecosystems)

Cricket Wireless LLC. • Bengaluru

On-site
INR 4,500,000 - 7,000,000
Career growth
Tech ownership
Impactful projects
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Devops)
Sr Specialist-Tier 2 Application Support Engineer (Kafka Admin, Devops)

Cricket Wireless LLC. • Hyderabad

On-site
INR 1,800,000 - 2,400,000
Senior App/Prod Support (SRE, Kafka, K8s, Azure Resource Management)
Senior App/Prod Support (SRE, Kafka, K8s, Azure Resource Management)

AT&T • Hyderabad

On-site
INR 3,000,000 - 6,000,000
Platform ownership
Cross-functional influence
Enterprise-scale impact
Senior App/Prod Support (SRE, Kafka, K8s, Azure Resource Management)
Senior App/Prod Support (SRE, Kafka, K8s, Azure Resource Management)

AT&T • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Sr Specialist System Engineering (4G/5G-ANT DevOps/Infra Engineer)
Sr Specialist System Engineering (4G/5G-ANT DevOps/Infra Engineer)

Cricket Wireless LLC. • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Specialist App/Prod Support- Tier 2 Kafka Administrator, AKS, cloud
Specialist App/Prod Support- Tier 2 Kafka Administrator, AKS, cloud

Cricket Wireless LLC. • Hyderabad

On-site
INR 1,500,000 - 2,200,000
Specialist System Engineer – Cloud Automation
Specialist System Engineer – Cloud Automation

Cricket Wireless LLC. • Bengaluru

On-site
INR 1,500,000 - 2,300,000
Specialist System Engineering - Mobile Packet Core Certification Automation Engineer
Specialist System Engineering - Mobile Packet Core Certification Automation Engineer

Cricket Wireless LLC. • Bengaluru

On-site
INR 1,800,000 - 3,200,000
Specialist System Engineering - Platform Automation Engineer
Specialist System Engineering - Platform Automation Engineer

Cricket Wireless LLC. • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Cybersecurity – Endpoint Security (SentinelOne) and Infrastructure Security
Senior Cybersecurity – Endpoint Security (SentinelOne) and Infrastructure Security

Cricket Wireless LLC. • Hyderabad

On-site
INR 3,500,000 - 5,500,000