Principal SRE: Proactive Reliability & Observability

AT&T

Atlanta (GA)

On-site

USD 120,000 - 180,000

Full time

2 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

401(k) plan
Paid Time Off and Holidays
Paid Caregiver Leave
Adoption Reimbursement
Disability Benefits
Employee Discounts

Job summary

AT&T seeks a seasoned Systems Engineer to analyze incidents end-to-end across cloud and on‑prem environments, delivering high‑quality postmortems and driving preventive actions. You will partner with engineering to implement permanent fixes, using AI-assisted analysis and data visualization to prevent recurrence.

You will apply systems thinking, QA/SRE background, and IA/DevOps practices to improve reliability and inform stakeholders across teams.

Qualifications

  • 7+ years in Systems Engineering, ITSM, RM/CM.
  • Background in SRE, Support or QA.
  • Hands-on experience with SRE tools: T-APM, T-Trace, CatchPoint, Grafana.
  • Experience with AI tech, data analytics and ITSM tools like ServiceNow.
  • Knowledge of Enterprise Release/Change Management practices.
  • Experience with SAFe, Agile, DevOps and Gen AI use cases.

Responsibilities

  • Analyze incidents end-to-end across applications, infra and cloud environments.
  • Turn incident insights into postmortems and drive corrective actions.
  • Partner with engineering to implement permanent fixes and preventive improvements.
  • Leverage data and AI-assisted analysis to shift from reactive to proactive reliability.
  • Communicate complex issues clearly to diverse stakeholders.

Skills

End-to-end system architecture
Observability tools
Pattern detection
Postmortems writing
Data analysis / AI-assisted methods
Systems thinking & problem solving
QA / SRE / automation background
AI technologies / Python / SQL
Data visualization (Power BI, Tableau)

Education

BS in Computer Science or related field

Tools

SAFe
Agile
DevOps
CI/CD
Power BI
SQL
Python

Job description

AT&T seeks a seasoned Systems Engineer to analyze incidents end-to-end across cloud and on‑prem environments, delivering high‑quality postmortems and driving preventive actions. You will partner with engineering to implement permanent fixes, using AI-assisted analysis and data visualization to prevent recurrence.

You will apply systems thinking, QA/SRE background, and IA/DevOps practices to improve reliability and inform stakeholders across teams.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal SRE: End-to-End Reliability & AI Insights
Principal SRE: End-to-End Reliability & AI Insights

AT&T • Dallas (TX)

On-site
USD 120,000 - 180,000
401(k) plan
Paid Time Off and Holidays
Paid Caregiver Leave
+5
Lead Software Engineer, SRE & Onboarding Automation
Lead Software Engineer, SRE & Onboarding Automation

AT&T • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Principal System Engineering - SRE
Principal System Engineering - SRE

AT&T • Atlanta (GA)

On-site
USD 120,000 - 180,000
401(k) plan
Paid Time Off and Holidays
Paid Caregiver Leave
+3
Lead SRE – AI-Driven Onboarding & Observability
Lead SRE – AI-Driven Onboarding & Observability

AT&T • Atlanta (GA)

On-site
USD 141,300 - 237,400
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+5
Principal System Engineering - SRE
Principal System Engineering - SRE

AT&T • Dallas (TX)

On-site
USD 120,000 - 180,000
401(k) plan
Paid Time Off and Holidays
Paid Caregiver Leave
+5
Lead Software Engineer - Reliability Engineering (SRE)
Lead Software Engineer - Reliability Engineering (SRE)

AT&T • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Senior Cloud Platform & Reliability Architect
Senior Cloud Platform & Reliability Architect

AT&T • City of Middletown (NY)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+5
Associate SRE: Automation & Observability
Associate SRE: Automation & Observability

Calabrio • United States

On-site
USD 90,000 - 130,000
Associate SRE: Automation & Reliability
Associate SRE: Automation & Reliability

Socket.dev • United States

Remote
USD 90,000 - 120,000
Senior SRE Lead: Reliability, Automation & Incidents
Senior SRE Lead: Reliability, Automation & Incidents

Shield AI • San Diego (CA)

On-site
USD 183,000 - 275,000
Equity
Bonus
Benefits