Principal System Engineering - SRE

AT&T

Atlanta (GA)

On-site

USD 120,000 - 180,000

Full time

4 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

401(k) plan
Paid Time Off and Holidays
Paid Caregiver Leave
Adoption Reimbursement
Disability Benefits
Employee Discounts

Job summary

AT&T seeks a seasoned Systems Engineer to analyze incidents end-to-end across cloud and on‑prem environments, delivering high‑quality postmortems and driving preventive actions. You will partner with engineering to implement permanent fixes, using AI-assisted analysis and data visualization to prevent recurrence.

You will apply systems thinking, QA/SRE background, and IA/DevOps practices to improve reliability and inform stakeholders across teams.

Qualifications

  • 7+ years in Systems Engineering, ITSM, RM/CM.
  • Background in SRE, Support or QA.
  • Hands-on experience with SRE tools: T-APM, T-Trace, CatchPoint, Grafana.
  • Experience with AI tech, data analytics and ITSM tools like ServiceNow.
  • Knowledge of Enterprise Release/Change Management practices.
  • Experience with SAFe, Agile, DevOps and Gen AI use cases.

Responsibilities

  • Analyze incidents end-to-end across applications, infra and cloud environments.
  • Turn incident insights into postmortems and drive corrective actions.
  • Partner with engineering to implement permanent fixes and preventive improvements.
  • Leverage data and AI-assisted analysis to shift from reactive to proactive reliability.
  • Communicate complex issues clearly to diverse stakeholders.

Skills

End-to-end system architecture
Observability tools
Pattern detection
Postmortems writing
Data analysis / AI-assisted methods
Systems thinking & problem solving
QA / SRE / automation background
AI technologies / Python / SQL
Data visualization (Power BI, Tableau)

Education

BS in Computer Science or related field

Tools

SAFe
Agile
DevOps
CI/CD
Power BI
SQL
Python

Job description

Join AT&T and reimagine the communications and technologies that connect the world. The Chief Information Office is responsible for advancing information technology performance and delivering solutions with a focus on maximizing ROI, increasing efficiency and enhancing the experience of end users. Guided by experienced leaders, Corporate Systems seamlessly integrate with advanced Technology and Operations to drive our enterprise forward. Our Systems Reliability and Software Delivery teams are unwavering in their commitment to excellence, ensuring every solution is robust and efficient. When you step into a career with AT&T, you won’t just imagine the future-you’ll create it.

What you’ll do:

In this role, you will focus on understanding why production incidents happen and how to prevent them from recurring. You will analyze incidents end-to-end across applications, infrastructure, and cloud environments, using observability data to identify root causes, patterns, and systemic weaknesses.

You will turn incident insights into high-quality postmortems and partner with engineering teams to drive corrective actions and long-term improvements. By combining system-level thinking with data, automation, and AI-assisted analysis, you will help shift the organization from reactive response to proactive reliability and incident prevention. You will partner with engineering and software development teams to implement permanent fix and preventive improvements

What you'll bring:

  • Strong understanding of end-to-end system architecture (cloud, web apps, APIs, databases, infrastructure)
  • Hands‑on experience with observability tools (logs, metrics, traces)
  • Ability to identify patterns and drive preventive actions
  • Experience writing clear, structured postmortems
  • Ability to analyze operational data using tools, queries, or AI-assisted methods
  • Strong systems thinking and problem‑solving skills
  • Background in QA, test engineering, or automation engineering (strong plus)
  • Experience using AI or advanced analytics for incident analysis or pattern detection
  • Understanding of distributed systems and failure modes
  • Experience with data analysis / visualization tools (e.g., Power BI, Tableau)
  • Mindset focused on eliminating recurring issues, not just fixing incidents
  • Strong communication skills to explain complex issues clearly

Required:

  • 7+ years in Systems Engineering, ITSM, RM/CM
  • Background in SRE, Support or QA
  • One or more of the following SRE Tools: T-APM, T-Trace, CatchPoint, Grafana
  • Hands‑on experience and understanding of concepts and tools such as SAFe, Agile, DevOps, CI/CD, Data Analytics, and building Gen AI use cases
  • Experience with AI technologies, Python, SQL, data analytics, Power BI and ITSM tools (e.g., ServiceNow)
  • Modern Enterprise Release Management/Change Management and ITSM

Preferred:

  • BS/BA in Computer Science
  • Preferred tools: modern Release Management processes for Agile and DevOps environments
  • Jira Align, JSM, Jira Cloud, Git for enterprise RM/CM
  • Relevant certifications (SAFe, Agile, DevOps, AI/ML)

Joining our team comes with amazing perks and benefits:

  • 401(k) plan
  • Paid Time Off and Holidays (based on date of hire, at least 23 days of vacation each year and 9 company-designated holidays)
  • Paid Caregiver Leave
  • Additional sick leave beyond what state and local law require may be available but is unprotected
  • Adoption Reimbursement
  • Disability Benefits (short term and long term)
  • Life and Accidental Death Insurance
  • Supplemental benefit programs: critical illness/accident hospital indemnity/group legal
  • Employee Assistance Programs (EAP)
  • Extensive employee wellness programs
  • Employee discounts up to 50% off on eligible AT&T mobility plans and accessories
  • AT&T internet (and fiber where available) and AT&T phone.

#LI-Onsite – Full-time office role

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal System Engineering - SRE
Principal System Engineering - SRE

AT&T • Dallas (TX)

On-site
USD 120,000 - 180,000
401(k) plan
Paid Time Off and Holidays
Paid Caregiver Leave
+5
Lead Software Reliability Engineering
Lead Software Reliability Engineering

AT&T • Atlanta (GA)

On-site
USD 141,300 - 237,400
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+5
Principal System Engineering
Principal System Engineering

AT&T • City of Middletown (NY)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+5
Principal System Engineering
Principal System Engineering

AT&T • Atlanta (GA)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+11
Principal System Engineering
Principal System Engineering

AT&T • Plano (TX)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement
+5
Principal System Engineering
Principal System Engineering

AT&T • Alpharetta (GA)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement
+7
Principal System Engineering
Principal System Engineering

AT&T • Dallas (TX)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Paid Time Off and Holidays
+6
Principal System Engineering
Principal System Engineering

AT&T • Bedminster Township (NJ)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+7
Principal System Engineering
Principal System Engineering

AT&T • Middletown (NJ)

On-site
USD 155,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+5
Lead Software Engineer - Reliability Engineering (SRE)
Lead Software Engineer - Reliability Engineering (SRE)

AT&T • Town of Texas (WI)

On-site
USD 140,000 - 190,000