Principal System Engineering - SRE

AT&T

Dallas (TX)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

401(k) plan
Paid Time Off and Holidays
Paid Caregiver Leave
Adoption Reimbursement
Disability Benefits
Life Insurance
Employee wellness programs
Employee discounts on AT&T plans

Job summary

AT&T in Dallas is seeking a senior Systems Reliability/ITSM professional to prevent recurring production incidents. You will analyze incidents across applications, infrastructure, and cloud environments, using observability data to identify root causes and systemic weaknesses.

You will create high-quality postmortems and lead engineering teams to implement permanent fixes and preventive improvements. Bring hands-on experience with AI-enabled incident analysis, Python, SQL, data analytics, and BI

Qualifications

  • 7+ years in Systems Engineering, ITSM, RM/CM
  • Background in SRE, Support or QA
  • One or more of the following SRE Tools: T-APM, T-Trace, CatchPoint, Grafana
  • Hands-on experience with SAFe, Agile, DevOps, CI/CD, Data Analytics, and Gen AI use cases

Responsibilities

  • Analyze production incidents end-to-end across applications, infra, and cloud environments using observability data
  • Turn incident insights into high-quality postmortems and drive corrective actions with engineering teams
  • Collaborate with software teams to implement permanent fixes and preventive improvements
  • Use data analytics and AI-assisted methods to identify systemic weaknesses and prevent recurrence

Skills

end-to-end system architecture
observability tools
identify patterns
postmortems
AI-assisted analysis
Power BI/Tableau
Python
SQL
DevOps CI/CD
QA or automation background

Education

BS in Computer Science

Tools

T-APM
T-Trace
CatchPoint
Grafana
ServiceNow
Jira
Git

Job description

Join AT&T and reimagine the communications and technologies that connect the world. The Chief Information Office is responsible for advancing information technology performance and delivering solutions with a focus on maximizing ROI, increasing efficiency and enhancing the experience of end users. Guided by experienced leaders, Corporate Systems seamlessly integrate with advanced Technology and Operations to drive our enterprise forward. Our Systems Reliability and Software Delivery teams are unwavering in their commitment to excellence, ensuring every solution is robust and efficient. When you step into a career with AT&T, you won’t just imagine the future-you’ll create it.

What you’ll do:

In this role, you will focus on understanding why production incidents happen and how to prevent them from recurring. You will analyze incidents end-to-end across applications, infrastructure, and cloud environments, using observability data to identify root causes, patterns, and systemic weaknesses.

You will turn incident insights into high-quality postmortems and partner with engineering teams to drive corrective actions and long-term improvements. By combining system-level thinking with data, automation, and AI-assisted analysis, you will help shift the organization from reactive response to proactive reliability and incident prevention. You will partner with engineering and software development teams to implement permanent fix and preventive improvements

What you'll bring:
  • Strong understanding of end-to-end system architecture (cloud, web apps, APIs, databases, infrastructure)
  • Hands‑on experience with observability tools (logs, metrics, traces)
  • Ability to identify patterns and drive preventive actions
  • Experience writing clear, structured postmortems
  • Ability to analyze operational data using tools, queries, or AI-assisted methods
  • Strong systems thinking and problem‑solving skills
  • Background in QA, test engineering, or automation engineering (strong plus)
  • Experience using AI or advanced analytics for incident analysis or pattern detection
  • Understanding of distributed systems and failure modes
  • Experience with data analysis / visualization tools (e.g., Power BI, Tableau)
  • Mindset focused on eliminating recurring issues, not just fixing incidents
  • Strong communication skills to explain complex issues clearly
Required:
  • 7+ years in Systems Engineering, ITSM, RM/CM
  • Background in SRE, Support or QA
  • One or more of the following SRE Tools: T-APM, T-Trace, CatchPoint, Grafana
  • Hands‑on experience and understanding of concepts and tools such as SAFe, Agile, DevOps, CI/CD, Data Analytics, and building Gen AI use cases
  • Experience with AI technologies, Python, SQL, data analytics, Power BI and ITSM tools (e.g., ServiceNow)
  • Modern Enterprise Release Management/Change Management and ITSM
Preferred:
  • BS/BA in Computer Science
  • Preferred tools: modern Release Management processes for Agile and DevOps environments
  • Jira Align, JSM, Jira Cloud, Git for enterprise RM/CM
  • Relevant certifications (SAFe, Agile, DevOps, AI/ML)
Joining our team comes with amazing perks and benefits:
  • 401(k) plan
  • Paid Time Off and Holidays (based on date of hire, at least 23 days of vacation each year and 9 company-designated holidays)
  • Paid Caregiver Leave
  • Additional sick leave beyond what state and local law require may be available but is unprotected
  • Adoption Reimbursement
  • Disability Benefits (short term and long term)
  • Life and Accidental Death Insurance
  • Supplemental benefit programs: critical illness/accident hospital indemnity/group legal
  • Employee Assistance Programs (EAP)
  • Extensive employee wellness programs
  • Employee discounts up to 50% off on eligible AT&T mobility plans and accessories
  • AT&T internet (and fiber where available) and AT&T phone.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Software Reliability Engineering
Lead Software Reliability Engineering

AT&T • Atlanta (GA)

On-site
USD 141,300 - 237,400
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+5
Lead Software Reliability Engineering
Lead Software Reliability Engineering

Cricket Wireless LLC. • Dallas (TX)

On-site
USD 141,000 - 237,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+9
Lead Software Engineer - Reliability Engineering (SRE)
Lead Software Engineer - Reliability Engineering (SRE)

AT&T • Town of Texas (WI)

On-site
USD 140,000 - 190,000
Principal System Engineer (Laravel/PHP Developer, AI Focus)
Principal System Engineer (Laravel/PHP Developer, AI Focus)

AT&T • Plano (TX)

On-site
USD 174,000 - 262,000
401(k) plan
Paid Time Off (23 days of vacation and 9 holidays)
Paid Caregiver Leave
+2
Systems Integration Engineer (Government)
Systems Integration Engineer (Government)

AT&T • Reston (VA)

On-site
USD 108,000 - 190,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement
+3
Systems Engineer (Government)
Systems Engineer (Government)

AT&T • Bellevue (NE)

On-site
USD 86,000 - 145,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+3
Systems Integration Engineer (Government)
Systems Integration Engineer (Government)

Cricket Wireless LLC. • Reston (VA)

On-site
USD 108,000 - 190,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+7
Systems Engineer
Systems Engineer

Southern Arkansas University • United States

On-site
USD 77,800 - 130,800
Tuition reimbursement
Paid Time Off (at least 23 days)
Paid Caregiver Leave
+4
Principal Enterprise Release Manager
Principal Enterprise Release Manager

AT&T • Atlanta (GA)

On-site
USD 117,000 - 196,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+6
Principal System Engineering
Principal System Engineering

AT&T • Plano (TX)

On-site
USD 198,000 - 261,000
Medical/Dental/Vision coverage
401(k) plan
Tuition reimbursement program
+6