Get more replies from employers
Send a job-specific resume in minutes.
AT&T in Dallas is seeking a senior Systems Reliability/ITSM professional to prevent recurring production incidents. You will analyze incidents across applications, infrastructure, and cloud environments, using observability data to identify root causes and systemic weaknesses.
You will create high-quality postmortems and lead engineering teams to implement permanent fixes and preventive improvements. Bring hands-on experience with AI-enabled incident analysis, Python, SQL, data analytics, and BI
Join AT&T and reimagine the communications and technologies that connect the world. The Chief Information Office is responsible for advancing information technology performance and delivering solutions with a focus on maximizing ROI, increasing efficiency and enhancing the experience of end users. Guided by experienced leaders, Corporate Systems seamlessly integrate with advanced Technology and Operations to drive our enterprise forward. Our Systems Reliability and Software Delivery teams are unwavering in their commitment to excellence, ensuring every solution is robust and efficient. When you step into a career with AT&T, you won’t just imagine the future-you’ll create it.
In this role, you will focus on understanding why production incidents happen and how to prevent them from recurring. You will analyze incidents end-to-end across applications, infrastructure, and cloud environments, using observability data to identify root causes, patterns, and systemic weaknesses.
You will turn incident insights into high-quality postmortems and partner with engineering teams to drive corrective actions and long-term improvements. By combining system-level thinking with data, automation, and AI-assisted analysis, you will help shift the organization from reactive response to proactive reliability and incident prevention. You will partner with engineering and software development teams to implement permanent fix and preventive improvements