A complete application in a minute — tailored resume and cover letter, ready to send.
TMUS Global Solutions, as part of T-Mobile US, seeks a Senior Site Reliability Engineer to strengthen hybrid infrastructure reliability and automate operations. You will drive incident response, build self-healing automation, and set SRE standards with measurable SLIs/SLOs.
You will mentor engineers, define observability maturity, and partner with cross-functional teams to elevate performance and resilience across multi-cloud environments.
T-Mobile US, Inc. (NASDAQ: TMUS), headquartered in Bellevue, Washington, is America’s supercharged Un-carrier, connecting millions through its strong nationwide network and flagship brands, T-Mobile and Metro by T-Mobile. Customers benefit from an unmatched combination of value, quality, and exceptional service experience.
TMUS Global Solutions is a world-class technology powerhouse accelerating the company’s global digital transformation. With a culture built on growth, inclusivity, and global collaboration, the teams here drive innovation at scale, powered by bold thinking.
TMUS India Private Limited operates as TMUS Global Solutions.
This role ensures the reliability and resilience of digital infrastructure to support efficient software development and deployment. It involves automating processes and reducing manual effort to prevent operational incidents and improve system performance. The role requires expertise in programming, scripting, incident response management, and various technical tools to maintain system robustness. Success is measured by system stability, incident reduction, and continuous improvement in operational efficiency. The work directly impacts organizational stability and customer experience by maintaining high-performing and reliable systems.
The Sr Engineer, Site Reliability is the core operations engineer, capable of resolving complex incidents, improving automation, and mentoring Engineer(s). They bridge operations and engineering by identifying recurring issues and creating scalable fixes.
Incident Command & Complex Troubleshooting:
Automation Framework Design (Infra & Ops):
Observability Strategy & Advanced Monitoring:
Database & Application Performance Engineering:
Cross-Domain SME Knowledge (Networking, Storage, APIs):