Sr Engineer, Software [T500-28132]

TMUS Global Solutions

Hyderabad

Hybrid

INR 1,800,000 - 3,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

TMUS Global Solutions in Hyderabad, India, is hiring a Senior Software Engineer – Resiliency to design and build complex failover and observability features for T-Mobile's production systems across multiple datacenters.

You will own end-to-end workstreams, collaborate with application owners and platform engineers, and mentor junior teammates while upholding strong documentation and runbooks in an async-first, cross‑time‑zone environment.

Qualifications

  • 5+ years of hands‑on software engineering experience across multiple technologies, languages, and system layers.
  • Strong first‑principles understanding of distributed systems, fault tolerance, and failure modes.
  • Demonstrated ability to take ownership of complex features and deliver them end‑to‑end with minimal hand‑holding.
  • AWS and Kubernetes experience at production scale.
  • Familiarity with secret management solutions (CyberArk, Vault).
  • Accountability mindset — you own problems end‑to‑end and you elevate with context and a path forward.
  • Strong documentation skills — ability to translate complex systems into clear, actionable guides.

Responsibilities

  • Design and build complex resiliency features and automation solutions that protect T-Mobile’s most critical systems across multiple datacenters
  • Take full ownership of workstreams end-to-end — from design through delivery — within the broader architecture and technical direction set by the team
  • Work across multiple technologies and applications, operating as a trusted technical partner to application owners and platform engineers
  • Contribute meaningfully to technical design discussions — bringing well‑reasoned proposals, raising risks, and helping the team make better decisions
  • Mentor junior and mid-level engineers through code reviews and hands‑on guidance
  • Proactively engage application owners and drive conversations to unblock delivery
  • Design and implement observability solutions — build monitoring dashboards, alerting, and health‑check mechanisms to provide real‑time visibility into failover readiness and execution
  • Recommend and contribute to best practices — evaluate current processes, identify gaps, and propose improvements for failover patterns, automation standards, and operational runbooks
  • Document everything — create clear, comprehensive technical documentation, architecture diagrams, runbooks, and onboarding guides that enable team scalability and knowledge transfer

Skills

AWS
Kubernetes
Distributed systems
Ownership
Documentation
First-principles thinking
Observability
Secret management
Ansible

Tools

Ansible

Job description

T-Mobile US, Inc. (NASDAQ: TMUS), headquartered in Bellevue, Washington, is America’s supercharged Un-carrier, connecting millions through its strong nationwide network and flagship brands, T-Mobile and Metro by T-Mobile. Customers benefit from an unmatched combination of value, quality, and exceptional service experience.


TMUS Global Solutions:

TMUS Global Solutions is a world-class technology powerhouse accelerating the company’s global digital transformation. With a culture built on growth, inclusivity, and global collaboration, the teams here drive innovation at scale, powered by bold thinking.


Senior Software Engineer – Resiliency

About the Role:


  • T-Mobile runs some of the most transaction-intensive systems in U.S. telecommunications — millions of payments, device activations, and customer interactions processed every day. When those systems fail, customers feel it immediately. Your job is to make sure they don’t.

  • We’re building the next generation of resiliency solutions including automated failover, cross-datacenter orchestration, observability pipelines, and AI Ops. This is hands‑on, high-ownership engineering work with real consequences at real scale. As a Senior Engineer, you’ll drive complex feature development, take ownership of key workstreams, and bring the technical depth needed to earn trust across application teams, DBAs, network engineers, and platform architects alike.

  • If you thrive in ambiguity, move fast, and hold yourself to a high bar — this role was built for you.


A Few Things Worth Knowing:


  • This team operates in an async-first model with regular sync touchpoints across U.S. and India time zones. You’ll have real ownership of features and workstreams — not just task execution. The systems you work on are production‑critical, and the team holds itself to high standards for reliability, documentation, and operational discipline.

  • If you’re looking for a role where you’ll be handed clean requirements and a clear path, this probably isn’t it. If you want to do work that matters, build things that run at scale, and grow fast in a high-trust environment — we’d like to talk.

  • We pride ourselves on encouraging a culture of innovation, agile ways of working, and transparency in all we do. Join us in embodying the spirit of the Un-carrier and make a tangible impact!


What You’ll Do:


  • Design and build complex resiliency features and automation solutions that protect T-Mobile’s most critical systems across multiple datacenters

  • Take full ownership of workstreams end-to-end — from design through delivery — within the broader architecture and technical direction set by the team

  • Work across multiple technologies and applications, operating as a trusted technical partner to application owners and platform engineers

  • Contribute meaningfully to technical design discussions — bringing well‑reasoned proposals, raising risks, and helping the team make better decisions

  • Mentor junior and mid-level engineers through code reviews and hands‑on guidance

  • Proactively engage application owners and drive conversations to unblock delivery

  • Design and implement observability solutions — build monitoring dashboards, alerting, and health‑check mechanisms to provide real‑time visibility into failover readiness and execution

  • Recommend and contribute to best practices — evaluate current processes, identify gaps, and propose improvements for failover patterns, automation standards, and operational runbooks

  • Document everything — create clear, comprehensive technical documentation, architecture diagrams, runbooks, and onboarding guides that enable team scalability and knowledge transfer


What You’ll Bring:

Must Have:


  • 5+ years of hands‑on software engineering experience across multiple technologies, languages, and system layers

  • Strong first‑principles understanding of distributed systems, fault tolerance, and failure modes — not just framework familiarity, but genuine depth earned through production experience

  • Demonstrated ability to take ownership of complex features and deliver them end‑to‑end with minimal hand‑holding

  • AWS and Kubernetes experience at production scale

  • Familiarity with secret management solutions (CyberArk, Vault)

  • Accountability mindset — you own problems end‑to‑end, you don’t wait to be unblocked, and you elevate with context and a proposed path forward

  • Strong documentation skills — ability to translate complex systems into clear, actionable guides


Who You Are:


  • Self‑driven — You take ownership, find answers yourself, and don’t wait to be told what to do next

  • First‑principles thinker — When something breaks in an unfamiliar system, you reason from fundamentals. You don’t just apply patterns — you understand why the pattern exists

  • Collaborative contributor — You actively participate in technical design conversations, bring well‑reasoned ideas, and work constructively within the direction the team has set

  • Fast learner — You ramp quickly on new tools and ecosystems with minimal guidance

  • Independent operator — You can engage app teams directly, extract what you need, and fill gaps through your own research

  • Fast, iterative, and comfortable with ambiguity — You ship something workable quickly, learn from it, and improve. You don’t need the perfect spec to start

  • Relationship builder — You build trust with stakeholders and drive conversations forward

  • Communicator with standards — You write clearly, document proactively, and treat your teammates’ time as valuable

  • Continuous improver — You don’t just execute — you identify what’s suboptimal and propose better ways of doing things, then follow through

  • Knowledge sharer — You believe documentation is a first‑class deliverable, not an afterthought


Nice‑to‑Have:


  • Ansible and failover engineering experience

  • Experience with observability platforms (Splunk, Grafana, Prometheus, OTEL)

  • Experience using AI coding tools (Claude, GitHub Copilot, ChatGPT) as a genuine productivity multiplier — not just having tried them, but having integrated them into your workflow

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Engineer, Software [T500-28131]
Engineer, Software [T500-28131]

TMUS Global Solutions • Hyderabad

On-site
INR 1,500,000 - 2,300,000
Principal Engineer - Cloud Security DevOps [T500-27143]
Principal Engineer - Cloud Security DevOps [T500-27143]

TMUS Global Solutions • Hyderabad

On-site
INR 4,000,000 - 9,000,000
Sr Engineer, Software - Java [T500-20426]
Sr Engineer, Software - Java [T500-20426]

TMUS Global Solutions • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Manager, Software Engineering [T500-28653]
Manager, Software Engineering [T500-28653]

TMUS Global Solutions • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Principal Engineer, Software - Cloud Network [T500-27161]
Principal Engineer, Software - Cloud Network [T500-27161]

TMUS Global Solutions • Hyderabad

On-site
INR 350,000 - 700,000
Sr Engineer, Software - Python Backend [T500-28120]
Sr Engineer, Software - Python Backend [T500-28120]

TMUS Global Solutions • Hyderabad

On-site
INR 350,000 - 550,000
Engineer, Site Reliability - Accounting Technology [T500-20169]
Engineer, Site Reliability - Accounting Technology [T500-20169]

ANSR • Hyderabad

On-site
INR 60,000 - 80,000
Sr. Software Engineer, Full Stack – CDP
Sr. Software Engineer, Full Stack – CDP

T-Mobile • Hyderabad

On-site
INR 4,000,000 - 6,000,000
Engineer, Software - Java [T500-20431]
Engineer, Software - Java [T500-20431]

TMUS Global Solutions • Hyderabad

On-site
INR 800,000 - 1,200,000
Principal Engineer, Software - Java Backend [T500-28298]
Principal Engineer, Software - Java Backend [T500-28298]

TMUS Global Solutions • Hyderabad

On-site
INR 3,500,000 - 5,500,000