- Lead 24x7 reliability, scalability, and operational excellence for a global portfolio of mission-critical applications and services
- Serve as executive escalation point during major incidents, coordinating rapid resolution and communication with senior stakeholders
- Partner with product engineering teams to embed reliability into product design, development, and release processes
- Define, implement, and improve observability, incident management, disaster recovery, and resilience engineering practices
- Champion automation and AI-driven operations to improve efficiency and reduce service disruption
- Shape and evolve the SRE strategy, operating model, and engineering behaviors across the Corporate Tax and Trade product portfolio
- Lead, mentor, and develop a global team of SRE managers and engineers
- Foster innovation, accountability, inclusion, continuous improvement, and service excellence
Requirements
- 10+ years of experience in site reliability engineering, software engineering, infrastructure, platform engineering, enterprise technology, service management, or a related technology leadership role
- Proven experience leading global teams supporting large-scale distributed systems, cloud-native architectures, and mission-critical enterprise applications
- Strong expertise in reliability engineering practices, observability tools, automation frameworks, incident management, disaster recovery, and operational readiness
- Experience operating within or closely partnering with Service Management functions, including incident, problem, change, availability, and continuity management
- Experience influencing product and technology roadmaps to prioritize reliability, scalability, performance, resilience, and service quality improvements
- Demonstrated ability to lead transformation in a complex enterprise technology environment, including adoption of automation, AI operations, and modern SRE practices
- Strong people leadership skills, including talent development, succession planning, performance management, and resource allocation across regions and technology domains
- Strong financial acumen, including experience managing large budgets and aligning investment decisions to business priorities
- Exceptional communication and stakeholder management skills, with ability to influence senior leaders, engineering teams, and business partners at all levels
Demonstrates extensive expertise in Site Reliability Engineering, focusing on reliability engineering practices, observability, incident management, and disaster recovery. Proven ability to lead global teams, drive automation and AI-driven operations, and influence technology roadmaps to enhance service quality and operational excellence.
Highest-signal resume keywords
- Site Reliability Engineering
- Reliability Engineering Practices
- Observability Tools
- Incident Management
- Automation Frameworks
ATS Optimization Keywords
Hard Skills
- Site Reliability Engineering
- Reliability Engineering Practices
- Observability Tools
- Incident Management
- Disaster Recovery
- Cloud-Native Architectures
- Automation
- AI Operations
- Service Management
- Operational Readiness
Soft Skills
- People Leadership
- Communication
- Stakeholder Management
- Talent DevelopmentPerformance Management
Industry Keywords
- Global Teams
- Large-Scale Distributed Systems
- Enterprise Technology
- Service Management Functions
- Financial Acumen