Job Title: Principal Service Reliability Engineer
The Principal Site Reliability Engineer will be responsible for ensuring the reliability, performance and scalability of our mission‑critical platforms. In this role, you will be safeguarding operational excellence in the assigned product, influence reliability strategies, integral in production incident response, and helping to improve operational metrics. The role will be collaborating closely with different teams, such as Development and Production Support teams, to make sure target SLOs are met, making adjustments where needed and designing/developing code to facilitate meeting such targets. The role is also expected to work on toil reduction projects, handle capacity planning/tuning activities, revisiting existing SOPs and designing/developing code for performance improvements. This is a hybrid position and would require you to be in the local office 2‑3 days a week.
Responsibilities
- Define and track Service Level Indicators (SLIs), Objectives (SLOs), and Error Budgets in partnership with engineering and product leads
- Collaborate with Operations and Development teams to drive service reliability, availability, and scalability
- Drive and participate in toil reduction projects to minimize if not eliminate recurring manual activities performed by the team
- Establish feedback loop with development teams for them to have visibility on the how stable and reliable their services are in client environments
- Drive production incident response and lead root cause analysis and continuous improvement
- Design/Develop operational improvement items with development teams working with them closely in prioritizing these improvements
- Provide input on process improvements to Change, Release, and Incident Management
- Create and implement support playbooks that resources can use as part of emergency response to production issues
Ideal Candidate
- Knowledgeable and experienced in utilizing different Azure resources such as VMs, Storage, Network, Functions, Logic Apps, App Services, AKS
- Strong technical expertise on Azure DevOps, developing in git and working on gitops repo and build/release pipelines
- Hands‑on experience in developing Azure PowerShell scripts, Azure Runbooks, or any other infrastructure automation tools
- Experienced with monitoring and logging tools (Grafana, Dynatrace, Splunk)
- Proven ability to adapt to emerging cloud technologies and industry leading DevOps applications such as Terraform, Docker Containers, and Kubernetes
- Knowledgeable in cloud implementation of Navitaire products across different cloud infrastructure models
- Understanding of production environments and processes and ways on how they can be further optimized through various Azure features and other cloud technologies/services
- Proven ability to drive problem‑solving efforts through effective issue analysis
- Ability to lead efforts to implement infrastructure changes to increase environment stability and support scalability
- Ability to drive collaborations with different Navitaire teams in enforcing environment standards and policies
- Effectively works in a team environment and contributes in building capabilities of team members
- Proficient in C#
- Proven ability to work in a dynamic, fast‑paced and multi‑cultural environment
- Willing to work on shifting schedules
Amadeus is an equal‑opportunity employer. All qualified applicants will receive consideration for employment without regard to gender, race, ethnicity, sexual orientation, age, beliefs, disability or any other characteristics protected by law.