Enable job alerts via email!

Site Reliability Engineer

Etihad Airways

Abu Dhabi

On-site

AED 200,000 - 300,000

Full time

5 days ago

Be an early applicant

Generate a tailored resume in minutes

Land an interview and earn more. Learn more

Start fresh or import an existing resume

Job summary

Etihad Airways is seeking a Site Reliability Engineer to lead an SRE squad focused on enhancing service reliability, performance, and scalability. This role involves automation, incident management, and optimizing infrastructure while ensuring compliance with IT governance standards. Ideal candidates will have extensive experience in distributed systems and software development, as well as strong communication and analytical skills.

Qualifications

7+ years experience in software development and 3+ years in DevOps or SRE role.
Ability to design and troubleshoot large-scale distributed systems.
Strong communication and analytical skills.

Responsibilities

Lead an SRE squad to enhance service reliability and performance.
Manage incident resolution and develop response playbooks.
Build monitoring systems and ensure alignment with SLAs.

Skills

Problem-solving

Communication

Analytical skills

Education

Bachelor's degree in Computer Science or related field

Press Tab to Move to Skip to Content Link

The Site Reliability Engineer (SRE) will lead an SRE squad focused on enhancing service reliability, performance, and scalability. They will drive automation to reduce toil, optimize system uptime, and manage incident resolution efforts. Responsible for building monitoring systems, optimizing infrastructure, and implementing safe deployment practices, the SRE will also ensure alignment with SLAs / SLOs and contribute to system development and code reviews. The role requires expertise in large-scale distributed systems, cloud infrastructure, and IT governance, with a focus on continuous service improvement and operational excellence.

Accountabilities

Team Leadership & Reporting : Lead an SRE squad handling operations and automation; represent team in senior management briefings; produce dashboards and progress reports.
Toil Reduction & Automation : Identify and eliminate toil through automation of repetitive tasks, enhancing team efficiency and service reliability.
Service Reliability & Uptime : Maintain and improve service availability by aligning with SLAs / SLOs, designing failover strategies, and hardening systems.
Performance & Latency Optimization : Enhance service performance and reduce latency using profiling tools, distributed tracing, load testing, and bottleneck analysis.
Change & Deployment Management : Implement safe deployment practices (e.g., canary releases, blue-green deployments), ensuring minimal risk and rapid rollback options
Monitoring & Observability : Build and manage real-time monitoring and alerting systems to ensure service health and proactively detect anomalies.
Incident Management & RCA : Lead incident resolution efforts, conduct root cause analyses (RCA), and develop response playbooks to reduce MTTR.
Capacity & Cost Optimization : Perform infrastructure capacity planning and cost-efficient scaling to meet service demands.
Development & Code Review : Contribute to system development, participate in design / code reviews, and ensure alignment with engineering best practices.
Governance, Compliance & Documentation : Enforce IT governance standards, maintain documentation, perform quality assessments, and contribute to architecture and risk committees.

Education & Experience

7+ years of experience with data structures / algorithms and software development in Two or more programming languages and operating and maintaining platforms with 3+ years of experience in a DevOps or SRE role.
Experience working in computing, distributed systems, storage, or networking.
Expertise in designing, analysing, and troubleshooting large-scale distributed systems.
Ability to debug, optimize code, and to automate routine tasks.
Systematic problem-solving approach, coupled with effective verbal and written communication skills.
Strong communication capability, able to articulate technical issues in terms of business risk and opportunity.
Knowledge of the technical aspects of cloud computing, data centres, networks and virtual infrastructure.
Strong analytical and problem-solving skills are necessary , TSM processes & tools

Etihad Airways, the national airline of the UAE, was formed in 2003 and quickly went on to become one of the world’s leading airlines. From its home in Abu Dhabi, Etihad flies to passenger and cargo destinations in the Middle East, Africa, Europe, Asia, Australia and North America. Together with Etihad’s codeshare partners, Etihad’s network offers access to hundreds of international destinations. In recent years, Etihad has received numerous awards for its superior service and products, cargo offering, loyalty programme and more.All this ties into Etihad’s ambitious Journey 2030 strategy. The airline plans to double its fleet size and triple the number of customers over the next six years as it sets out to be the airline everyone wants to fly!

Beware of fraudulent job offers from individuals or organizations claiming to represent the Etihad group. We will never ask for personal information, bank details, or payment during the recruitment process. Interviews are conducted face-to-face or via video / telephone before any formal offer. If you are asked for money, please treat it as fraudulent.

J-18808-Ljbffr

Site Engineer • Abu Dhabi, Abu Dhabi Emirate, United Arab Emirates

Get your free, confidential resume review.

or drag and drop a PDF, DOC, DOCX, ODT, or PAGES file up to 5MB.

Site Reliability Engineer

Etihad Airways

Abu Dhabi

On-site

AED 200,000 - 300,000