SRE/DevOps Specialist

Remotedxb

Dubai

On-site

AED 300,000 - 520,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

iFood is seeking a Senior SRE/Incident Response engineer to lead critical incident investigations in AWS and Kubernetes. You will develop internal APIs and workers using Go and Python, and contribute to standardization efforts for SLOs/SLIs and incident post-mortems.

You will collaborate with SREs and engineering teams to refine internal tools, prepare for high-impact events, and codify standards via RFCs. Hybrid or on-site in Dubai is expected.

Qualifications

  • Experience investigating and responding to critical incidents in AWS and Kubernetes.
  • Ability to develop internal solutions using Go and Python.
  • Experience with metric standardization (SLOs/SLIs).
  • Ability to lead technical refinements and define standards via RFCs.

Responsibilities

  • Investigate and respond to critical incidents, providing technical leadership in AWS environments, Kubernetes clusters, and network components.
  • Develop internal solutions such as APIs and workers using Go and Python/LangGraph for AI solutions.
  • Participate in metric standardization projects, including the definition of SLOs/SLIs, error budget alerts, and burn-rate.
  • Collaborate with engineering teams and SREs to evolve internal solutions.
  • Prepare applications for high-impact events.
  • Contribute to technical refinement to increase team maturity.
  • Define and govern standards through RFCs/IRCs.
  • Lead deep Post-Incident Reviews (PIRs) and GameDays (Chaos Engineering).

Skills

Go
Python
AWS
Kubernetes
SLOs/SLIs
RFCs

Tools

Istio
Datadog
Litmus

Job description

About the Company

iFood is a leading Brazilian technology company in Latin America. We connect thousands of restaurants to millions of consumers daily through innovative solutions, including food delivery, grocery, pharmacy, and pet markets, as well as our fintech arm, iFood Pago.

Responsibilities
  • Investigate and respond to critical incidents, providing technical leadership in AWS environments, Kubernetes clusters, and network components.
  • Develop internal solutions such as APIs and workers using Go and Python/LangGraph for AI solutions.
  • Participate in metric standardization projects, including the definition of SLOs/SLIs, error budget alerts, and burn-rate.
  • Collaborate with engineering teams and SREs to evolve internal solutions.
  • Prepare applications for high-impact events.
  • Contribute to technical refinement to increase team maturity.
  • Define and govern standards through RFCs/IRCs.
  • Lead deep Post-Incident Reviews (PIRs) and GameDays (Chaos Engineering).
Requirements
  • Experience investigating and responding to critical incidents in AWS and Kubernetes.
  • Ability to develop internal solutions using Go and Python.
  • Experience with metric standardization (SLOs/SLIs).
  • Ability to lead technical refinements and define standards via RFCs.
Preferred Qualifications
  • Proficiency in Go (Golang) and Python.
  • Knowledge of Chaos Engineering tools (e.g., Litmus) and performance testing (e.g., K6).
  • Familiarity with Service Mesh (Istio) and Datadog.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE/DevOps Lead - Incident Response, Kubernetes & Go/Python
SRE/DevOps Lead - Incident Response, Kubernetes & Go/Python

Remotedxb • Dubai

On-site
AED 300,000 - 520,000
Staff Cloud Security Engineer
Staff Cloud Security Engineer

Remotedxb • Dubai

On-site
AED 350,000 - 650,000
Senior DevOps Engineer
Senior DevOps Engineer

Good co India • United Arab Emirates

On-site
AED 240,000 - 420,000
DevOps Engineer
DevOps Engineer

Client of Huzzle • Abu Dhabi

On-site
AED 150,000 - 200,000
Competitive salary package
Opportunity to work with cutting-edge technologies
Strong career progression opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Good co India • United Arab Emirates

On-site
AED 240,000 - 480,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Open Innovation AI • Abu Dhabi Emirate

On-site
AED 150,000 - 210,000
Software Development Engineer
Software Development Engineer

Client of Salt • Dubai

On-site
AED 300,000 - 420,000
Senior Platform Engineer - AWS
Senior Platform Engineer - AWS

Client of Michael Page • United Arab Emirates

On-site
AED 240,000 - 420,000
Cloud Security Architect — AWS & Kubernetes Expert
Cloud Security Architect — AWS & Kubernetes Expert

Remotedxb • Dubai

On-site
AED 350,000 - 650,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Loft Orbital Solutions • Abu Dhabi

On-site
AED 450,000 - 750,000