Site Reliability Engineer (SRE) - Azure focus

Dicetek LLC

Dubai

On-site

AED 300,000 - 550,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Dicetek LLC in Dubai seeks an experienced Site Reliability Engineer to ensure reliability, availability, performance and scalability of critical applications and infrastructure. You will monitor systems, design observability, automate deployments, and lead incident response, collaborating with DevOps and development teams.

The role requires hands-on experience with Kubernetes, cloud platforms, IaC and CI/CD pipelines; strong troubleshooting skills and ability to operate in large-scale

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or Production Support Engineer.
  • Strong experience in cloud infrastructure, automation, monitoring, and incident management.
  • Hands-on experience with Kubernetes and containerized environments.
  • Strong troubleshooting and root cause analysis skills.
  • Experience working in highly available, large-scale production environments.
  • Excellent communication and collaboration skills.

Responsibilities

  • Monitor and maintain the availability performance and reliability of applications and infrastructure
  • Design and implement effective monitoring logging and alerting solutions
  • Manage and troubleshoot production incidents and perform root cause analysis
  • Automate operational and deployment processes to improve system reliability and efficiency
  • Work closely with development infrastructure and DevOps teams to improve application performance and resilience
  • Implement and maintain CI CD pipelines and Infrastructure as Code IaC
  • Support containerized environments and cloud-based infrastructure
  • Develop scripts and automation tools to reduce manual operational activities
  • Implement observability solutions including monitoring logging and distributed tracing
  • Ensure proper documentation of operational procedures incidents and system configurations

Skills

Monitoring
Automation
Cloud technologies
Incident management
DevOps practices
CI/CD pipelines
Communication

Education

Bachelor's degree in Computer Science, Information Technology, or a related field

Tools

Prometheus
Grafana
ELK/Elastic Stack
Splunk
Dynatrace
AppDynamics
New Relic
AWS
Azure
GCP
Docker
Kubernetes
OpenShift
Jenkins
GitLab CI/CD
Terraform
Ansible
Git
ServiceNow
PagerDuty
OpenTelemetry
Jaeger
Bash
Python
PowerShell
PostgreSQL
Oracle
SQL Server

Job description

We are seeking an experienced strong Site Reliability Engineer SRE strong to ensure the reliability availability performance and scalability of critical applications and infrastructure The ideal candidate will have strong experience in monitoring automation cloud technologies incident management and DevOps practices

Key Responsibilities
  • Monitor and maintain the availability performance and reliability of applications and infrastructure
  • Design and implement effective monitoring logging and alerting solutions
  • Manage and troubleshoot production incidents and perform root cause analysis
  • Automate operational and deployment processes to improve system reliability and efficiency
  • Work closely with development infrastructure and DevOps teams to improve application performance and resilience
  • Implement and maintain CI CD pipelines and Infrastructure as Code IaC
  • Support containerized environments and cloud-based infrastructure
  • Develop scripts and automation tools to reduce manual operational activities
  • Implement observability solutions including monitoring logging and distributed tracing
  • Ensure proper documentation of operational procedures incidents and system configurations
Technology Stack
  • Monitoring: Prometheus, Grafana, Zabbix
  • Logging: ELK/Elastic Stack, Splunk
  • APM: Dynatrace, AppDynamics, New Relic
  • Cloud: AWS, Microsoft Azure, or GCP
  • Containers: Docker, Kubernetes, OpenShift
  • CI/CD: Jenkins, GitLab CI/CD, Azure DevOps
  • Infrastructure as Code: Terraform, Ansible
  • Version Control: Git, GitHub, GitLab
  • Incident Management: ServiceNow, PagerDuty
  • Distributed Tracing: OpenTelemetry, Jaeger
  • Scripting: Bash, Python, PowerShell
  • Databases: PostgreSQL, Oracle, SQL Server
  • Web & APIs: IIS, Nginx, Apache, REST APIs
Qualifications & Experience
  • Bachelor's degree in Computer Science, Information Technology, or a related field.
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, or Production Support Engineer.
  • Strong experience in cloud infrastructure, automation, monitoring, and incident management.
  • Hands-on experience with Kubernetes and containerized environments.
  • Strong troubleshooting and root cause analysis skills.
  • Experience working in highly available, large-scale production environments.
  • Excellent communication and collaboration skills.

Skillscloud infrastructuremonitoringincident managemen

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Ras Al Khaimah

On-site
AED 469,000 - 603,000
Senior DevOps / Site Reliability Engineer (SRE)
Senior DevOps / Site Reliability Engineer (SRE)

Stellar Technologies • Abu Dhabi

On-site
AED 360,000 - 540,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Abu Dhabi

On-site
AED 180,000 - 250,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Epergne Solutions • Dubai

On-site
AED 200,000 - 300,000
Site Reliability Engineer (SRE) - Azure AI
Site Reliability Engineer (SRE) - Azure AI

Dicetek LLC • Abu Dhabi

On-site
AED 260,000 - 420,000
Head of Site Reliability Engineering (SRE)
Head of Site Reliability Engineering (SRE)

Client of Mark Williams • Dubai

On-site
AED 600,000 - 1,200,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

31 CONCEPT • United Arab Emirates

On-site
AED 300,000 - 460,000
Platform Site Reliability Engineer
Platform Site Reliability Engineer

Dicetek LLC • Abu Dhabi

Hybrid
AED 120,000 - 180,000
Site Reliability Engineering Manager
Site Reliability Engineering Manager

Open Innovation AI • Abu Dhabi Emirate

On-site
AED 150,000 - 210,000
Lead Site Reliability Engineer at HCLTech
Lead Site Reliability Engineer at HCLTech

HCLTech • United Arab Emirates

On-site
AED 120,000 - 160,000