Senior DevOps Engineer

Mirai, a Scopely company

Riyadh

On-site

SAR 420,000 - 640,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Mirai, a Scopely company, is building the infra that powers our Generative AI products in Riyadh. This role owns infrastructure as code, CI/CD, observability, and security on AWS, setting the operational standard for a growing team.

You will run GPU-enabled AI workloads, manage container platforms (EKS/ECS/Fargate), API gateways, and secrets, plus build monitoring and incident response into every service.

Strong Terraform/OpenTofu, Ansible, Kubernetes, and GitHub Actions experience are required.

Qualifications

  • Eight or more years in DevOps, SRE, or infrastructure engineering with production AI/ML workloads.
  • Terraform or OpenTofu with module design and remote state.
  • Terraform Cloud (HCP) experience is a plus.
  • Ansible for configuration management.
  • AWS experience across compute, networking, IAM, and storage.
  • GitHub Actions CI/CD with reusable workflows and credentials handling.
  • Docker/Kubernetes (EKS) and a registry (ECR).
  • API gateways experience (Kong or AWS API Gateway) with auth and routing.
  • RDS/Aurora, PostgreSQL, and cache (Redis/ElastiCache) operations.
  • HashiCorp Vault or AWS Secrets Manager for secrets.
  • Observability tools: Prometheus, Grafana, CloudWatch, OpenTelemetry.
  • Linux scripting with Bash and Python.

Responsibilities

  • Write and maintain infrastructure as code for reproducible environments.
  • Own CI/CD pipelines for building, testing, scanning, and deploying applications and model-serving services.
  • Run the container platform (EKS, ECS, Fargate) and GitOps workflows.
  • Set up runtime for AI workloads: GPU capacity and model serving (vLLM, Triton, or similar).
  • Manage API gateways, networking, load balancing, DNS, and certificates.
  • Own secrets, identity, and least-privilege access across environments.
  • Operate production databases: clustering, replication, backups, failover, recovery.
  • Build monitoring for token usage, GPU utilization, and service level objectives.
  • Lead reliability and security: incident response, policy as code, scans, and cost discipline.

Skills

DevOps & SRE
AWS
CI/CD
Observability
Security & IAM
AI workloads

Tools

Terraform/OpenTofu
Ansible
Kubernetes (EKS)
GitHub Actions
Vault / Secrets Manager
Prometheus/Grafana
RDS/Aurora
vLLM/Triton model serving

Job description

This role builds and runs the infrastructure our Generative AI products depend on: the pipelines that ship code, the platforms that run services and models, and the controls that keep all of it secure and reliable. AI workloads bring their own demands. GPUs, model serving, inference autoscaling, and token cost all shape the work, and you have run workloads like these before. You should be comfortable owning infrastructure as code, CI/CD, observability, and security on AWS, and ready to set the operational standards a growing team will lean on.

What You Will Do
  • Write and maintain infrastructure as code so environments are reproducible, reviewable, and quick to recover
  • Own CI/CD: the pipelines that build, test, scan, and deploy applications, agents, and model-serving services
  • Run the container platform (EKS, ECS, or Fargate) and the deployment workflows on top of it, including GitOps where it fits
  • Stand up the runtime for AI workloads: GPU capacity, model serving such as vLLM, Triton, or TGI, inference autoscaling, and the gateways and caching that sit in front of the models
  • Manage API gateways, networking, load balancing, DNS, and certificates so services are exposed safely and predictably
  • Own secrets, identity, and least-privilege access across every environment
  • Run databases in production: clustering, replication, failover, backups, and recovery
  • Build monitoring into everything, including token usage and GPU utilisation, with alerting and clear service objectives
  • Lead reliability and security practice: incident response, policy as code, vulnerability and container scanning, and cost discipline, which matters once GPUs are in the mix
Requirements
  • Eight or more years in DevOps, SRE, or infrastructure engineering overall. That includes hands-on experience supporting AI or ML workloads in production, which can be a more recent part of your backgroun.
  • Strong infrastructure as code with Terraform or OpenTofu, including module design and remote state.
  • Strong infrastructure as code with Terraform or OpenTofu, including module design and remote state. Experience with HCP Terraform (formerly Terraform Cloud) is a plus.
  • Configuration management with Ansible
  • Solid AWS experience across compute, networking (VPC, subnets, security groups, load balancers, Route 53), IAM, and storage
  • Strong CI/CD with GitHub Actions, including reusable workflows and careful handling of credentials
  • Containers and orchestration: Docker with Kubernetes (EKS preferred), Helm, and a registry such as ECR
  • API gateway experience with Kong or Amazon API Gateway, including auth, rate limiting, and routing
  • Database operations including clustering and high availability, with RDS or Aurora, PostgreSQL, and a cache such as Redis or ElastiCache
  • Secrets management with HashiCorp Vault, AWS Secrets Manager, or Parameter Store
  • Observability with Prometheus, Grafana, CloudWatch, and OpenTelemetry, or close equivalents
  • Comfort in Linux and scripting with Bash and Python
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior DevOps Engineer
Senior DevOps Engineer

Mirai Arabian International Company Limited • Riyadh

On-site
SAR 180,000 - 230,000
AI Engineer
AI Engineer

Latitude • Riyadh

On-site
SAR 420,000 - 660,000
AI Engineer
AI Engineer

Saudi Azm عزم السعودية • Riyadh

On-site
SAR 260,000 - 460,000
Senior Data Engineer
Senior Data Engineer

Mirai Arabian International Company Limited • Riyadh

On-site
SAR 120,000 - 150,000
MLOps & Devops Engineers
MLOps & Devops Engineers

Devoteam • Riyadh

On-site
SAR 190,000 - 290,000
GenOps Engineer
GenOps Engineer

atmaal • Riyadh

On-site
SAR 180,000 - 300,000
Sr. AI Engineer (GCP)
Sr. AI Engineer (GCP)

Total-TECH Co. • Jeddah

On-site
MLOps & Devops Engineers
MLOps & Devops Engineers

Devoteam Middle East • Riyadh

On-site
SAR 240,000 - 420,000
Senior AI Infra Engineer: GPUs, CI/CD & Cloud Ops
Senior AI Infra Engineer: GPUs, CI/CD & Cloud Ops

Mirai, a Scopely company • Riyadh

On-site
SAR 420,000 - 640,000
AI/ML Associate Manager
AI/ML Associate Manager

Accenture Middle East • Riyadh

On-site
SAR 360,000 - 600,000