Principal DevOps & SRE — AI Efficiency Platform Lead

Riot Games

Singapore

On-site

SGD 180,000 - 280,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Full relocation support
Comprehensive health insurance
Open paid time off
Retirement benefits with company match
Life insurance and parental leave

Job summary

Riot Games is seeking a Principal DevOps / Site Reliability Engineer for the AI Efficiency team in Singapore. The role focuses on building and operating scalable AI-enabled platforms, improving observability, and reducing operational toil.

You will collaborate with software engineers, data scientists, and security teams to ensure production readiness and reliability. You will lead incident responses, define SLOs/SLIs, and champion governance for automated operational systems while mentoring

Qualifications

  • Bachelor’s degree in Computer Science or a related field or equivalent professional experience.
  • 5+ years in SRE, DevOps, or equivalent roles supporting production systems.
  • Strong programming and automation skills in Python, Go, JavaScript, or TypeScript.
  • Experience designing and operating cloud-based production systems in AWS, GCP, Azure, or similar.
  • Experience building and maintaining CI/CD pipelines and deployment automation.
  • Strong observability, metrics, logging, tracing, dashboards, and alert design.
  • Experience incident response, on-call support, root-cause analysis, and post-incident improvements.
  • Experience improving reliability, availability, scalability, and performance of distributed systems.
  • Strong knowledge of containers and orchestration (Kubernetes, Docker, etc.).
  • Experience with infrastructure-as-code tools (Terraform, Pulumi, CloudFormation).
  • Linux, networking, DNS, load balancing, authentication, secrets management, and cloud security fundamentals.
  • Ability to drive improvements across teams and collaborate across organizational boundaries.
  • Experience mentoring engineers and establishing engineering standards.

Responsibilities

  • Own and improve reliability, availability, scalability, performance, and health of the AI Efficiency platform and tools.
  • Design, build, and maintain infrastructure, deployment systems, and operational foundations for production AI services.
  • Improve CI/CD pipelines, release engineering, environment management, and deployment automation.
  • Set production-readiness standards with monitoring, ownership, and rollback strategies.
  • Define and operationalize SLOs, SLIs, and reliability metrics to guide priorities.
  • Build observability across applications, infrastructure, and service dependencies.
  • Establish on-call practices, escalation paths, and runbooks for production.
  • Lead diagnosis and resolution of incidents, with blameless post-incident reviews.
  • Build automation to reduce toil and improve MTTR.
  • Implement safe deployment patterns: canary, feature flags, health checks.
  • Perform capacity planning, load testing, and performance analysis for AI services.
  • Design backup, recovery, failover, and disaster-recovery strategies.
  • Identify single points of failure and systemic risks across systems.
  • Improve developer experience with self-service workflows and reusable infrastructure.
  • Maintain infrastructure-as-code, secrets-management, and environment governance.

Skills

Programming languages
Cloud platforms
CI/CD pipelines
Observability
Incident response
Container orchestration
Infrastructure as code
Security fundamentals
Cross-team collaboration

Education

Bachelor’s degree in Computer Science or related field

Tools

Terraform
Pulumi
CloudFormation
Kubernetes
Docker
ECS

Job description

Riot Games is seeking a Principal DevOps / Site Reliability Engineer for the AI Efficiency team in Singapore. The role focuses on building and operating scalable AI-enabled platforms, improving observability, and reducing operational toil.

You will collaborate with software engineers, data scientists, and security teams to ensure production readiness and reliability. You will lead incident responses, define SLOs/SLIs, and champion governance for automated operational systems while mentoring

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Engineer, AI Efficiency & Reliability
Senior Platform Engineer, AI Efficiency & Reliability

Riot Games • Singapore

On-site
SGD 100,000 - 140,000
Full relocation support
Comprehensive health insurance
Open paid time off
+4
Senior Site Reliability Engineer: AI-Powered Scale (Hybrid)
Senior Site Reliability Engineer: AI-Powered Scale (Hybrid)

United States Digital Space LLC • Singapore

Hybrid
SGD 120,000 - 190,000
ESOP
High-Impact Environment
Health & Retirement Benefits
+3
Senior Site Reliability Engineer — AI-Driven Resilience
Senior Site Reliability Engineer — AI-Driven Resilience

Singapore Exchange Limited • Singapore

On-site
SGD 180,000 - 250,000
Principal Software Engineer - DevOps / Site Reliability Engineer
Principal Software Engineer - DevOps / Site Reliability Engineer

Riot Games • Singapore

On-site
SGD 180,000 - 280,000
Full relocation support
Comprehensive health insurance
Open paid time off
+2
VP, Site Reliability Engineering — AI-Driven Platform & Reliability
VP, Site Reliability Engineering — AI-Driven Platform & Reliability

Singapore Exchange Limited • Singapore

On-site
SGD 120,000 - 160,000
Principal Software Engineer - ML Platform Engineer
Principal Software Engineer - ML Platform Engineer

Riot Games • Singapore

On-site
SGD 100,000 - 140,000
Full relocation support
Comprehensive health insurance
Open paid time off
+4
Senior SRE: AI-Driven Reliability & Automation
Senior SRE: AI-Driven Reliability & Automation

SGX Group • Singapore

On-site
SGD 180,000 - 300,000
Vice President SRE — AI-Driven Platform & Automation
Vice President SRE — AI-Driven Platform & Automation

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000
Senior LLMOps Engineer: Production AI Reliability Lead
Senior LLMOps Engineer: Production AI Reliability Lead

NCS Pte Ltd • Singapore

On-site
SGD 120,000 - 180,000
Health insurance
Learning & development
Senior AI Deployment & Infra Engineer
Senior AI Deployment & Infra Engineer

United States Digital Space LLC • Singapore

On-site
SGD 90,000 - 150,000