Senior SRE (DevOps)

MOZAT PTE LTD

Singapore

On-site

SGD 120,000 - 170,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Competitive compensation
Performance-based bonuses
Growth opportunities
Inclusive environment

Job summary

MOZAT PTE LTD is seeking a Senior Site Reliability Engineer to own the reliability of global gaming and livestreaming systems. You will manage day-to-day operations, implement automation, and partner with engineering teams to maintain high availability and scalable infrastructure.

You will optimize performance, monitor systems, and contribute to disaster recovery planning. The role requires 5-7 years in DevOps/SRE, strong Linux skills, and experience with Docker/Kubernetes and cloud platforms

Qualifications

  • 5-7 years of experience in DevOps, SRE, or infrastructure operations in the internet, gaming, or livestreaming industry.
  • Strong Linux system administration and troubleshooting skills.
  • Proficient in infrastructure scripting (Shell/Python) and automation.
  • Solid experience with production database management (e.g., MySQL/PostgreSQL), including tuning, scaling, and disaster recovery.
  • Familiar with cloud providers such as AWS and AliCloud.
  • Experience building network observability and monitoring systems for global markets.
  • Knowledge of container technologies (Docker, Kubernetes) and CI/CD pipelines.
  • Experience in 24/7 mission-critical environments or on-call duty is a plus.

Responsibilities

  • Manage day-to-day operations, deployment, monitoring, and incident response for global gaming/livestreaming systems.
  • Collaborate with engineering, QA, and product teams to diagnose and resolve production issues, ensuring high service availability (SLA compliance).
  • Analyze system performance and optimize network quality across global regions.
  • Oversee production database health: backups, recovery, tuning, and capacity planning.
  • Implement and maintain monitoring and alerting systems for observability.
  • Automate operational tasks and workflows using Shell or Python.
  • Support capacity planning, cost optimization, and disaster recovery preparedness.
  • Participate in on-call rotation to support 24/7 uptime.

Skills

Linux Admin
Shell/Python scripting
Database tuning
AWS
AliCloud
Docker
Kubernetes
CI/CD
Observability
On-call

Tools

Docker
Kubernetes
MySQL
PostgreSQL

Job description

Senior Site Reliability Engineer (SRE / DevOps)

As a Senior SRE, you will be the backbone of our infrastructure, ensuring our global gaming and livestreaming systems are fast, reliable, and scalable for millions of users. You will manage day-to-day operations, drive automation, and collaborate closely with engineering teams to maintain the highest levels of service availability.

Key Responsibilities
  • Manage day-to-day operations, deployment, monitoring, and incident response for global gaming/livestreaming systems.
  • Collaborate with engineering, QA, and product teams to quickly diagnose and resolve production issues, ensuring high service availability (SLA compliance).
  • Analyze system performance and optimize network quality across global regions.
  • Oversee production database health: conduct routine inspections, manage backups and recovery, optimize slow queries, and plan for capacity.
  • Implement and maintain monitoring and alerting systems to ensure infrastructure observability.
  • Automate operational tasks and workflows using scripting languages such as Shell or Python.
  • Support capacity planning, cost optimization, and disaster recovery preparedness.
  • Participate in an on-call rotation to support 24/7 system uptime as needed.
Requirements
  • 5-7 years of relevant experience in DevOps, SRE, or infrastructure operations in the internet, gaming, or livestreaming industry.
  • Strong Linux system administration and troubleshooting skills.
  • Proficient in infrastructure scripting (Shell/Python) and automation.
  • Solid experience with production database management (e.g., MySQL/PostgreSQL), including tuning, scaling, and disaster recovery.
  • Familiar with global cloud infrastructure providers such as AWS and AliCloud.
  • Experience building network observability and monitoring systems for overseas markets.
  • Working knowledge of container technologies (e.g., Docker, Kubernetes) and CI/CD pipelines.
  • Experience supporting 24/7 mission-critical environments or participating in on-call duty is a strong advantage.
Nice to Have
  • Experience with real-time systems (gaming, live streaming, WebRTC).
  • Familiarity with GPU infrastructure for AI/ML workloads.
  • Certifications in AWS/GCP (Solutions Architect, DevOps Engineer, etc.).
What We Offer
  • Be part of a high-impact global product team targeting emerging markets.
  • Competitive compensation and performance-based bonuses.
  • Opportunities to grow into infrastructure leadership roles.
  • Dynamic, inclusive, and tech-forward working environment.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer(Senior SRE)
Site Reliability Engineer(Senior SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
Technical Leadership
Career Growth
High-Performance Team
+1
Senior Platform Engineer / Site Reliability Engineer (SRE)
Senior Platform Engineer / Site Reliability Engineer (SRE)

AMBITION GROUP SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior SRE: Global Infra, Automation & 24/7 Uptime
Senior SRE: Global Infra, Automation & 24/7 Uptime

MOZAT PTE LTD • Singapore

On-site
SGD 120,000 - 170,000
Competitive compensation
Performance-based bonuses
Growth opportunities
+1
Site Reliability Engineer
Site Reliability Engineer

IDEMIA Public Security • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 200,000
Software Engineer/ Site Reliability Engineer
Software Engineer/ Site Reliability Engineer

United States Digital Space LLC • Singapore

On-site
SGD 90,000 - 150,000
Senior DevOps Engineer
Senior DevOps Engineer

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 80,000 - 120,000
Technical leadership opportunities
Access to mentorship and certifications
Modern DevOps environment
Sr. SRE
Sr. SRE

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000
On-site in Singapore (3 days/wk)
Software Engineer - Engineering Enablement (SRE Focus)
Software Engineer - Engineering Enablement (SRE Focus)

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000