Site Reliability Engineer(Senior SRE)

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

XIAOMI TECHNOLOGIES SINGAPORE PTE. LTD.

is seeking a seasoned Site Reliability Engineer to maintain high availability of overseas production environments, improve system reliability, and ensure business continuity through proactive capacity planning, automation, and incident response. You will work with R&D, product, security, and infrastructure teams, build observability with Prometheus, Grafana, and ELK, and explore AI tools to boost automation and fault analysis while driving best practices

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering, Information Technology, or related field.
  • 5+ years of experience in SRE, Computer Systems Administration, Platform/Cloud Engineer roles.
  • Proficiency in at least one programming language (Python/Go/Java/C++).
  • Familiarity with cloud services; multi-cloud or hybrid cloud is a plus.
  • Strong Linux, networking, and distributed systems knowledge, plus HA architectures.

Responsibilities

  • Ensure stability, reliability, and high availability of overseas production environment.
  • Manage provisioning, capacity planning, monitoring, change management, incident response, and daily operations.
  • Review system architecture, identify risks, and implement mitigations for stability and performance.
  • Participate in on-call rotations and respond to production incidents.
  • Build and enhance observability with monitoring, logging, and tracing.
  • Develop automation platforms and engineering efficiency tools.
  • Explore AI technologies in operations to improve automation and R&D efficiency.
  • Collaborate with R&D, product, security, and infrastructure teams to drive stability initiatives.

Skills

SRE experience
Python
Go
Java
C++
Cloud platforms
Linux
Networking
Observability
AI in ops
Bash scripting

Education

Bachelor's degree in Computer Science or related field

Tools

Prometheus
Grafana
ELK
CI/CD tooling
Kubernetes
AWS
Azure
GCP

Job description

Job Responsibilities:


  • Ensure the stability, reliability, and high availability of the company’s overseas production environment, continuously improving system availability and service quality.

  • Manage resource provisioning, capacity planning, monitoring, change management, incident response, and daily operations to maintain business continuity and stability.

  • Review system architecture and technical solutions, identify potential risks, and implement mitigations to optimize system stability, performance, and resource efficiency.

  • Participate in on-call rotations, responding promptly to production incidents to safeguard business operations.

  • Build and enhance observability systems, including monitoring, logging, and distributed tracing, to improve monitoring capabilities and fault detection efficiency.

  • Develop and optimize automation platforms and engineering efficiency tools to advance operational automation and team delivery effectiveness.

  • Explore and promote the application of AI technologies in operations scenarios, leveraging AI tools to improve automation, fault analysis, knowledge management, and R&D efficiency.

  • Collaborate closely with R&D, product, security, and infrastructure teams to drive stability initiatives, implement best practices, and support ongoing business development.<br>


Requirements:


  • Bachelor’s degree or above in Computer Science, Software Engineering, Information Technology, or a related field, with 5+ years of experience in Site Reliability Engineering (SRE), Computer Systems Administrator, Platform Engineer, or Cloud Engineer.

  • Proficiency in at least one programming language (e.g., Python, Go, Java, or C++), with strong software development and automation skills.

  • Familiarity with cloud computing services; experience with multi-cloud or hybrid cloud platforms (e.g., Alibaba Cloud, Azure, AWS, GCP) is a plus.

  • Solid understanding of Linux, computer networking, load balancing, distributed systems, and high-availability architectures.

  • Ability to quickly diagnose issues, communicate across teams, and drive solutions—developing system optimization and stability plans aligned with business goals, including dependency management, traffic governance, and disaster recovery planning.

  • Experience with system monitoring and observability tools (e.g., Prometheus, Grafana, ELK, or similar), along with scripting knowledge (Bash or Python) and familiarity with CI/CD concepts is preferred.

  • Familiarity with AI tools and their applications in software development, automated operations, or R&D efficiency—understanding of AI Agents or AIOps technologies; practical experience is a plus.

  • Strong communication, teamwork, and project management skills; ability to adapt to a fast-paced technical environment and continuously learn and apply new technologies.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

IDEMIA Public Security • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 200,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
Technical Leadership
Career Growth
High-Performance Team
+1
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Purview Asia Pacific • Singapore

On-site
SGD 120,000 - 180,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

PURVIEW ASIA PACIFIC PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Software Engineer/ Site Reliability Engineer
Software Engineer/ Site Reliability Engineer

United States Digital Space LLC • Singapore

On-site
SGD 90,000 - 150,000
Site Reliability Engineer
Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 180,000 - 300,000
Site Reliability Engineer
Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 180,000 - 250,000
Senior SRE (DevOps)
Senior SRE (DevOps)

MOZAT PTE LTD • Singapore

On-site
SGD 120,000 - 170,000
Competitive compensation
Performance-based bonuses
Growth opportunities
+1