Senior Site Reliability Engineer

landi international (singapore) pte. ltd.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

LANDI Global is seeking a Senior Site Reliability Engineer to drive reliability across our platform infrastructures. You will define standards, lead incident response, and push automation across cloud and on‑prem environments.

You will collaborate with R&D and platform teams to ensure availability, resilience, and operational excellence while mentoring junior engineers and shaping SRE practices across the organization.

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering or a related field.
  • Minimum 7 years of experience as a Site Reliability Engineer, DevOps Engineer, or in a similar role.
  • Strong verbal and written communication skills in English and Mandarin.
  • Strong experience designing and operating distributed systems at scale.
  • Proven ability to improve reliability across multiple services or platforms.
  • Deep understanding of system failure modes, scalability, and performance trade-offs.
  • Experience defining and implementing SLOs, SLIs, and observability practices.
  • Ability to lead incident response and drive systemic improvements.
  • Strong communication skills with the ability to influence without authority.
  • Demonstrated mentorship and technical leadership experience.

Responsibilities

  • Design, build, and optimize LANDI Global’s platform infrastructures across development, staging, and production environments, with a focus on scalability and resilience.
  • Collaborate with R&D and platform teams to define architecture patterns and reliability standards that ensure availability and operational excellence.
  • Lead platform readiness for new client onboarding, ensuring scalability, repeatability, and operational sustainability.
  • Define and drive improvements in monitoring, logging, and alerting systems to ensure high signal quality and proactive issue detection.
  • Lead incident response for high severity events, and drive high-quality root cause analysis (RCA) with a focus on systemic improvements.
  • Design, evolve, and validate Disaster Recovery (DR) and business continuity strategies, ensuring systems meet recovery objectives.
  • Participate in and help evolve the 24/7 standby model to improve operational effectiveness and sustainability.
  • Analyze platform performance metrics and lead optimization strategies across cloud and on-prem environments.
  • Drive improvements in automated testing, CI/CD pipelines, and deployment workflows to enhance release safety, speed, and reliability.
  • Identify and eliminate operational toil through automation and engineering solutions.
  • Establish and standardize operational runbooks and procedures across services.
  • Provide advanced troubleshooting and support for complex production issues, guiding teams toward effective resolution.
  • Lead continuous improvement initiatives to enhance platform resilience, scalability, and operational efficiency.
  • Act as a key escalation point for critical platform issues and reliability concerns.
  • Mentor Associate SREs and SREs through guidance, reviews, and knowledge sharing.
  • Influence engineering teams without direct authority to adopt best practices in reliability and operations.
  • Act as a bridge between SRE, platform, and R&D teams to align on scalable and sustainable engineering practices.

Education

Bachelor's degree in Computer Science, Software Engineering or a related field

Job description

As a Senior Site Reliability Engineer at LANDI Global, you will play a critical role in defining and advancing the reliability, scalability, and performance of our platform infrastructures. You will work closely with cross‑functional teams to establish reliability standards, drive automation strategy, and lead continuous improvement initiatives across our environments.

Infrastructure & Platform Operations

Design, build, and optimize LANDI Global’s platform infrastructures across development, staging, and production environments, with a focus on scalability and resilience.

Collaborate with R&D and platform teams to define architecture patterns and reliability standards that ensure availability and operational excellence.

Lead platform readiness for new client onboarding, ensuring scalability, repeatability, and operational sustainability.

Monitoring, Reliability & Incident Management

Define and drive improvements in monitoring, logging, and alerting systems to ensure high signal quality and proactive issue detection.

Lead incident response for high severity events, and drive high-quality root causeanalysis (RCA) with a focus on systemic improvements.

Design, evolve, and validate Disaster Recovery (DR) and business continuity strategies, ensuring systems meet recovery objectives.

Participate in and help evolve the 24/7 standby model to improve operational effectiveness and sustainability.

Performance, Optimization& Automation

Analyze platform performance metrics and lead optimization strategies across cloud and on-prem environments.

Drive improvements in automated testing, CI/CD pipelines, and deployment workflows to enhance release safety, speed, and reliability.

Identify and eliminate operational toil through automation and engineering solutions.

Establish and standardize operational runbooks and procedures across services.

Operational Support

Provide advanced troubleshooting and support for complex production issues, guiding teams toward effective resolution.

Lead continuous improvement initiatives to enhance platform resilience, scalability, and operational efficiency.

Act as a key escalation points for critical platform issues and reliability concerns.

Technical Leadership & Collaboration

Mentor Associate SREs and SREs through guidance, reviews, and knowledge sharing.

Influence engineering teams without direct authority to adopt best practices in reliability and operations.

Act as a bridge between SRE, platform, and R&D teams to align on scalable and sustainable engineering practices.

REQUIREMENTS
  • Bachelor's degree in Computer Science, Software Engineering or a related field.
  • Minimum 7 years of experienceas a Site Reliability Engineer, DevOps Engineer, or in a similar role.
  • Strong verbal and written communication skills in English and Mandarin.
  • Strong experience designing and operating distributed systems at scale
  • Proven ability to improve reliability across multiple services or platforms
  • Deep understanding of system failure modes, scalability, and performance trade-offs
  • Experience defining and implementing SLOs, SLIs, and observability practices
  • Ability to lead incident response and drive systemic improvements
  • Strong communication skills with the ability to influence without authority
  • Demonstrated mentorship and technical leadership experience.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior SRE: Scale, Leadership & Automation
Senior SRE: Scale, Leadership & Automation

landi international (singapore) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

Longbridge Singapore • Singapore

On-site
SGD 90,000 - 130,000
Site Reliability Engineer, Enterprise Technology Services
Site Reliability Engineer, Enterprise Technology Services

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 200,000
Sr. SRE
Sr. SRE

United States Digital Space LLC • Singapore

On-site
SGD 120,000 - 180,000
On-site in Singapore (3 days/wk)
Site Reliability Engineer
Site Reliability Engineer

SGX Group • Singapore

On-site
SGD 180,000 - 300,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

VANGUARD SOFTWARE PTE. LTD. • Singapore

On-site
SGD 100,000 - 150,000
Technical Leadership
Career Growth
High-Performance Team
+1
SRE Engineer
SRE Engineer

BOUNTEOUSXACCOLITE SINGAPORE PTE. LTD. • Singapore

On-site
SGD 90,000 - 180,000
AVP - Site Reliability Engineer
AVP - Site Reliability Engineer

Singapore Exchange Limited • Singapore

On-site
SGD 140,000 - 210,000
Senior Platform Engineer / Site Reliability Engineer (SRE)
Senior Platform Engineer / Site Reliability Engineer (SRE)

AMBITION GROUP SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000