Senior SRE: Scale, Leadership & Automation

landi international (singapore) pte. ltd.

Singapore

On-site

SGD 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

LANDI Global is seeking a Senior Site Reliability Engineer to drive reliability across our platform infrastructures. You will define standards, lead incident response, and push automation across cloud and on‑prem environments.

You will collaborate with R&D and platform teams to ensure availability, resilience, and operational excellence while mentoring junior engineers and shaping SRE practices across the organization.

Qualifications

  • Bachelor's degree in Computer Science, Software Engineering or a related field.
  • Minimum 7 years of experience as a Site Reliability Engineer, DevOps Engineer, or in a similar role.
  • Strong verbal and written communication skills in English and Mandarin.
  • Strong experience designing and operating distributed systems at scale.
  • Proven ability to improve reliability across multiple services or platforms.
  • Deep understanding of system failure modes, scalability, and performance trade-offs.
  • Experience defining and implementing SLOs, SLIs, and observability practices.
  • Ability to lead incident response and drive systemic improvements.
  • Strong communication skills with the ability to influence without authority.
  • Demonstrated mentorship and technical leadership experience.

Responsibilities

  • Design, build, and optimize LANDI Global’s platform infrastructures across development, staging, and production environments, with a focus on scalability and resilience.
  • Collaborate with R&D and platform teams to define architecture patterns and reliability standards that ensure availability and operational excellence.
  • Lead platform readiness for new client onboarding, ensuring scalability, repeatability, and operational sustainability.
  • Define and drive improvements in monitoring, logging, and alerting systems to ensure high signal quality and proactive issue detection.
  • Lead incident response for high severity events, and drive high-quality root cause analysis (RCA) with a focus on systemic improvements.
  • Design, evolve, and validate Disaster Recovery (DR) and business continuity strategies, ensuring systems meet recovery objectives.
  • Participate in and help evolve the 24/7 standby model to improve operational effectiveness and sustainability.
  • Analyze platform performance metrics and lead optimization strategies across cloud and on-prem environments.
  • Drive improvements in automated testing, CI/CD pipelines, and deployment workflows to enhance release safety, speed, and reliability.
  • Identify and eliminate operational toil through automation and engineering solutions.
  • Establish and standardize operational runbooks and procedures across services.
  • Provide advanced troubleshooting and support for complex production issues, guiding teams toward effective resolution.
  • Lead continuous improvement initiatives to enhance platform resilience, scalability, and operational efficiency.
  • Act as a key escalation point for critical platform issues and reliability concerns.
  • Mentor Associate SREs and SREs through guidance, reviews, and knowledge sharing.
  • Influence engineering teams without direct authority to adopt best practices in reliability and operations.
  • Act as a bridge between SRE, platform, and R&D teams to align on scalable and sustainable engineering practices.

Education

Bachelor's degree in Computer Science, Software Engineering or a related field

Job description

LANDI Global is seeking a Senior Site Reliability Engineer to drive reliability across our platform infrastructures. You will define standards, lead incident response, and push automation across cloud and on‑prem environments.

You will collaborate with R&D and platform teams to ensure availability, resilience, and operational excellence while mentoring junior engineers and shaping SRE practices across the organization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

landi international (singapore) pte. ltd. • Singapore

On-site
SGD 120,000 - 180,000
Senior Site Reliability Engineer - Scale & Resilience Leader
Senior Site Reliability Engineer - Scale & Resilience Leader

Kidentify • Singapore

On-site
SGD 120,000 - 180,000
SRE Lead: Scale, Reliability & Observability on AWS
SRE Lead: Scale, Reliability & Observability on AWS

kidentify pte. ltd. • Singapore

On-site
SGD 80,000 - 120,000
Senior Platform Engineer / Site Reliability Engineer (SRE)
Senior Platform Engineer / Site Reliability Engineer (SRE)

AMBITION GROUP SINGAPORE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Lead Platform Site Reliability Engineer
Lead Platform Site Reliability Engineer

JPMorgan Chase & Co. • Singapore

On-site
SGD 120,000 - 190,000
Cloud-Scale SRE: Reliability, Observability & Incident Mastery
Cloud-Scale SRE: Reliability, Observability & Incident Mastery

SEVEN HILLS CONSULTING PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Senior SRE: Global Incident Commander & Automation
Senior SRE: Global Incident Commander & Automation

Goldman Sachs • Singapore

On-site
SGD 150,000 - 210,000
Senior SRE Leader: Reliability, Automation & Cloud Ops
Senior SRE Leader: Reliability, Automation & Cloud Ops

REOLINK TECHNOLOGY PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Insurance Coverage
Yearly Bonus
Performance Bonus
Global SRE & Reliability Transformation Leader
Global SRE & Reliability Transformation Leader

DBS BANK LTD. • Singapore

On-site
SGD 300,000 - 420,000
Senior SRE: Global Infra, Automation & 24/7 Uptime
Senior SRE: Global Infra, Automation & 24/7 Uptime

MOZAT PTE LTD • Singapore

On-site
SGD 120,000 - 170,000
Competitive compensation
Performance-based bonuses
Growth opportunities
+1