Get more replies from employers
Send a job-specific resume in minutes.
LANDI Global is seeking a Senior Site Reliability Engineer to drive reliability across our platform infrastructures. You will define standards, lead incident response, and push automation across cloud and on‑prem environments.
You will collaborate with R&D and platform teams to ensure availability, resilience, and operational excellence while mentoring junior engineers and shaping SRE practices across the organization.
As a Senior Site Reliability Engineer at LANDI Global, you will play a critical role in defining and advancing the reliability, scalability, and performance of our platform infrastructures. You will work closely with cross‑functional teams to establish reliability standards, drive automation strategy, and lead continuous improvement initiatives across our environments.
Design, build, and optimize LANDI Global’s platform infrastructures across development, staging, and production environments, with a focus on scalability and resilience.
Collaborate with R&D and platform teams to define architecture patterns and reliability standards that ensure availability and operational excellence.
Lead platform readiness for new client onboarding, ensuring scalability, repeatability, and operational sustainability.
Define and drive improvements in monitoring, logging, and alerting systems to ensure high signal quality and proactive issue detection.
Lead incident response for high severity events, and drive high-quality root causeanalysis (RCA) with a focus on systemic improvements.
Design, evolve, and validate Disaster Recovery (DR) and business continuity strategies, ensuring systems meet recovery objectives.
Participate in and help evolve the 24/7 standby model to improve operational effectiveness and sustainability.
Analyze platform performance metrics and lead optimization strategies across cloud and on-prem environments.
Drive improvements in automated testing, CI/CD pipelines, and deployment workflows to enhance release safety, speed, and reliability.
Identify and eliminate operational toil through automation and engineering solutions.
Establish and standardize operational runbooks and procedures across services.
Provide advanced troubleshooting and support for complex production issues, guiding teams toward effective resolution.
Lead continuous improvement initiatives to enhance platform resilience, scalability, and operational efficiency.
Act as a key escalation points for critical platform issues and reliability concerns.
Mentor Associate SREs and SREs through guidance, reviews, and knowledge sharing.
Influence engineering teams without direct authority to adopt best practices in reliability and operations.
Act as a bridge between SRE, platform, and R&D teams to align on scalable and sustainable engineering practices.