Synthlane is a leading IT services and consulting firm that helps businesses, large enterprises, and government institutions solve complex technology challenges with strategic, scalable solutions. The company specializes in enterprise IT consulting, digital transformation, cybersecurity and risk management, cloud and infrastructure services, and custom software development. Synthlane focuses on aligning technology initiatives with business objectives, enabling clients to modernize operations, strengthen security, and improve productivity. Its experienced team of consultants, engineers, and strategists works closely with clients to design future‑ready solutions that deliver measurable results. Organizations seeking a trusted partner for reliable, innovation‑driven IT services rely on Synthlane to support long‑term success.
Role Description
We are seeking an accomplished Senior Site Reliability Engineer (SRE) to lead the
design, implementation, and evolution of highly available, scalable, and resilient systems
across our multi‑cloud infrastructure. In this senior role, you will drive architectural
decisions, establish reliability standards, and mentor teams while ensuring operational
excellence across complex distributed systems. You will partner with engineering
leadership, development teams, and product stakeholders to shape infrastructure
strategy, implement sophisticated automation, and champion a culture of reliability
As a Senior SRE, you'll tackle sophisticated, large‑scale challenges using cutting‑edge
technologies across AWS and Azure platforms. You will lead critical initiatives that
impact system reliability at scale, architect solutions for complex infrastructure problems,
and guide teams in adopting industry‑leading practices that drive meaningful
improvements across our entire technology ecosystem.
Requirements
- Architect and implement highly reliable, scalable, and cost‑effective infrastructure
- solutions for mission‑critical applications across multi‑cloud environments (AWS and
- Lead the definition and refinement of service level objectives (SLOs), service level
- indicators (SLIs), and error budgets, establishing reliability standards across the
- organization.
- Design and implement sophisticated Infrastructure as Code (IaC) solutions using
- Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.
- Drive automation strategies to eliminate toil, improve operational efficiency, and enable
- self‑service capabilities for development teams.
- Lead incident response efforts, conduct thorough post‑incident reviews, and implement
- systemic improvements to prevent recurrence.
- Champion cloud‑native architectures and modern reliability practices, serving as a
- technical advisor for infrastructure and platform decisions.
- Participate in and help optimize the on‑call rotation, ensuring sustainable practices and
- Establish and maintain comprehensive documentation standards, runbooks, and
- knowledge repositories that enable team autonomy and effective incident response.
- Design and implement advanced monitoring, logging, and alerting strategies using
- observability platforms to enable proactive issue detection and resolution.
- Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement
- sophisticated deployment strategies including blue‑green, canary, and progressive
- delivery patterns.
- Ensure security, compliance, and governance standards are embedded throughout the
- infrastructure lifecycle, implementing security‑as‑code practices.
- Drive capacity planning, performance optimization, and cost management initiatives
- Architect and implement highly reliable, scalable, and cost‑effective infrastructure
- solutions for mission‑critical applications across multi‑cloud environments (AWS and
- Lead the definition and refinement of service level objectives (SLOs), service level
- indicators (SLIs), and error budgets, establishing reliability standards across the
- organization.
- Design and implement sophisticated Infrastructure as Code (IaC) solutions using
- Terraform, Ansible, and Azure Resource Manager (ARM) templates or Bicep.
- Drive automation strategies to eliminate toil, improve operational efficiency, and enable
- self‑service capabilities for development teams.
- Lead incident response efforts, conduct thorough post‑incident reviews, and implement
- systemic improvements to prevent recurrence.
- Champion cloud‑native architectures and modern reliability practices, serving as a
- technical advisor for infrastructure and platform decisions.
- Participate in and help optimize the on‑call rotation, ensuring sustainable practices and
- Establish and maintain comprehensive documentation standards, runbooks, and
- knowledge repositories that enable team autonomy and effective incident response.
- Design and implement advanced monitoring, logging, and alerting strategies using
- observability platforms to enable proactive issue detection and resolution.
- Lead container orchestration initiatives using Kubernetes (AKS, EKS) and implement
- sophisticated deployment strategies including blue‑green, canary, and progressive
- delivery patterns.
- Ensure security, compliance, and governance standards are embedded throughout the
- infrastructure lifecycle, implementing security‑as‑code practices.
- Drive capacity planning, performance optimization, and cost management initiatives
- Collaborate with architecture and security teams to establish platform standards,
- reference architectures, and best practices.