Senior Cloud Site Reliability Engineer AP

Lighthouse Document Technologies Inc.

India

On-site

INR 4,200,000 - 6,800,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Lighthouse Document Technologies Inc. in Bengaluru, India, is seeking a Senior Cloud Site Reliability Engineer who will own reliability, scalability, and security of our cloud platforms and product infrastructure.

You will drive SRE practices, build observability, automate tasks, and work with cross‑functional teams to improve MTTR and production readiness. This is a full‑time, on-site role based in Bengaluru.

Qualifications

  • Bachelor’s degree or equivalent experience in computer science or engineering.
  • Strong knowledge of cloud platforms, Kubernetes, and DevOps practices.
  • Experience with automation, observability, and incident management.
  • Proficiency in scripting (Python, Bash, PowerShell) and IaC tools.

Responsibilities

  • Drive and implement SRE best practices across cloud platforms and services.
  • Define and improve SLIs, SLOs, SLAs, and error budgets.
  • Improve reliability, scalability, and operational efficiency.
  • Conduct RCA and postmortems to drive improvements.
  • Participate in on‑call rotations and incident response.

Skills

Kubernetes
Cloud platforms
DevOps
Automation
Observability
Python
Terraform
Ansible
ARM/Bicep
PowerShell

Education

Bachelor’s degree in Computer Science, Engineering, or related field

Tools

Grafana
Prometheus
Azure Monitor
Log Analytics
ELK Stack
ElasticSearch
Power BI
Azure DevOps
Jenkins
GitHub Actions
Terraform
ARM/Bicep

Job description

Full Time Nagar, Bengaluru, Karnataka, IN

30+ days ago Requisition ID: 2569

What is special about Lighthouse?

Lighthouse is built on a foundation of unique, compassionate, highly driven individuals. We elevate the strengths and talents of those around us while leveraging opportunities for growth. We offer the experience of solving complex problems while continuing to grow multiple facets of your career. Lighthouse is where innovation meets support and where collaboration is the key ingredient to success. We grow together and are stronger together.

What’s unique about this role?

The Senior Cloud Site Reliability Engineer (Senior Cloud SRE) is responsible for ensuring the reliability, scalability, availability, performance, security, and operational excellence of Lighthouse’s cloud platforms and critical product infrastructure.

This role combines software engineering, cloud engineering, automation, observability, and operational governance practices to build highly resilient and self-healing platforms across hybrid and cloud-native environments. The ideal candidate will drive SRE best practices, improve service reliability through automation, establish observability standards, and partner closely with Engineering, Product, Security, DBA, and DevEx teams to improve operational maturity across the organization.

The role requires deep expertise in cloud infrastructure, Kubernetes, DevOps/SRE principles, telemetry, incident management, monitoring, and automation, along with strong collaboration and communication skills.

What will this person do?
Site Reliability Engineering & Operational Excellence
  • Drive and implement Site Reliability Engineering (SRE) best practices across cloud platforms and services.
  • Define, maintain, and improve:
    • Service Level Indicators (SLIs)
    • Service Level Objectives (SLOs)
    • Service Level Agreements (SLAs)
    • Error Budgets
  • Improve service reliability, resiliency, scalability, and operational efficiency.
  • Establish operational standards, reliability governance, and production readiness practices.
  • Conduct Root Cause Analysis (RCA), postmortems, and reliability improvement initiatives.
  • Participate in on-call rotations, incident management, and major incident resolution activities.
  • Continuously improving operational processes to reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR)
Observability, Monitoring & Telemetry
  • Design, implement, and maintain enterprise observability and telemetry platforms.
  • Build operational dashboards, reliability scorecards, and service health monitoring solutions.
  • Configure proactive alerting, anomaly detection, and incident correlation mechanisms.
  • Implement centralized monitoring and telemetry using:
    • Grafana
    • Prometheus
    • Azure Monitor
    • Log Analytics
    • ELK Stack / ElasticSearch
    • Power BI dashboards
  • Develop actionable operational metrics and telemetry reporting for engineering and leadership teams.
  • Enhance visibility into infrastructure, application, Kubernetes, and platform health.
  • Drive automation-first operational practices across infrastructure and platform services.
  • Develop Infrastructure-as-Code (IaC) solutions using:
    • Terraform
    • ARM/Bicep
    • Ansible
  • Build operational automation scripts using:
    • Python
    • Bash
    • PowerShell
  • Develop self-healing and auto-remediation capabilities for recurring operational incidents.
  • Automate infrastructure provisioning, monitoring, scaling, backup, recovery, and deployment workflows.
  • Reduce manual operational effort and improve engineering productivity through intelligent automation.
Collaboration & Engineering Partnership
  • Collaborate closely with:
    • Cloud Engineering teams
    • Product Engineering teams
    • DevEx teams
    • Security teams
    • DBA teams
    • Operations teams
  • Support engineering teams in improving production readiness and operational maturity.
  • Contribute to continuous improvement initiatives, reliability reviews, and operational excellence programs.
Bring your passion and together we will shine. It would also be great if you had the following:
  • Bachelor’s degree in Computer Science, Engineering, or related field (or equivalent experience/certification).
  • Knowledge of Python, scripting, or Infrastructure-as-Code tools (e.g., Terraform, Ansible, ARM/Bicep).
  • Experience managing cloud platforms (e.g., Azure, AKS, Pivotal Cloud Foundry, or equivalent).
  • Strong understanding of Kubernetes and containerization concepts.
  • Experience with application packaging, deployment automation, and release management.
  • Solid knowledge of relational databases (MS-SQL) and exposure to NoSQL technologies (e.g., Redis, ElasticSearch, MongoDB).
  • Experience with CI/CD tools (Azure DevOps, Jenkins, GitHub Actions, or similar).
  • Familiarity with monitoring and logging tools (Grafana, ELK stack, Prometheus, PowerBI, etc.).
  • Proficiency with Git and modern branching/merging workflows.
  • Strong Linux administration and troubleshooting skills.
  • Excellent problem-solving, communication, and teamwork skills.
Work Environment and Physical Demands
  • Duties are performed in a typical office environment while at a desk or computer table.
  • Duties require the ability to use a computer, communicate over the telephone, and read printed material, in a quiet and professional setting.
  • Duties may require being on call periodically and working outside normal working hours (evenings and weekends).

Lighthouse celebrates and thrives on diversity and is an Equal Opportunity Employer. We hire, train, and promote regardless of race, religion, color, national origin, sex, disability, age, veteran status, and other protected status as required by applicable law. We welcome any talents and contributions you can bring to the team and are deeply committed to growing an environment where everyone can feel safe, is respected, and can show up as themselves. Come as you are!

As required by applicable pay transparency laws, Lighthouse complies with compensation disclosure requirements for roles that may be hired in locations under these requirements. Factors that may be used to determine your actual salary may include a wide array of factors, including: your specific skills and experience, geographic location, or other relevant factors. The salary range for this position may be tailored to be lower or higher in different talent markets.

This role will be eligible to participate in an annual bonus or incentive program.

As a trailblazer and catalyst for change, Lighthouse rises to each opportunity to help our clients, and our people do what they do best—shine.

This position will work for and be employed by Lighthouse's India subsidiary, which is an independent company located in India.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Cloud Site Reliability Engineer AP
Senior Cloud Site Reliability Engineer AP

Lighthouse • Bengaluru

On-site
INR 1,500,000 - 2,000,000
Annual bonus or incentive program
Diversity and Inclusion initiatives
Flexible working hours
Cloud Infrastructure Engineer AP
Cloud Infrastructure Engineer AP

Lighthouse • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Annual bonus or incentive program
Technical Analyst AP
Technical Analyst AP

Lighthouse • India

On-site
INR 400,000 - 640,000
Technical Analyst AP
Technical Analyst AP

Lighthouse • Bengaluru

On-site
INR 500,000 - 700,000
Solution Consultant AP
Solution Consultant AP

Lighthouse • Mumbai

On-site
INR 1,500,000 - 2,300,000
Annual bonus or incentive program
Managed Review Operations Analyst - AP
Managed Review Operations Analyst - AP

Lighthouse • Bengaluru

On-site
INR 500,000 - 750,000
Site Reliability Engineer
Site Reliability Engineer

London Stock Exchange Group • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Healthcare
Retirement planning
Paid volunteering days
+1
Associate Database Administrator
Associate Database Administrator

CL 88e84fc2 cfc3 4bec b5cf 08a3bda081bc • Nagar

On-site
INR 350,000 - 480,000
Annual bonus program
Associate Database Administrator
Associate Database Administrator

Lighthouse • Bengaluru

On-site
INR 300,000 - 420,000
Lead SRE
Lead SRE

UST • Bengaluru

On-site
INR 4,000,000 - 7,000,000