Senior Cloud Site Reliability Engineer

Augusta Infotech

Bengaluru

Hybrid

INR 1,500,000 - 2,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Augusta Infotech is looking for a Senior Cloud Site Reliability Engineer to ensure the reliability and operational excellence of cloud platforms. The ideal candidate should have deep expertise in cloud infrastructure and SRE principles, coupled with solid experience in Kubernetes and automation.

Your responsibilities will include implementing SRE best practices, designing observability platforms, and collaborating with diverse engineering teams to enhance operational maturity. A Bachelor's degree in Computer Science is required, along with strong programming skills.

Qualifications

  • Bachelors degree in Computer Science or related field.
  • Experience managing cloud platforms like Azure or Kubernetes.
  • Knowledge of scripting and Infrastructure-as-Code tools.

Responsibilities

  • Ensure reliability and performance of cloud platforms.
  • Implement Site Reliability Engineering best practices.
  • Develop automation scripts for operational efficiency.

Skills

Cloud infrastructure
Kubernetes
Python
CI/CD tools
Linux administration

Education

Bachelor's degree in Computer Science or related field

Tools

Terraform
Grafana
Prometheus
Azure Monitor
ELK Stack

Job description

Role & responsibilities

The Senior Cloud Site Reliability Engineer (Senior Cloud SRE) is responsible for ensuring the reliability, scalability, availability, performance, security, and operational excellence for cloud platforms and critical product infrastructure.

This role combines software engineering, cloud engineering, automation, observability, and operational governance practices to build highly resilient and self-healing platforms across hybrid and cloud-native environments. The ideal candidate will drive SRE best practices, improve service reliability through automation, establish observability standards, and partner closely with Engineering, Product, Security, DBA, and DevEx teams to improve operational maturity across the organization.

The role requires deep expertise in cloud infrastructure, Kubernetes, DevOps/SRE principles, telemetry, incident management, monitoring, and automation, along with strong collaboration and communication skills.

Site Reliability Engineering & Operational Excellence
  • Drive and implement Site Reliability Engineering (SRE) best practices across cloud platforms and services.
  • Define, maintain, and improve:
    • Service Level Indicators (SLIs)
    • Service Level Objectives (SLOs)
    • Service Level Agreements (SLAs)
    • Error Budgets
  • Improve service reliability, resiliency, scalability, and operational efficiency.
  • Establish operational standards, reliability governance, and production readiness practices.
  • Conduct Root Cause Analysis (RCA), postmortems, and reliability improvement initiatives.
  • Participate in on-call rotations, incident management, and major incident resolution activities.
  • Continuously improving operational processes to reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
Observability, Monitoring & Telemetry
  • Design, implement, and maintain enterprise observability and telemetry platforms.
  • Build operational dashboards, reliability scorecards, and service health monitoring solutions.
  • Configure proactive alerting, anomaly detection, and incident correlation mechanisms.
  • Implement centralized monitoring and telemetry using:
    • Grafana
    • Prometheus
    • Azure Monitor
    • Log Analytics
    • ELK Stack / ElasticSearch
    • Power BI dashboards
  • Develop actionable operational metrics and telemetry reporting for engineering and leadership teams.
  • Enhance visibility into infrastructure, application, Kubernetes, and platform health.
Automation & Auto-Healing Engineering
  • Drive automation-first operational practices across infrastructure and platform services.
  • Develop Infrastructure-as-Code (IaC) solutions using:
    • Terraform
    • ARM/Bicep
    • Ansible
  • Build operational automation scripts using:
    • Python
    • Bash
    • PowerShell
  • Develop self-healing and auto-remediation capabilities for recurring operational incidents.
  • Automate infrastructure provisioning, monitoring, scaling, backup, recovery, and deployment workflows.
  • Reduce manual operational effort and improve engineering productivity through intelligent automation.
Collaboration & Engineering Partnership
  • Collaborate closely with:
    • Cloud Engineering teams
    • Product Engineering teams
    • DevEx teamsSecurity teams
    • DBA teams
    • Operations teams
  • Support engineering teams in improving production readiness and operational maturity.
  • Contribute to continuous improvement initiatives, reliability reviews, and operational excellence programs.
Experience
  • Bachelors degree in Computer Science, Engineering, or related field (or equivalent experience/certification).
  • Knowledge of Python, scripting, or Infrastructure-as-Code tools (e.g., Terraform, Ansible, ARM/Bicep).
  • Experience managing cloud platforms (e.g., Azure, AKS, Pivotal Cloud Foundry, or equivalent).
  • Strong understanding of Kubernetes and containerization concepts.
  • Experience with application packaging, deployment automation, and release management.
  • Solid knowledge of relational databases (MS-SQL) and exposure to NoSQL technologies (e.g., Redis, ElasticSearch, MongoDB).
  • Experience with CI/CD tools (Azure DevOps, Jenkins, GitHub Actions, or similar).
  • Familiarity with monitoring and logging tools (Grafana, ELK stack, Prometheus, PowerBI, etc.).
  • Proficiency with Git and modern branching/merging workflows.
  • Strong Linux administration and troubleshooting skills.
  • Excellent problem-solving, communication, and teamwork skills.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Visa Consolidated Support Services India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Five9 • Bengaluru

On-site
INR 4,000,000 - 7,000,000
SRE Engineer
SRE Engineer

Prodapt Solutions Private Limited • Chennai District

On-site
INR 1,800,000 - 3,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

iLink Digital • Chennai

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Lonvec Technologies Private Limited • Hyderabad

On-site
INR 3,000,000 - 5,000,000