Staff S/W Engg , SRE & Platform Automation Engg

Aziro

Bengaluru

On-site

INR 4,000,000 - 7,000,000

Full time

9 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Aziro in Bengaluru, India is seeking a Staff Software Engineer to lead Site Reliability Platform Automation within the SaaS Platform Engineering team. You will own end-to-end delivery of a functional area across global cloud networking and platform reliability, hands-on with code and automation.

You will design production software for reliability workflows, drive SLI/SLO definitions, toil automation, and self-service tooling, collaborating with DevOps, CloudOps, security, and product management

Qualifications

  • 8+ years of software engineering experience in infrastructure, platform reliability, DevOps, or related systems.

Responsibilities

  • Own a functional area end to end, including architecture, implementation, rollout, adoption, operational health, and evolution

Skills

Go
Python
Terraform
Kubernetes
AWS/GCP
CI/CD
Observability
DevOps
Automation

Tools

Prometheus
Grafana
OpenTelemetry
ELK
PagerDuty
Jenkins
GitHub Actions
Argo

Job description

Job Summary

Staff Software Engineer, Site Reliability Platform Automation We have an opportunity for a Staff Software Engineer, Site Reliability Platform Automation to join our SaaS Platform Engineering team in Bangalore, India, reporting to the Sr. Manager, Site Reliability Platform Engineering. You will own the technical delivery of a functional area across global cloud networking and SaaS platforms. You will work hands-on across reliability measurement or toil automation, building systems that reduce operational effort and enable product engineering teams to operate more independently. One opening is anchored in reliability measurement, including SLO and SLI definition across the service catalog, error budget policy and burn alerting, and the observability practice and evidence store that make reliability claims verifiable. The second is anchored in toil automation, including automated CVE remediation across four production realms, test automation for third-party provider software, and automated EKS and RDS upgrades. You will collaborate with DevOps, CloudOps, product engineering, architecture, security, and product management teams. You will also help apply AI-assisted engineering and operational tools responsibly to improve productivity, analytics, automation, and incident decision support.

Responsibilities
  • Own a functional area end to end, including architecture, implementation, rollout, adoption, operational health, and evolution
  • Design and build production software that automates infrastructure and reliability workflows; spend most of your time writing and reviewing software
  • For the reliability measurement opening, define and implement SLIs, SLOs, error budgets, burn-rate alerting, observability standards, and a durable evidence store across critical and core services
  • For the toil automation opening, build automated workflows for CVE remediation across four realms, EKS and RDS upgrades, third-party provider testing, and recurring infrastructure operations
  • Work directly with NA DevOps, IN DevOps, and CloudOps engineers to understand manual workflows before replacing them with reliable, measurable automation
  • Build self-service capabilities wherever product engineering teams can safely operate them independently, using guardrails rather than approval gates
  • Establish measurement baselines for toil, reliability, adoption, developer experience, and operational cost before changing systems or workflows
  • Build golden paths, Terraform modules, policy-as-code controls, and reusable platform components that shift DevOps ownership toward the teams that build and own the services
  • Partner with product engineering teams to improve production readiness, observability, incident response, capacity planning, change safety, and operational ownership
  • Land adoption across peer engineering and product engineering teams; treat usage, satisfaction, and reduced manual effort as part of delivery
Requirements
  • 8+ years of software engineering experience, with meaningful depth in infrastructure, platform, reliability, DevOps, or related systems
  • Experience owning a significant production system, including its failure modes, operational cost, reliability, and evolution
  • Strong production coding ability in Go, Python, or a comparable language, together with infrastructure-as-code fluency, particularly Terraform
  • Deep practical Kubernetes experience, including operating, debugging, and improving real production clusters
  • Experience with AWS or GCP, CI/CD systems, distributed systems, and production operations
  • Operational judgment demonstrated through on-call experience, incident participation, production troubleshooting, and sound decision-making under pressure
  • A track record of replacing recurring manual work with software and explaining the measured improvement before and after automation
  • Experience building internal platforms, developer tooling, or reliability capabilities that other engineering teams adopted
  • For the reliability measurement opening: experience defining SLIs and SLOs for services owned by other teams and negotiating meaningful targets with service owners
  • For the toil automation opening: experience with vulnerability remediation at scale, automated infrastructure upgrades, or test automation for third-party software that you cannot modify
  • Nice to have Policy-as-code experience with Kyverno, OPA, or Gatekeeper
  • Observability stack experience with Prometheus, Grafana, Loki, Cortex, OpenTelemetry, ELK, Datadog, PagerDuty, or comparable technologies
  • CI/CD platform engineering experience, including Jenkins, GitHub Actions, GitOps, Argo, or large-scale CI/CD platform migration
  • Experience with multi-tenant, multi-region, or regulated environments, including FedRAMP, SOC 2, or ISO-controlled platforms
  • Internal developer platform or developer experience work with measurable adoption outcomes
  • Experience applying LLM-based or agentic tools to engineering or operational workflows, including human approval and controlled remediation
  • Experience with EKS, RDS, PostgreSQL, cloud networking, or disaster recovery testing

Location: Bangalore, India

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr Engg Manager, SRE & Platform Automation
Sr Engg Manager, SRE & Platform Automation

Aziro • Bengaluru

On-site
INR 4,500,000 - 7,000,000
Sr S/W Engg, SRE & Platform Automation
Sr S/W Engg, SRE & Platform Automation

Aziro • Bengaluru

On-site
INR 3,000,000 - 5,000,000
AWS Delivery Manager
AWS Delivery Manager

EPAM Systems Inc • India

Hybrid
INR 6,000,000 - 9,000,000
Lead Site Reliability Engineer or Platform Engineer
Lead Site Reliability Engineer or Platform Engineer

Weekday (YC W21) • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

BuildxPartners • Bengaluru Urban

Hybrid
INR 2,400,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
MTS 2, Platform Reliability Engineer
MTS 2, Platform Reliability Engineer

The Networker • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Senior Software Engineer
Senior Software Engineer

NVIDIA • Bengaluru

On-site
INR 4,000,000 - 7,000,000