Principal SRE

Azira

Bengaluru

On-site

INR 4,000,000 - 9,000,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Azira is seeking a seasoned Site Reliability Engineer to lead the evolution of our cloud infrastructure and reliability practices. You will architect scalable, secure systems, drive incident response, and advance IaC and automation across global products and AI workloads.

The role requires deep AWS experience, strong Kubernetes and Linux expertise, and a track record of mentoring engineers while balancing performance, security, and cost. Bengaluru-based with a global scope.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related field.
  • 12–15 years of experience in Site Reliability Engineering, Cloud Infrastructure, Platform Architecture, or related roles at senior levels.
  • Deep expertise in AWS; experience with GCP is valuable.
  • Strong understanding of cloud-native architecture (Kubernetes, containers, Linux, storage, networks).
  • Proficiency with IaC tools (Terraform/OpenTofu) and automation.

Responsibilities

  • Define and evolve Azira’s cloud infrastructure and reliability strategy for global products and AI workloads.
  • Design scalable, secure, and highly available systems with disaster-recovery capabilities.
  • Improve reliability and performance of large-scale production systems by addressing failures and capacity constraints.
  • Develop SLIs/SLOs and error-budget practices to set measurable reliability targets.
  • Strengthen observability via monitoring, logging, tracing, and alerting across apps and infra.
  • Lead incident response as a senior escalation point and implement lessons learned.
  • Advance infrastructure-as-code, automation, and environment management for consistency and scalability.
  • Collaborate with Engineering, Data, AI, Product, and Security to embed reliability and security in the development lifecycle.
  • Enhance cloud cost visibility and balance performance, capacity, and spend; support business continuity.
  • Mentor SREs and engineers, communicating risks and trade-offs to leaders and stakeholders.

Skills

SRE leadership
cloud infra
DevOps
observability
incident response
communication
mentoring

Education

Bachelor’s or Master’s degree in CS/Engineering

Tools

Kubernetes
Terraform/OpenTofu
AWS
GCP
GitHub/Bitbucket

Job description

Azira is a data-first media and insights company on a mission to reinvent how brands use data to make smarter decisions, from where to open their next location to how they connect with customers in the real world. We blend marketing, location analytics, and strategy into a single platform, helping leading brands take action with confidence. We move fast, think boldly, and care deeply about building things that matter.

Why This Role Matters

This role will contribute to the Development and Improvement of Azira’s Cloud infrastructure, addressing key areas such as scalability, observability, security, and cost efficiency. The role will collaborate with teams across Engineering, AI, Product, Data, and Security to improve infrastructure standards, solve complex reliability challenges, and provide technical guidance and mentorship across the organization.

What you’ll do

  • Define and evolve Azira’s cloud infrastructure and reliability strategy, ensuring it supports global products, data platforms, and AI initiatives.
  • Design and implement scalable, resilient, secure, and highly available systems, including architecture standards, best practices, and disaster-recovery capabilities.
  • Improve the reliability and performance of distributed, high-volume production systems by identifying and addressing single points of failure, capacity constraints, and recurring sources of instability.
  • Develop and mature service-level indicators (SLIs), service-level objectives (SLOs), and error-budget practices to establish measurable reliability standards.
  • Strengthen observability across applications and infrastructure through effective monitoring, logging, tracing, and alerting.
  • Lead the technical response to complex production incidents, contribute as a senior escalation point when needed, and ensure incident learnings result in lasting improvements.
  • Advance infrastructure-as-code, automation, and environment management practices to improve consistency, repeatability, scalability, and operational efficiency.
  • Partner with Engineering, Data, AI, Product, and Security teams to embed reliability, security, and operational readiness throughout the development lifecycle, including supporting the infrastructure needs of AI workloads.
  • Improve cloud cost visibility and efficiency while balancing performance, capacity, scalability, and spend; strengthen business continuity, backup, and disaster-recovery practices.
  • Mentor SREs and engineers, raise technical standards, and communicate infrastructure risks, trade-offs, and recommendations clearly to technical leaders and business stakeholders.

What you Bring

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 12–15 years of experience in Site Reliability Engineering, Cloud Infrastructure, Platform Architecture, or related roles, with experience operating at a Principal, Staff-plus, Architect, or equivalent level.
  • Significant experience in Site Reliability Engineering, cloud infrastructure, platform engineering, DevOps, or a related discipline, with experience operating at a Principal, Staff-plus, Architect, or equivalent level.
  • Deep expertise in AWS, including experience architecting and operating infrastructure for products involving big data, high-volume workloads, or distributed systems. Experience with GCP in similar environments is also valuable.
  • Strong understanding of cloud-native architecture, including Kubernetes, containers, networking, Linux, storage, databases, Amazon EMR, and cloud security.
  • Advanced experience with Infrastructure-as-Code tools such as Terraform or OpenTofu, with a focus on building consistent, repeatable, and scalable infrastructure.
  • Strong scripting and programming skills in Python, Bash, Go, or comparable languages, along with experience working with CI/CD and source-control platforms such as GitHub or Bitbucket.
  • Experience with centralized authentication and authorization systems, including concepts such as OIDC and RBAC, as well as a solid understanding of infrastructure-level security and compliance requirements.
  • Experience designing and operating distributed, data-intensive, or highly available systems at scale, with a strong understanding of observability across metrics, logs, traces, and alerting.
  • Proven experience leading complex incident response, root-cause analysis, and reliability improvement initiatives, with the ability to balance immediate operational needs with long-term architectural improvements.
  • Strong technical judgment and communication skills, with the ability to evaluate trade-offs across reliability, performance, security, speed, and cost, and communicate recommendations clearly across teams and regions.
  • A collaborative leadership approach with a track record of mentoring engineers, influencing without formal authority, constructively challenging existing approaches, and taking ownership of problems through to resolution.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Namely • India

On-site
INR 1,500,000 - 2,500,000
Lead SRE
Lead SRE

Cvent, Inc. • Gurugram District

On-site
INR 4,000,000 - 8,000,000
Lead SRE
Lead SRE

Cvent, Inc. • India

On-site
INR 2,500,000 - 4,500,000
Senior SRE Engineer
Senior SRE Engineer

Epam Systems • Chennai District

On-site
INR 2,500,000 - 4,000,000
Lead SRE
Lead SRE

Cvent • Gurugram District

On-site
INR 4,000,000 - 7,000,000
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Senior Data Engineer
Senior Data Engineer

Azira • Bengaluru

On-site
INR 1,400,000 - 2,500,000
Health and wellness benefits
Flexible work environment
High impact role
+1
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior SRE
Senior SRE

CloudRaft, Inc. • India

On-site
INR 2,500,000 - 5,000,000
Competitive salary
Premium health insurance
GPU infrastructure projects
Senior Consultant - Site Reliability Engineer
Senior Consultant - Site Reliability Engineer

Darwinbox Digital Solutions Pvt. Ltd. • Hyderabad

On-site
INR 3,000,000 - 5,200,000