Lead SRE

United States Digital Space LLC

Karnataka

On-site

INR 900,000 - 1,400,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC is seeking an experienced Site Reliability Engineer (SRE) to strengthen cloud infrastructure, platform reliability, automation, observability, and production support. This role emphasizes improving service reliability with SLIs/SLOs and reducing toil through automation while partnering with global teams to ensure resilient, scalable platforms.

The ideal candidate will own capacity planning, performance tuning, and incident management, and contribute to CI/CD

Qualifications

  • 8-12 years of experience in Site Reliability Engineering.
  • Strong hands-on experience with: Azure, Kubernetes, Docker, Jenkins, Terraform, Ansible.
  • Proficiency in Python, Go, or Bash.
  • Strong understanding of Networking & Security.
  • Experience in L2/L3 production support, incident management, and root cause analysis.
  • Experience working with global teams across multiple time zones.
  • Strong communication, stakeholder management, and ownership mindset.
  • Willingness to work from the Bangalore ITPL office, 5 days a week.

Responsibilities

  • Define, measure, and report SLIs, SLOs, and error budgets for critical services.
  • Drive service reliability improvements and reduce operational toil through automation.
  • Own capacity planning, performance tuning, and scalability initiatives.
  • Lead blameless postmortems, root cause analyses, and corrective action tracking.
  • Design and maintain CI/CD pipelines using Azure DevOps and Jenkins.
  • Operate and manage Azure cloud infrastructure and Kubernetes platforms.
  • Deploy and support containerized applications using Docker, Kubernetes, and Helm.
  • Automate infrastructure provisioning and configuration using Terraform and Ansible.
  • Manage artifacts and repositories using JFrog Artifactory.
  • Implement monitoring, logging, and observability using Prometheus, Grafana, Loki, and OpenTelemetry.
  • Provide L2/L3 production support, incident management, troubleshooting, and RCA.
  • Participate in on-call rotations supporting critical production services.
  • Support PostgreSQL, Redis, and RabbitMQ environments, including high availability, backups, replication, and performance tuning.
  • Partner with Development, QA, Product, Operations, and global engineering teams.
  • Ensure platform availability, scalability, security, and performance.

Skills

Azure
Kubernetes
Docker
Jenkins
Terraform
Ansible
Python
Go
Bash
Networking & Security

Education

Bachelor’s degree in Computer Science or related field

Tools

Azure DevOps
JFrog Artifactory
OpenTelemetry
Prometheus
Grafana
Loki

Job description

Role description
Who we are:

At the company, we help the world’s best organizations grow and succeed through transformation. Bringing together the right talent, tools, and ideas, we work with our client to co-create lasting change. Together, with over 30,000 employees in 30+ countries, we build for boundless impact—touching billions of lives in the process. Visit us at .

Job Description
Role Summary

We are seeking an experienced Site Reliability Engineer (SRE) with strong hands-on expertise in cloud infrastructure, platform reliability, automation, observability, and production support. This role focuses on improving service reliability through SLIs/SLOs and error budgets, reducing operational toil through automation, and partnering with global engineering teams to ensure resilient, scalable, and secure platforms.

Key Responsibilities
Reliability Engineering
  • Define, measure, and report SLIs, SLOs, and error budgets for critical services.
  • Drive service reliability improvements and systematically reduce operational toil through automation.
  • Own capacity planning, performance tuning, and scalability initiatives.
  • Lead blameless postmortems, root cause analyses, and corrective action tracking.
Platform Reliability & Automation
  • Design and maintain CI/CD pipelines using Azure DevOps and Jenkins.
  • Operate and manage Azure cloud infrastructure and Kubernetes platforms.
  • Deploy and support containerized applications using Docker, Kubernetes, and Helm.
  • Automate infrastructure provisioning and configuration using Terraform and Ansible.
  • Manage artifacts and repositories using JFrog Artifactory.
Observability & Production Support
  • Implement monitoring, logging, and observability using Prometheus, Grafana, Loki, and OpenTelemetry.
  • Provide L2/L3 production support, incident management, troubleshooting, and RCA.
  • Participate in on-call rotations supporting critical production services.
  • Support PostgreSQL, Redis, and RabbitMQ environments, including high availability, backups, replication, and performance tuning.
Collaboration
  • Partner with Development, QA, Product, Operations, and global engineering teams.
  • Ensure platform availability, scalability, security, and performance.
Mandatory Skills & Experience
  • 8-12 years of experience in Site Reliability Engineering
  • Strong hands-on experience with: Azure, Kubernetes, Docker, Jenkins, Terraform, Ansible
  • Proficiency in Python, Go, or Bash.
  • Strong understanding of Networking & Security.
  • Experience in L2/L3 production support, incident management, and root cause analysis.
  • Experience working with global teams across multiple time zones.
  • Strong communication, stakeholder management, and ownership mindset.
  • Willingness to work from the Bangalore ITPL office, 5 days a week.
Good to Have
  • Experience with AI/GenAI concepts and AIOps practices.
  • Exposure to chaos engineering and resilience testing.
  • Azure (AZ-104/AZ-400), CKA, or CKAD certifications.
  • Bachelor’s degree in computer science or a related field.
Key Skills

Azure Kubernetes Docker Helm Azure DevOps Jenkins Terraform Ansible JFrog Artifactory Prometheus Grafana Loki OpenTelemetry PostgreSQL Redis RabbitMQ Python/Go/Bash SLIs/SLOs Incident Management RCA Networking & Security.

What we believe:

We’re proud to embrace the same values that have shaped the company since the beginning. Since day one, we’ve been building enduring relationships and a culture of integrity. And today, it's those same values that are inspiring us to encourage innovation from everyone, to champion diversity and inclusion and to place people at the centre of everything we do.

Humility

We will listen, learn, be empathetic and help selflessly in our interactions with everyone.

Humanity

Through business, we will better the lives of those less fortunate than ourselves.

Integrity

We honour our commitments and act with responsibility in all our relationships.

Equal Employment Opportunity Statement

the company is an Equal Opportunity Employer. We believe that no one should be discriminated against because of their differences, such as age, disability, ethnicity, gender, gender identity and expression, religion, or sexual orientation.

All employment decisions shall be made without regard to age, race, creed, colour, religion, sex, national origin, ancestry, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by federal, state, or local law.

the company reserves the right to periodically redefine your roles and responsibilities based on the requirements of the organization and/or your performance.

  • To support and promote the values of the company.
  • Comply with all Company policies and procedures
Skills

AIOps, Adobe Experience Manager, Azure DevOps, Chaos Engineering

About the company

the company is a global digital transformation solutions provider. For more than 20 years, the company has worked side by side with the world’s best companies to make a real impact through transformation. Powered by technology, inspired by people and led by purpose, the company partners with their clients from design to operation. With deep domain expertise and a future-proof philosophy, the company embeds innovation and agility into their clients’ organizations. With over 30,000 employees in 30 countries, the company builds for boundless impact—touching billions of lives in the process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead SRE
Lead SRE

UST • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Maharashtra

On-site
INR 1,800,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Lead SRE & Support Engineer
Lead SRE & Support Engineer

Providence Global Center • Hyderabad

On-site
INR 3,500,000 - 5,500,000
Competitive Pay
Supportive Reporting Relation
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Site Reliability Engineer
Site Reliability Engineer

LSEG • Port Blair

On-site
INR 1,400,000 - 2,200,000
Healthcare and retirement planning
Paid volunteering days
Wellbeing initiatives
Senior Site Reliability Engineer
Senior Site Reliability Engineer

VMC Soft Technologies, Inc • Hyderabad

Hybrid
INR 1,500,000 - 2,000,000
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering
Senior Team Lead | Engineering, AI & Data - Engineering | Site Reliability Engineering

Deloitte & Touche GmbH Wirtschaftsprüfungsgesellschaft • Bengaluru

On-site
INR 2,000,000 - 3,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000