Staff Site Reliability Engineer

United States Digital Space LLC

Bengaluru

On-site

INR 6,000,000 - 12,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC seeks a Staff Site Reliability Engineer to lead reliability architecture for a new AI-focused platform from Bengaluru. You will shape multi-cloud strategy, own on-call practices, and guide capacity planning across regions.

The role emphasizes automation, collaboration with platform teams, and mentoring a growing SRE group in a hybrid work model.

Qualifications

  • 10+ years of experience in distributed systems with deep Kubernetes expertise and multi-cluster design.
  • Proficiency in Python, Go, or a similar programming language.
  • Understand workload isolation at the system level: containers, VMs, and trade-offs for untrusted code.
  • Customer-focused mindset and strong automation bias.
  • Experience with at least one major cloud provider (AWS, GCP, or Azure).
  • Track record of driving infrastructure architecture across teams and mentoring engineers.

Responsibilities

  • Own the reliability architecture of the platform across regions and cloud providers.
  • Collaborate with platform teams to provide operability guidance on capacity and best practices.
  • Set operational standards: on-call quality, incident response, and SLO discipline.
  • Mentor and develop the SRE team.
  • Participate in a 24/7 on-call rotation to resolve platform infra issues.

Skills

Kubernetes
Distributed systems
Python
Go
On-call ownership

Tools

Kubernetes
Docker
Terraform

Job description

We are seeking a Staff Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using the company.

As a Staff Site Reliability Engineer on this new team, you will be responsible for providing technical leadership for the operational foundations that enable deployment at scale of AI applications. You will own the reliability architecture of the platform as it expands across regions and cloud providers, and set the technical direction for how the platform is operated, including capacity planning, multi-cloud expansion, incident response, and SLO discipline. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. The company engineering teams pride themselves on building high-quality software and living the company cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition.

We are looking to speak to candidates who are based in Bengaluru for our hybrid working model.

Position Expectations
  • Own the reliability architecture of the platform across regions and cloud providers
  • Collaborate with the teams building the platform, providing internal support and guidance on operability, capacity, and best practices
  • Set operational standards for the team: on-call quality, incident response, SLO discipline
  • Mentor and technically develop the SRE team
  • Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure
Qualifications
  • 10+ years of experience working on software and operating distributed systems, with deep Kubernetes expertise, including designing or evolving multi-cluster platforms.
  • Proficiency in Python, Go, or a similar programming language
  • Understand workload isolation at the systems level: containers, virtual machines, and the trade-offs between them for running untrusted code
  • Possess a customer-focused mindset
  • Value efficiency in processes and operations, and display a strong preference for automation over manual processes
  • Be intimately familiar with the infrastructure primitives of at least one of AWS, GCP, or Azure, and comfortable reasoning about differences between them
  • Have a track record of driving infrastructure architecture across teams and mentoring engineers
About the company

the company is built for change, empowering our customers and our people to innovate at the speed of the market. We have redefined the data platform for the AI era, enabling builders to create, transform, and disrupt industries with software. the company’s unified data platform, the most widely available, globally distributed data platform on the market, helps organizations modernize legacy workloads, embrace innovation, and unleash AI. Our cloud-native platform, the company Atlas, is the only globally distributed, multi-cloud data platform and is available across AWS, Google Cloud, and Microsoft Azure.

With offices worldwide and over 67,000 customers, including 75% of the Fortune 100 and AI-native startups, relying on the company for their most important applications, we’re powering the next era of software.

Our compass at the company is our Leadership Commitment, guiding how and why we make decisions, show up for each other, and win. It’s what makes us the company.

To drive the personal growth and business impact of our employees, we’re committed to developing a supportive and enriching culture for everyone. From employee affinity groups, to fertility assistance and a generous parental leave policy, we value our employees’ wellbeing and want to support them along every step of their professional and personal journeys. Learn more about what it’s like to work at the company, and help us make an impact on the world!

the company is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter.

the company is an equal opportunities employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Gurugram District

Hybrid
INR 3,500,000 - 7,000,000
Staff Engineer
Staff Engineer

United States Digital Space LLC • Bengaluru

Hybrid
INR 3,500,000 - 5,500,000
Senior Software Engineer
Senior Software Engineer

The Consulting Solutions • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Hybrid working model
Career development opportunities
Staff Site Reliability Engineer
Staff Site Reliability Engineer

MongoDB • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Manager, Engineering
Manager, Engineering

United States Digital Space LLC • Bengaluru

Hybrid
INR 6,000,000 - 9,000,000
Lead, Engineering
Lead, Engineering

The Consulting Solutions • Gurugram District

On-site
INR 2,500,000 - 3,500,000
Employee affinity groups
Fertility assistance
Generous parental leave policy
Lead SRE
Lead SRE

United States Digital Space LLC • Karnataka

On-site
INR 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

Arch Systems • Hyderabad

On-site
INR 2,800,000 - 4,200,000
Software Engineer 3
Software Engineer 3

The Consulting Solutions • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AlleyCorp • Gurugram District

Hybrid
INR 3,000,000 - 4,200,000