Senior Site Reliability Engineer

United States Digital Space LLC

Gurugram District

On-site

INR 3,500,000 - 7,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

United States Digital Space LLC in Gurugram invites a Senior Site Reliability Engineer to shape architecture and build operational foundations for a new AI platform. You will drive scalable deployment, improve performance, and ensure reliability of distributed systems while collaborating with a product & technology team.

The role requires hands-on experience with Kubernetes, cloud platforms, and Linux networking, along with a strong bias for automation and efficiency.

Qualifications

  • Strong background in software development and operating distributed systems.
  • 6+ years of experience building and operating distributed systems with Python/Go or similar.
  • Production experience operating Kubernetes and debugging at node level.
  • Expertise in cloud platforms AWS, GCP, or Azure.
  • Strong understanding of Linux internals, TCP/IP, DNS, TLS and routing.
  • Excellent verbal and written technical communication, with collaboration mindset.

Responsibilities

  • Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads.
  • Build for reliability with fault tolerance and self-healing capabilities.
  • Identify and configure metrics to detect incidents and quantify service health, availability, and performance.
  • Participate in a 24/7 on-call rotation to resolve platform issues.
  • Mentor junior SREs and contribute to evolving operational practices.

Skills

Distributed systems
On-call experience
Python/Go
Cloud infrastructure
Linux networking

Tools

Kubernetes
AWS
GCP
Azure

Job description

We are seeking a Senior Site Reliability Engineer to join our growing Gurugram Products & Technology team to provide technical direction, shape architecture, and build key operational foundations of a new platform we are building to make it easier for customers to build AI applications using the company.

As a Senior Site Reliability Engineer on this new team, you will be responsible for enabling deployment at scale of AI applications and improving the performance, scalability, and reliability of the distributed systems infrastructure for this new product. The platform's SRE team owns the operational foundations: the Kubernetes fleet, networking, observability and alerting, and tenant isolation. the company engineering teams pride themselves on building high-quality software and living the company cultural values every day – we value intellectual curiosity and honesty, and building together in an environment that prioritizes collaboration over competition.

We are looking to speak to candidates who are based in Gurugram for our hybrid working model.

Position Expectations
  • Operate and improve the multi-tenant Kubernetes infrastructure that runs customer workloads
  • Build for reliability, making services and infrastructure available, resilient, fault-tolerant, and self-healing
  • Identify and configure key metrics to detect incidents and quantify service health, availability, and performance
  • Participate in a 24/7 on-call rotation to resolve issues involving platform infrastructure
  • Mentor early-career SREs and contribute to the team’s operational practices as it grows
Qualifications
  • Strong background in software development and operating distributed systems
  • 6+ years of experience building and operating distributed systems, with proficiency in Python, Go, or a similar programming language
  • Experience operating Kubernetes in production and debugging below the abstraction layer, including scheduling, cluster networking, and node-level issues
  • Expertise in cloud infrastructure platforms, including AWS, Google Cloud Platform (GCP), or Azure
  • Strong understanding of Linux operating system internals and networking concepts such as TCP/IP, DNS, TLS, and routing
  • Customer-focused mindset and strong verbal and written technical communication skills, with a desire to collaborate with colleagues
  • Strong bias for efficient processes, operational simplicity, and automation over manual work
  • Bonus points for experience with Kubernetes networking, such as Istio or Cilium, service mesh or edge load balancing in production, secure multi-tenant runtime environments at scale, multi-cloud infrastructure management, and virtualization or workload isolation technologies
  • Eager to learn, with a strong technical background

the company is committed to providing any necessary accommodations for individuals with disabilities within our application and interview process. To request an accommodation due to a disability, please inform your recruiter.

the company is an equal opportunities employer.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Bengaluru

Hybrid
INR 6,000,000 - 12,000,000
Senior Platform Engineer
Senior Platform Engineer

The Consulting Solutions • Gurugram District

On-site
INR 1,200,000 - 1,800,000
Senior Software Engineer
Senior Software Engineer

The Consulting Solutions • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Hybrid working model
Career development opportunities
Staff Engineer
Staff Engineer

United States Digital Space LLC • Bengaluru

Hybrid
INR 3,500,000 - 5,500,000
Software Engineer 3
Software Engineer 3

The Consulting Solutions • Gurugram District

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

Arch Systems • Hyderabad

On-site
INR 2,800,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F-Prime Capital • Pune District

On-site
INR 1,500,000 - 2,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AlleyCorp • Gurugram District

Hybrid
INR 3,000,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000