Senior SRE

TechGrove by Banyan Software

Bengaluru

On-site

INR 2,500,000 - 4,200,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

TechGrove by Banyan Software in Bengaluru, India seeks a Senior SRE to own operational excellence for modernized SaaS apps. You will manage 24x7 on-call coverage, automate deployments with Terraform and CI/CD pipelines, and ensure secure, highly available production systems on AWS/Azure.

The ideal candidate brings 5-7 years in SRE/Software Engineering, container expertise, and experience with AI-assisted tooling like Claude Code, Datadog, and resilient incident response practices.

Qualifications

  • 5-7 years of progressive experience in Software Engineering or SRE, operating distributed systems.
  • Deep expertise in container technologies (Docker/Kubernetes) for scalable distributed systems.
  • Infrastructure-as-Code with Terraform at scale.
  • Hands-on experience with AWS and/or Azure cloud services.
  • Experience with CI/CD platforms (GitHub Actions, GitLab CI) and DevSecOps practices.
  • Experience with AI-assisted engineering tools such as Claude Code or similar.
  • Familiarity with APM tooling (Datadog/New Relic/Dynatrace).
  • Incident response and disaster recovery procedures.
  • Strong communication and collaboration skills.
  • Bachelor's degree in Computer Science or a related field.

Responsibilities

  • Operate 24x7 with on-call rotations to ensure availability.
  • Serve as Tier 1 SRE for modernized SaaS applications across AWS and Azure.
  • Implement and maintain observability tooling for performance and reliability.
  • Respond to security incidents and follow runbooks for remediation.
  • Use Terraform and CI/CD pipelines to automate deployments.
  • Develop AI agents to scale DevSecOps and incident response.
  • Troubleshoot infrastructure, network, and automation issues across multi-tenant environments.

Skills

Containerization
CI/CD
Observability
Security incident response
Communication & collaboration
AI-fluent engineering

Education

Bachelor's degree in Computer Science or related field

Tools

Docker
Kubernetes
Terraform
GitHub Actions
GitLab CI
Datadog

Job description

TechGrove is the Centre of Excellence for Banyan Software, based in Chennai, India. It plays a key role in supporting Banyan's global businesses through technology, security, and software development. TechGrove brings together India's deep pool of technical talent with Banyan's long-term approach to growth, creating a trusted, developer-focused environment where people can do their best work.

Job Title: Senior SRE (Site Reliability Engineer) - Modernized Application Operations
Overview

We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).

You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos' distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds - Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.

Key Responsibilities
  • 24x7 Operations & On-Call: Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.
  • Tier 1 SRE & Operations: Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds - AWS and Azure - to keep production systems healthy, performant, and secure.
  • Performance & Availability Monitoring/Observability: Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.
  • Security Incident Response: Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation
  • Automation & Infrastructure-as-Code : Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.
  • AI Agents & DevSecOps Scale: Build scale in our DevSecOps practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation.
  • Hands-on Problem Solving: Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.
Required Qualifications & Experience
  • Experience: 5-7 years of progressive experience in Software Engineering, and/or Site Reliability Engineering, with a focus on operating distributed systems.
  • Containerization: Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems.
  • Infrastructure-as-Code with Terraform: Have experience working with modules at scale. This is a requirement for the role.
  • Cloud Native Services: hands-on experience operating production workloads on Amazon Web Services (AWS) (e.g., EC2, Lambda, EKS, S3, RDS) and / or Microsoft Azure (e.g., Container Apps, AKS, Container Storage).
  • CI/CD & Automation: Deep history of hands-on work with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices directly into operational workflows.
  • Operations, Monitoring & Observability: Experience with application level logging, troubleshooting, and tracing tools, with a proven track record operating highly available production systems.
  • AI-Fluent Engineering: Experience with AI-assisted engineering tools such as Claude Code or similar
  • Application Performance Management (APM): Familiarity with APM tooling and practices (e.g., Datadog, New Relic, Dynatrace, or similar) to instrument, profile, and optimize application performance in production.
  • Incident & Security Response: Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures.
  • Communication & Collaboration: Exceptional communication, presentation, and collaboration skills, with a proven ability to coordinate across teams.
  • Education: Bachelor's degree in Computer Science or a related technical field.
Preferred Skills (A Plus)

Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Solutions By Text • Bengaluru

On-site
INR 800,000 - 1,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Senior Manager - Site Reliability Engineer|NR-2026-0246
Senior Manager - Site Reliability Engineer|NR-2026-0246

Media.net • Bengaluru

On-site
INR 6,000,000 - 8,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Clarus Advisers • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Pune District

On-site
INR 2,250,000 - 2,750,000
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

New Era Technology • Gurugram District

On-site
INR 1,500,000 - 2,100,000
Site Reliability Engineer (SRE) – Core IT Infrastructure
Site Reliability Engineer (SRE) – Core IT Infrastructure

TECEZE • Chennai District

On-site
INR 1,000,000 - 2,000,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000