TechGrove is the Centre of Excellence for Banyan Software, based in Chennai, India. It plays a key role in supporting Banyan's global businesses through technology, security, and software development. TechGrove brings together India's deep pool of technical talent with Banyan's long-term approach to growth, creating a trusted, developer-focused environment where people can do their best work.
Job Title: Senior SRE (Site Reliability Engineer) - Modernized Application Operations
Overview
We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos).
You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos' distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds - Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.
Key Responsibilities
- 24x7 Operations & On-Call: Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.
- Tier 1 SRE & Operations: Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds - AWS and Azure - to keep production systems healthy, performant, and secure.
- Performance & Availability Monitoring/Observability: Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.
- Security Incident Response: Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation
- Automation & Infrastructure-as-Code : Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.
- AI Agents & DevSecOps Scale: Build scale in our DevSecOps practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation.
- Hands-on Problem Solving: Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.
Required Qualifications & Experience
- Experience: 5-7 years of progressive experience in Software Engineering, and/or Site Reliability Engineering, with a focus on operating distributed systems.
- Containerization: Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems.
- Infrastructure-as-Code with Terraform: Have experience working with modules at scale. This is a requirement for the role.
- Cloud Native Services: hands-on experience operating production workloads on Amazon Web Services (AWS) (e.g., EC2, Lambda, EKS, S3, RDS) and / or Microsoft Azure (e.g., Container Apps, AKS, Container Storage).
- CI/CD & Automation: Deep history of hands-on work with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices directly into operational workflows.
- Operations, Monitoring & Observability: Experience with application level logging, troubleshooting, and tracing tools, with a proven track record operating highly available production systems.
- AI-Fluent Engineering: Experience with AI-assisted engineering tools such as Claude Code or similar
- Application Performance Management (APM): Familiarity with APM tooling and practices (e.g., Datadog, New Relic, Dynatrace, or similar) to instrument, profile, and optimize application performance in production.
- Incident & Security Response: Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures.
- Communication & Collaboration: Exceptional communication, presentation, and collaboration skills, with a proven ability to coordinate across teams.
- Education: Bachelor's degree in Computer Science or a related technical field.
Preferred Skills (A Plus)
Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.