Senior Site Reliability Engineer

THG Ingenuity

Pune District

On-site

INR 3,200,000 - 5,400,000

Full time

6 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

THG Ingenuity is seeking a Senior SRE in India to own production services end-to-end, architect scalable cloud-native platforms on GCP, and drive reliability improvements.

You will lead incident response, implement SLOs, and mentor engineers while collaborating with multiple teams to reduce toil and improve efficiency.

Qualifications

  • Proven ability to own production systems at scale.
  • Strong focus on automation, reliability and capacity planning.
  • Experience with cloud-native architectures and Kubernetes-based platforms.

Responsibilities

  • Own services and domains end-to-end across design, build and operate phases.
  • Lead automation and platform component design and set best practices for infrastructure.
  • Drive toil-reduction projects and evaluate new tooling with buy/build recommendations.
  • Design and operate CI/CD pipelines with zero-downtime deployment strategies.

Skills

Golang
Python
GKE
Kubernetes
Terraform
Ansible
GitOps
CI/CD
Prometheus
Grafana
PostgreSQL
GCP
Cloud Architecture

Education

Software Engineering degree

Tools

Terraform
Ansible
Kubernetes
GCP
GitHub Actions
Flux CD
Argo CD

Job description

THG Ingenuity is a fully integrated digital commerce ecosystem, designed to power brands without limits. Our global end-to-end tech platform is comprised of three products: THG Commerce, THG Studios, and THG Fulfilment. Each represents a single, unified solution, overcoming challenges and taking brands direct-to-consumer. Our client portfolio includes globally recognised brands such as Coca-Cola, Nestle, Elemis, Homebase, and Proctor & Gamble. Technology is the driving force behind THG Ingenuity, and it starts with our people. We are ambitious with our goals and challenge conventional thinking.

THG Ingenuity is different because we support every single person to make a massive impact and drive their own work. Our people are always learning, and we work every day to ensure our technology, from our software platforms to our hosting services, to our AI capability and beyond, is world class. This enables us to keep powering THG Ingenuity and our partners on a global scale.

Tech at THG Ingenuity

Technology is the driving force behind THG Ingenuity, and it starts with our people. We are ambitious with our goals and challenge conventional thinking. THG Ingenuity is different because we support every single person to make a massive impact and drive their own work. Our people are always learning, and we work every day to ensure our technology, from our software platforms to our hosting services, to our AI capability and beyond, is world class. This enables us to keep powering THG Ingenuity and our partners on a global scale.

About the Team and Your Role

Our primary Kubernetes-as-a-Service platform, built end-to-end on Google Cloud Platform, provides THG Ingenuity's engineering teams with scalable, cloud-native infrastructure. Leveraging GCP's managed services, including GKE, we deliver comprehensive Kubernetes capabilities including multi-tenant clusters, dedicated clusters and App Runtime integration. GCP is our strategic home for container orchestration and cloud infrastructure, delivering scalability, reliability and reduced operational overhead through GCP's managed offerings.

We maintain seamless integration with App Runtime, THG Ingenuity's internal developer platform owned by a separate team. This partnership enables our GCP-based clusters to support App Runtime's simplified deployment interface, allowing developers to deploy applications without complex YAML configurations or approval workflows. Our K8saaS on GCP platform provides the underlying infrastructure and cluster management that powers App Runtime's developer experience.

What will I be doing?

As a Senior SRE, you will own services and problem domains end-to-end: designing, building, operating and improving them. You will multiply the effectiveness of the team through technical leadership as well as your own delivery.

In a developmental capacity, you will be responsible for:

  • Leading the design of automation and platform components and setting best practice for how infrastructure is built and operated.
  • Owning problem domains end-to-end, from design through build to operation.
  • Identifying sources of toil across the platform and driving projects that engineer them away.
  • Evaluating new tooling and approaches and making well-reasoned adopt/build/buy recommendations.
  • Contributing to upstream tooling where appropriate and representing the team in those communities.

In an operational capacity, you will be responsible for:

  • Acting as incident commander for major incidents, leading blameless post-mortems, and ensuring corrective actions of land.
  • Defining and refining SLIs and SLOs with tenant teams, and using error budgets to guide the balance between reliability work and feature work.
  • Leading complex upgrades and migrations across the GKE estate using zero-downtime strategies.
  • Capacity planning and performance engineering for the platform.
  • Designing and testing backup, failover and disaster recovery procedures against agreed RTO/RPO targets.

In a wider capacity, you will be part of the team and helping to:

  • Mentor and grow mid-level and junior engineers.
  • Plan and communicate significant pieces of work to stakeholders, including timelines and risk.
  • Evangelise App Runtime, GCP and Kubernetes within THG Ingenuity, and feed developer requirements into the platform roadmap.
What would success in this role look like?

Success is a platform whose reliability is designed rather than firefought. Services have meaningful SLOs, error budgets inform prioritisation, and the areas you own visibly reduce their operational load quarter on quarter.

Success is a new GCP region being brought online and the entire stack deploying to it in a short period of time, because the automation and GitOps pipelines you helped design make it routine.

Success is major incidents being rarer, shorter and better-learned-from because of the incident and post-mortem practice you drive.

Success is engineers around you getting measurably better because you mentor them, review their designs, and raise the bar on how the team works.

Role Requirements

The ideal candidate has a proven track record of owning production systems at scale, encompasses the DevOps and SRE mindset of wanting to automate, maintain a good service, practise what they preach, address technical debt, and is ready to lead as well as build.

A Software Engineering background is required, with focus on availability, performance, monitoring, capacity planning and change management.

Desired Technologies and Skills
  • Strong programming experience in Golang (primary) and Python (secondary), including designing maintainable tooling and services, not just scripts.
  • Deep public cloud expertise with a strong preference towards GCP: designing solutions around GCP managed services, IAM, and understanding their cost, performance and reliability trade-offs.
  • Infrastructure as Code at scale with Terraform: module design, state management and multi-environment patterns.
  • Experience with configuration management tools and concepts, for example: Ansible.
  • Deep Kubernetes (GKE) and Container expertise, including Helm chart authoring.
  • Debugging, troubleshooting, and extending CRDs and controllers.
  • Designing and operating Ingress controllers such as Nginx or HAP Roxy, and the Gateway API as their successor, including migration between the two.
  • Designing CI/CD pipelines in GitHub Actions and selecting appropriate zero-downtime deployment strategies (blue-green, canary, rolling) for a given service.
  • GitOps at scale with Flux CD and Argo CD, including multi-cluster patterns.
  • Observability architecture: operating Prometheus, Alert Manager and Grafana at scale, GCP's Cloud Operations suite, and designing SLO-based alerting that keeps pages actionable.
  • Operating, tuning and recovering PostgreSQL and etc..
  • Object storage design, including lifecycle and cost management.
  • Messaging systems: operating ActiveMQ in production and designing with GCP Pub/Sub.
  • End-to-end networking design: load balancing, firewall rules, DNS, CDN, VPC design, and automating the TLS certificate lifecycle.
  • Familiarity with different open-source licenses, and their implications.
  • Effective and appropriate use of AI tooling: using AI assistants and agents productively in day-to-day engineering, and building skills, agents and sub-agents with proper verification, validation and guardrails in place. Establishing patterns and guardrails for how the team uses these tools is expected at this level.
Nice to See
  • Cloud-native workload migrations.
  • Experience running services subject to compliance regimes such as PCI DSS.
Operational Skills
  • Defining SLIs, SLOs and SLAs with stakeholders, and operating an error budget in practice.
  • Incident command, on-call leadership, and driving a strong blameless post-mortem culture.
  • Security best practices: IAM, secrets management, compliance (PCI DSS for ecommerce).
  • Disaster recovery ownership: backup strategies, failover procedures, RTO/RPO planning and regular testing.
  • Performance optimization: load testing, capacity planning, autoscaling design.
  • Cost optimization: resource rights-sizing, committed use discounts, budget monitoring.
  • Documentation: setting the standard for runbooks, architecture diagrams, and migration plans.
Soft Skills
  • Mentoring: The ability to guide and grow junior and mid-level engineers.
  • Project management: migration timeline planning, stakeholder communication.
  • Cross-team collaboration: working with development, QA, and business teams.
  • Customer-facing platform management.
What’s in it for me?
  • Build solutions using the latest technology.
  • Continuous development through THG Academy, our in-house L&D team.
Equal Opportunity Statement

THG Ingenuity is an equal opportunity employer. We are committed to creating an inclusive, respectful and merit-based workplace where all employment decisions are made without discrimination on the basis of race, color, religion, caste, gender, gender identity or expression, sexual orientation, disability, age, marital status, pregnancy, nationality, veteran status, or any other status protected under applicable Indian laws. We encourage applications from candidates of all backgrounds and are committed to providing reasonable accommodation throughout the recruitment process, where required.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Enterprise Support Engineer (GOC) - THG India
Enterprise Support Engineer (GOC) - THG India

THG Ingenuity • Pune District

On-site
INR 600,000 - 900,000
Senior Software Engineer - Java
Senior Software Engineer - Java

ESP Engineered • Pune District

Hybrid
INR 800,000 - 1,200,000
Career advancement opportunities
Innovative work environment
Senior Database Administrator
Senior Database Administrator

THG Ingenuity • Pune District

On-site
INR 1,400,000 - 2,000,000
Lead, Site Reliability Engineer
Lead, Site Reliability Engineer

Toyota Connected India • Chennai

On-site
INR 1,500,000 - 2,000,000
Yearly gym membership reimbursement
Free catered lunches
Flexible dress code
+1
Senior Staff Site Reliability Engineer
Senior Staff Site Reliability Engineer

ServiceTitan, Inc. • India

On-site
INR 4,500,000 - 6,500,000
Sr. Devops Engineer
Sr. Devops Engineer

METRO Global Solution Center IN • Maharashtra

On-site
INR 3,500,000 - 7,000,000
Director Cloud & Infrastructure Architect (Multi-Cloud | Datacenter | SRE)
Director Cloud & Infrastructure Architect (Multi-Cloud | Datacenter | SRE)

Mancer Consulting Services • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Developer II - DevOps Engineering
Developer II - DevOps Engineering

UST • Maharashtra

On-site
INR 1,800,000 - 3,400,000
Tazapay - Staff DevOps Engineer
Tazapay - Staff DevOps Engineer

Tazapay • Chennai District

On-site
INR 3,500,000 - 6,000,000
Senior Platform Engineer
Senior Platform Engineer

IG Infotech • Bengaluru

On-site
INR 1,500,000 - 2,000,000