Lead Site Reliability Engineer (SRE/DBA)

Kochava India

Bengaluru

On-site

INR 4,500,000 - 7,500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Kochava India is seeking an experienced Site Reliability Engineer to join our fast‑growing team in Bengaluru. You will be hands‑on, owning production infrastructure, databases, and shared platforms across cloud and on‑prem environments.

You will drive reliability and efficiency through automation, observability, and best practices. You will collaborate with platform, development and infrastructure teams to improve uptime, performance, and resilience in a hybrid data center setup, with a focus on

Qualifications

  • Minimum 8+ years in SRE/DevOps/Platform engineering or equivalent infrastructure role.

Responsibilities

  • Design, build, automate, troubleshoot, and operate production infrastructure and database systems in a 24x7 environment across GCP, AWS, and on‑prem.
  • Own production services from implementation to operations, incident response, and optimization.
  • Define and maintain SLOs/SLIs/SLAs with cross‑functional teams.
  • Improve observability, reliability, and performance of platforms.
  • Provide on‑call coverage and ensure high availability and disaster recovery for databases.

Skills

Kubernetes
Terraform
Linux
MySQL
AWS
GCP
Observability
CI/CD
Go
Python
Java

Tools

Terraform
Kubernetes

Job description

What the job involves

You will be joining a fast-growing team of motivated and talented engineers as a hands‑on Individual Contributor (IC) responsible for ensuring the reliability, scalability, and performance of Kochavas core infrastructure and database platforms.

As part of the Site Reliability Engineering team, you will design, build, manage, and optimize systems that support mission‑critical production services across cloud and on‑prem environments. In addition to core SRE responsibilities, this role requires strong database administration expertise and the ability to provide operational coverage for critical database infrastructure.

Working closely with platform, development and infrastructure teams, you will contribute to improving system availability, performance and operational efficiency by enhancing existing platforms and enabling new capabilities.

You will play a key role in managing shared infrastructure that supports databases, message queues, monitoring systems, networking and security layers across hybrid environments. This is a highly hands‑on Individual Contributor role where you will spend the majority of your time designing, building, troubleshooting, automating, and operating production systems.

Who you are
Required Skills:
  • 10+ years of experience working as an SRE, DevOps or Platform Engineer in large‑scale production environments.
  • Must be a hands‑on Individual Contributor with recent experience designing, implementing, troubleshooting, and operating production infrastructure.
  • Strong hands‑on experience managing containerized workloads in Kubernetes.
  • Proven expertise in Infrastructure as Code tools such as Terraform.
  • Deep understanding of Linux systems, from OS‑level operations to performance troubleshooting.
  • Experience supporting production environments across cloud platforms such as AWS or GCP.
  • Strong troubleshooting capability across infrastructure, application and networking layers.
  • Experience working in 24x7 production environments.
  • Strong hands‑on database administration experience with MySQL in large‑scale production environments, including performance tuning, replication, backup and recovery, capacity planning, and troubleshooting.
  • Experience supporting and maintaining highly available database infrastructure, with the ability to provide DBA coverage and operational ownership when required.
  • Working knowledge of complementary data platforms such as Redis and Elasticsearch from an operational standpoint.
  • Experience implementing observability, monitoring and reliability best practices.
  • Familiarity with CI/CD workflows and production deployment pipelines.
  • Experience managing database reliability, availability, and performance in mission‑critical production environments.
  • Strong understanding of database monitoring, query optimization, incident response, and disaster recovery best practices.
Key Responsibilities
  • Personally design, build, automate, troubleshoot, and operate production infrastructure, shared platforms, and database systems in a 24x7x365 environment across GCP, AWS, and physical data centers.
  • Take end‑to‑end ownership of assigned production services from implementation through operations, incident response, troubleshooting, performance optimization, and automation.
  • Work with cross‑functional teams to define and maintain SLOs, SLIs and SLAs
  • Evangelize best practices related to system performance, reliability and availability.
  • Continuously improve observability to ensure uptime and resilience of applications and infrastructure.
  • Troubleshoot issues across the entire stack infrastructure, application, network and hardware layers.
  • Provide on‑call support for shared services and production infrastructure. Manage, maintain, and optimize production database environments, ensuring high availability, performance, scalability, and reliability.
  • Perform database administration activities including replication management, backup and recovery validation, capacity planning, performance tuning, and incident troubleshooting.
  • Partner with engineering teams to improve database architecture, resiliency, observability, and operational excellence.
  • Act as a backup resource for DBA responsibilities and participate in database‑related on‑call and incident response activities.
Preferred/Optional Skills:
  • Exposure to distributed systems such as Kafka, Memcached or InfluxDB or similar large‑scale data platforms.
  • Experience managing hybrid infrastructure across cloud and physical data centers.
  • Programming exposure in Go, Python or Java.
  • Experience supporting high‑availability and failover‑ready environments.
  • Experience operating large‑scale data platforms and stateful workloads within Kubernetes environments.
Minimum Qualifications:
  • You have spend at least 8+ years as an SRE, DevOps engineer Platform Engineer, DBA/SRE, or equivalent infrastructure‑focused role.
  • Strong hands‑on experience administering production MySQL environments, including replication, backup and recovery, performance tuning, and high‑availability configurations.
  • In‑depth knowledge of containerization and managing complex workloads in Kubernetes
  • Proven track record of personally designing, building, troubleshooting, and operating large‑scale production infrastructure in a hands‑on engineering role.
  • Ability to work effectively as part of a distributed team
Soft Skills:
  • You enjoy solving reliability and performance challenges across infrastructure and database platforms.
  • You take ownership of infrastructure stability database health, and service availability.
  • You can work independently in production‑critical environments and make sound operational decisions during incidents.
  • You collaborate effectively with distributed engineering, platform, and development. teams.
  • You are comfortable operating in fast‑paced environments with evolving priorities.
  • You approach operational excellence with a strong focus on automation, observability, and continuous improvement.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Falabella India • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

Smart Ims • Bengaluru

Hybrid
INR 1,200,000 - 2,000,000
Senior Site Reliability Engineer - Infrastructure
Senior Site Reliability Engineer - Infrastructure

S&P Global Market Intelligence • Bengaluru, Hyderabad

On-site
INR 1,500,000 - 2,500,000
Sr. Site Reliability Engineer
Sr. Site Reliability Engineer

Visa Consolidated Support Services India • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F-Prime Capital • Pune District

On-site
INR 1,500,000 - 2,000,000