Site Reliability Engineer (SRE)

C1X

Chennai District

On-site

INR 1,800,000 - 3,200,000

Full time

10 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

C1X is seeking a Site Reliability Engineer to enhance the availability, performance, scalability, and resilience of our production services in Chennai. You will define SLIs/SLOs, design monitoring and observability, and automate operations to reduce toil.

Collaborate with developers, DevOps and cloud teams; work on Kubernetes, Terraform, CI/CD and incident response. Strong problem solving and programming in Python/Go/Java are required to build automation and improve workflows.

Qualifications

  • Define and monitor SLIs, SLOs and error budgets to measure reliability and expectations.
  • Design and maintain monitoring, alerting and observability solutions for visibility into health.
  • Automate repetitive tasks using scripting and infrastructure automation practices.
  • Investigate incidents, troubleshoot issues and conduct post-incident reviews to prevent recurrence.
  • Collaborate with software developers, DevOps, cloud teams and stakeholders on reliability.
  • Programming in Python, Go or Java to build automation tools and improve workflows.

Responsibilities

  • Provide support for production releases and improve system performance.
  • Ensure infrastructure scales with changing workloads and business needs.
  • Maintain runbooks and incident-management procedures for reliability.

Skills

SLI/SLOs
Monitoring
Observability
Linux
Networking
Scripting
Kubernetes
Terraform
CI/CD
Incident response
Capacity planning
Disaster recovery
Python/Go/Java

Tools

Prometheus
Grafana
ELK
OpenTelemetry

Job description

We are looking for a skilled Site Reliability Engineer (SRE) to improve the availability, performance, scalability, and operational resilience of our production services. The role focuses on building reliable systems, reducing operational complexity, and ensuring that business‑critical applications remain stable, efficient, and highly available.

The candidate will be responsible for defining and monitoring Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to measure service reliability and establish clear performance expectations. You will design and maintain monitoring, alerting, and observability solutions to provide visibility into application and infrastructure health.

The role involves automating repetitive operational tasks and reducing manual intervention through scripting, infrastructure automation, and reliable deployment practices. You will investigate production incidents, troubleshoot system and application issues, coordinate recovery activities, and conduct blameless post‑incident reviews to identify root causes and prevent recurring problems.

Strong knowledge of Linux, networking, scripting, distributed systems, cloud platforms, Kubernetes, and Terraform is required. Experience with observability and monitoring tools such as Prometheus, Grafana, ELK, OpenTelemetry, or similar technologies is expected. The candidate should also have a good understanding of CI/CD, incident response, capacity planning, disaster recovery, and reliability engineering practices.

You will work closely with software developers, DevOps engineers, cloud teams, and other technical stakeholders to design reliable services, improve deployment safety, and identify potential reliability risks throughout the application lifecycle. The role also includes supporting production releases, improving system performance, and ensuring that infrastructure can handle changing workloads and business requirements.

The ideal candidate should have strong problem‑solving and troubleshooting skills, an automation‑focused mindset, and the ability to work effectively during production incidents. Programming experience with Python, Go, Java, or a comparable language is required to build automation tools and improve operational workflows.

You will also contribute to capacity planning, disaster‑recovery strategies, system resilience, and continuous reliability improvements, while maintaining clear operational documentation, runbooks, monitoring standards, and incident‑management procedures.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Recro • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Site Reliability Engineer
Site Reliability Engineer

Snapmint • Gurugram District

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer
Site Reliability Engineer

Yantran • Chennai District

On-site
INR 900,000 - 1,400,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Senior Site Reliability Engineer (SRE)
Senior Site Reliability Engineer (SRE)

Lonvec Technologies Private Limited • Hyderabad

On-site
INR 3,000,000 - 5,000,000
Site Reliability Engineer (SRE) – DevOps Infrastructure
Site Reliability Engineer (SRE) – DevOps Infrastructure

PQAngels Technologies Pvt. Ltd. • Bengaluru

On-site
INR 1,200,000 - 2,100,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer
Site Reliability Engineer

3across • Bengaluru

Hybrid
INR 3,500,000 - 6,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Bahwan CyberTek • Hyderabad

On-site
INR 1,200,000 - 1,800,000