Site Reliability Engineer (SRE) - Engineering Productivity

Jobgether

India

On-site

INR 1,800,000 - 3,000,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Jobgether in India is seeking a Site Reliability Engineer (SRE) – Engineering Productivity to scale and secure our production systems across hybrid cloud environments. You will design, build, and operate resilient infrastructure while collaborating with software teams to remove bottlenecks and improve the developer experience.

The role combines software engineering, automation, observability, and production operations, with ownership from deployment to incident response.

Qualifications

  • Bachelor’s or Master’s degree in CS/Engineering or related field, or equivalent professional experience.
  • Proficient in Go, Python, or shell scripting with automation workflows.
  • Strong Linux/UNIX admin and debugging skills.
  • Experience operating software systems, infrastructure, or complex applications at scale.
  • Experience with infrastructure-as-code and automation tools.
  • Familiarity with CI/CD tools and observability platforms.
  • Experience with databases such as MariaDB, PostgreSQL, MongoDB.
  • Strong communication and collaboration skills.

Responsibilities

  • Design, build, deploy, and operate production systems with a focus on scalability, reliability, observability, performance, and security.
  • Develop automation that reduces operational toil and improves production workflows.
  • Monitor infrastructure and services, improve alerting, and implement automated responses.
  • Create and improve incident response procedures and runbooks.
  • Collaborate with product teams to identify bottlenecks and design solutions in infrastructure.
  • Investigate and resolve platform issues, supporting engineering teams with triage and debugging.
  • Write post-incident reviews and implement corrective measures to prevent recurrences.

Skills

Go
Python
Shell scripting
Linux/UNIX administration
Infrastructure automation
CI/CD tooling

Education

Bachelor’s or Master’s degree in CS/Engineering or related field

Tools

Docker
Kubernetes
Prometheus/Grafana
Elasticsearch
Ansible
Git

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer (SRE) - Engineering Productivity based in India.

As a Site Reliability Engineer, you will help build and operate the infrastructure that powers software engineering and product development at scale.

You will design secure, resilient, and highly available systems across hybrid cloud environments.

The role combines software engineering, infrastructure, automation, observability, and production operations.

You will work closely with development teams to remove technical bottlenecks and improve the overall developer experience.

You will have ownership of critical systems, from deployment and monitoring through incident response and continuous improvement.

The environment is highly engineering-focused, collaborative, and designed to encourage learning and experimentation with modern technologies.

This is an opportunity to make a direct impact on engineering productivity while working with large-scale, business‑critical infrastructure.

Accountabilities
  • Design, build, deploy, and operate critical production systems with a strong focus on scalability, reliability, observability, performance, and security.
  • Develop automation that reduces operational toil and improves the efficiency of production systems and engineering workflows.
  • Monitor infrastructure and services proactively, improve alerting, and implement automated responses where appropriate.
  • Create, maintain, and continuously improve incident response procedures and operational runbooks.
  • Build and deploy new systems incrementally using staged rollout practices to minimize operational risk.
  • Investigate and resolve platform and infrastructure issues, supporting software engineering teams with technical triage and troubleshooting.
  • Collaborate with third-party vendors when required to diagnose and resolve infrastructure or platform-related issues.
  • Write post-incident reviews and implement corrective measures to prevent recurring incidents.
  • Plan and communicate production maintenance activities and maintenance windows.
  • Partner with product development teams to identify infrastructure-related bottlenecks and design effective solutions.
  • Implement fault‑tolerance, performance improvements, and scaling strategies to increase system availability and resilience.
  • Research and adopt infrastructure and platform best practices to maintain secure, scalable, and fault‑tolerant environments.
  • Develop a strong understanding of open‑source and industry‑standard systems to improve troubleshooting, diagnosis, and issue resolution.
  • Support and enhance the overall developer experience across internal platforms and services.
Requirements
  • Hold a Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or possess equivalent professional experience, with approximately 5+ years of relevant experience.
  • Have working knowledge of Go, Python, and/or shell scripting, with the ability to develop medium‑complexity automation workflows.
  • Demonstrate strong Linux or UNIX administration and debugging capabilities.
  • Have hands‑on experience operating software systems, infrastructure, or complex applications at scale.
  • Have experience with server provisioning, particularly from storage and networking perspectives.
  • Possess strong software troubleshooting, analytical, and problem‑solving skills.
  • Have practical experience with infrastructure‑as‑code and automated infrastructure management.
  • Be comfortable working across multiple technologies and adapting quickly to new tools and platforms.
  • Experience with databases such as MariaDB, PostgreSQL, MongoDB, or similar technologies is desirable.
  • Experience with Docker, virtualization, or related technologies such as KVM, QEMU, or Kata Containers is a plus.
  • Experience managing monitoring and observability platforms such as Prometheus, Loki, Tempo, InfluxDB, Grafana, or Thanos is beneficial.
  • Experience managing Elasticsearch clusters, Artifactory, container registries, or similar infrastructure platforms is advantageous.
  • Familiarity with CI/CD technologies such as ArgoCD or Spinnaker is desirable.
  • Experience with version‑control systems such as Perforce or Gerrit is a plus.
  • Experience with infrastructure‑as‑code frameworks such as Ansible is beneficial.
  • Experience managing large Java applications or storage infrastructure such as NAS, SAN, or Ceph is an advantage.
  • Demonstrate strong communication and collaboration skills, with the ability to work effectively with software engineers, infrastructure teams, and external technology partners.
Benefits
  • Opportunity to work on large‑scale, business‑critical infrastructure and engineering productivity systems.
  • Exposure to a broad range of modern cloud, automation, observability, CI/CD, database, storage, and infrastructure technologies.
  • Significant ownership and autonomy over engineering projects and technical solutions.
  • Collaborative environment with opportunities to work across different technical domains and development teams.
  • Hybrid cloud engineering experience across scalable and fault‑tolerant systems.
  • Opportunities for continuous learning, experimentation, and adoption of infrastructure best practices.
  • Engineering‑focused culture that emphasizes technical excellence, automation, quality, and innovation.
  • Opportunity to contribute to systems and tools that directly improve the productivity and experience of software development teams.
  • Access to a globally distributed engineering environment with opportunities for cross‑functional and international collaboration.
Data Privacy Notice

By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre‑contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Arch Systems • Hyderabad

On-site
INR 2,800,000 - 4,200,000
Site Reliability Engineer
Site Reliability Engineer

SourcingXPress • Mumbai

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer SRE DevOps Engineer
Site Reliability Engineer SRE DevOps Engineer

New Era Technology • India

On-site
INR 1,800,000 - 3,000,000
Competitive benefits
Continuous training
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

New Era Technology • Delhi

On-site
INR 1,200,000 - 1,800,000
Site Reliability Engineer
Site Reliability Engineer

New Era Technology • Gurugram District

On-site
INR 1,400,000 - 2,000,000
Site Reliability Engineer 3 (Remote - India)
Site Reliability Engineer 3 (Remote - India)

Jobgether • India

Remote
INR 1,800,000 - 2,500,000
Competitive salary and bonuses
Flexible remote work
Paid time off and wellbeing days
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Lead SRE
Lead SRE

UST • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Right Advisors • Gurgaon

On-site
INR 1,200,000 - 2,100,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Employ • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Remote-first
Flexible scheduling
Paid time off
+1