Site Reliability / DevOps Engineer - Python & Kubernetes - Gurgaon - 20 LPA - Immediate JOINER

datavruti

Gurugram District

On-site

INR 1,200,000 - 2,000,000

Full time

14 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

datavruti in Gurgaon seeks a hands-on Site Reliability Engineer (SRE) / DevOps Engineer to drive automation and reliability. The role starts with development and automation work, then expands to CI/CD, cloud infrastructure, observability, and incident management across production systems.

The candidate should be proficient with Python development, Docker, Kubernetes, and cloud platforms, and capable of improving reliability through tooling and automation.

Qualifications

  • Hands-on Site Reliability Engineer with strong programming and automation skills.
  • Role involves development and automation, moving towards broader DevOps and SRE responsibilities including CI/CD, cloud infra, observability, incident management, and automation.
  • Ideal candidate comfortable with application code and production systems and using engineering to improve reliability.

Responsibilities

  • Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.
  • Write clean, production-ready code with proper testing and reviews.
  • Build, maintain, and improve CI/CD pipelines and deployment processes.
  • Work with Docker and Kubernetes for deployment and operations.
  • Support on-prem and cloud deployments and ensure reliable production environments.
  • Implement monitoring, logging, alerting, and observability solutions.
  • Contribute to defining SLIs, SLOs, and error budgets; perform RCA when issues occur.
  • Address recurring problems through automation and engineering improvements.

Skills

Python automation
REST APIs
Linux/Unix
Networking concepts
SRE concepts (SLIs/SLOs/RCA)
Cloud platforms (AWS/Azure/GCP)

Tools

Docker
Kubernetes
CI/CD (GitHub Actions/GitLab CI/Jenkins/Azure DevOps)
Terraform
Grafana/Prometheus

Job description

Hiring for: A US-based, AI-first data solutions company founded by seasoned technology leaders.

Positions: 1

Experience: 3 to 8 years

Location(s): Gurgaon

Type: On-site / Permanent

Salary: Up to INR 20 LPA (Based on fitment)

Notice Period: 15 days

About the Role

We are looking for a hands‑on Site Reliability Engineer (SRE) / DevOps Engineer with strong programming and automation skills.

The role will initially involve development and automation work, helping the engineer build a strong understanding of the applications and platform. Over time, the role will expand into broader DevOps and SRE responsibilities, including CI/CD, cloud infrastructure, observability, production reliability, incident management, and operational automation.

The ideal candidate should be comfortable working with both application code and production systems and should use engineering and automation to improve reliability and reduce manual effort.

Key Responsibilities
  • Develop and enhance internal applications, automation tools, APIs, utilities, and platform capabilities using Python.
  • Write clean, maintainable, testable, and production-ready code.
  • Participate in code reviews, debugging, testing, and technical discussions.
  • Build, maintain, and improve CI/CD pipelines and automated deployment processes.
  • Work with Docker and Kubernetes for application deployment and operations.
  • Support on prem and cloud-based application and infrastructure deployments.
  • Maintain reliable, scalable, secure, and highly available production environments.
  • Implement and manage monitoring, logging, alerting, and observability solutions.
  • Contribute to defining and tracking SLIs, SLOs, and error budgets.
  • Troubleshoot application and production issues and perform Root Cause Analysis (RCA).
  • Identify recurring operational problems and address them through automation and engineering improvements.
  • Support incident response, change management, deployment governance, and disaster recovery practices.
  • Maintain runbooks, SOPs, incident documentation, and technical documentation.
  • Collaborate with Engineering, Product, Platform, Security, Operations, and external teams.
Technical Skills
  • Strong hands‑on experience with Python for development and automation.
  • Experience developing scripts, APIs, integrations, utilities, or backend services.
  • Good understanding of software engineering principles, debugging, logging, testing, and exception handling.
  • Experience with REST APIs, JSON, Git, pull requests, and code reviews.
  • Strong knowledge of Linux/Unix environments and basic Windows administration.
  • Good understanding of networking concepts including DNS, TCP/IP, HTTP/HTTPS, load balancing, and firewalls.
  • Experience with at least one cloud platform: AWS, Azure, or GCP.
  • Hands‑on experience with Docker and Kubernetes.
  • Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or equivalent.
  • Familiarity with Infrastructure-as-Code tools such as Terraform is preferred.
  • Experience with monitoring and observability tools such as Grafana, Prometheus, Power BI, or equivalent.
  • Ability to analyze logs, metrics, alerts, and traces for troubleshooting.
  • Understanding of SRE concepts including SLIs, SLOs, availability, reliability, error budgets, and RCA.
  • Experience with JIRA, ServiceNow, and Confluence is desirable.
Preferred Experience
  • 3–6 years of experience in SRE, DevOps, Platform Engineering, Cloud Engineering, or related roles.
  • Experience supporting applications across development, deployment, and production environments.
  • Exposure to cloud-native, distributed, or production‑grade systems.
  • Understanding of security and compliance best practices.
  • Familiarity with AI‑assisted engineering tools such as GitHub Copilot, Claude Code, or similar tools.
Soft Skills
  • Strong analytical and troubleshooting skills.
  • Engineering and automation mindset.
  • Good written and verbal communication skills.
  • Effective cross‑functional collaboration.
  • Ownership‑driven approach to problem solving.
  • Ability to remain structured during production incidents.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

New Era Technology • Gurugram District

On-site
INR 1,400,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Right Advisors • Gurgaon

On-site
INR 1,200,000 - 2,100,000
Site Reliability Engineer
Site Reliability Engineer

MNR Solutions Pvt. Ltd. • Bengaluru

On-site
INR 900,000 - 1,500,000
Site Reliability Engineer
Site Reliability Engineer

Tech Mahindra • Bengaluru

On-site
INR 1,500,000 - 2,200,000
Senior Site Reliability Engineer (SRE) / DevOps Engineer
Senior Site Reliability Engineer (SRE) / DevOps Engineer

Umanist Staffing LLC • Maharashtra

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

United States Digital Space LLC • Gurugram District

Hybrid
INR 3,500,000 - 7,000,000
Staff Site Reliability Engineer
Staff Site Reliability Engineer

United States Digital Space LLC • Bengaluru

Hybrid
INR 6,000,000 - 12,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Site Reliability Engineer
Site Reliability Engineer

ScaleneWorks People Solutions LLP • Pune District

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

F-Prime Capital • Pune District

On-site
INR 1,500,000 - 2,000,000