Application SRE

Cloudxtreme

Pune District

On-site

INR 1,200,000 - 1,800,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cloudxtreme in Pune, Maharashtra, India is seeking a skilled SRE/Production Support professional to drive reliability improvements and manage complex incidents. You will design SLIs/SLOs, implement robust monitoring, and support cloud deployments across major platforms.

The role emphasizes DevOps practices, ITIL-aligned incident handling, and collaboration with cross-functional teams to ensure service levels are met and stakeholders are kept informed.

Qualifications

  • Experienced in SRE, Production Support, Incident Management, Release Management.
  • Good understanding of DevOps and Agile processes.
  • Proven ability to communicate with stakeholders and coordinate across teams.

Responsibilities

  • Drive improvements in system reliability through proactive monitoring, automation, and incident response.
  • Design and implement SLIs/SLOs and error budgets to measure service health.
  • Develop and maintain monitors, dashboards, and synthetic transactions using Datadog, Splunk, Grafana, or equivalent.
  • Use APM and telemetry tools to ensure applications meet service levels.
  • Troubleshoot distributed systems on cloud platforms; focus on Java and PostgreSQL.
  • Automate with Python/Go/Java/Shell; perform Linux admin tasks.
  • Support cloud apps across GCP, Azure, AWS, or PCF environments.
  • Manage production deployments and resolve deployment failures.
  • Collaborate with CI/CD tools (Bitbucket Pipelines, Bamboo, Harness, GitHub Actions).
  • Apply ITIL processes (Incident, Problem, Change, Release Management).
  • Use ITSM tools (JIRA, Remedy, ServiceNow) for incident tracking and service mgmt.
  • Coordinate with cross-functional teams and communicate clearly with stakeholders.

Skills

SRE
Production Support
Incident Management
Release Management
DevOps
Agile Process
Stakeholder Communication

Job description

Role & responsibilities



Drive improvements in system reliability through proactive monitoring, automation, and incident response.



  • Design and implement Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to measure and maintain service health.

  • Develop and maintain monitors, dashboards, and synthetic transactions using tools such as Datadog, Splunk, AppDynamics, Grafana, or equivalent.

  • Utilize Application Performance Monitoring (APM) and telemetry tools to ensure applications meet expected service levels.

  • Troubleshoot and debug distributed systems deployed on cloud platforms, primarily involving Java and PostgreSQL.

  • Perform Linux system administration and develop automation scripts using Python, Go, Java, Shell, or similar languages.

  • Provide support for cloud applications across GCP, Azure, AWS, or PCF environments.

  • Manage production deployments and resolve deployment failures efficiently.

  • Work with CI/CD tools such as Bitbucket Pipelines, Bamboo, Harness, or GitHub Actions to streamline development workflows.

  • Apply strong knowledge of ITIL processes including Incident, Problem, Change, and Release Management.

  • Use ITSM tools like JIRA, Remedy, ServiceNow, etc., for issue tracking and service management.

  • Collaborate effectively with cross-functional teams and communicate clearly with stakeholders.


Preferred candidate profile



Process Skills:


  • Experienced in SRE, Production Support, Incident Management, Release Management

  • Deep understanding in working on Agile process.

  • Good understanding in DevOps process

  • Contribute to Production Releases, Incident Triage & Resolution

  • Contribute to team documents, incident history, known-error database.

  • Prepare & review Impact Analysis of Incidents, status reports to stakeholders.


Behavioral Skills:


  • Effectively collaborates and communicates with stakeholders and ensure client satisfaction

  • Resolve technical issues of incidents and explore workarounds for immediate resolution, identify permanent solutions.

  • Participates as a team member and fosters teamwork by inter-group coordination within the modules of the project.

  • Work under the direction of SRE Architect

  • Develop model and set up proactive monitoring for transaction flows

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineering Lead (Application SRE Lead)
Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Developer II - DevOps Engineering
Developer II - DevOps Engineering

UST • Maharashtra

On-site
INR 1,800,000 - 3,400,000
Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer (SRE) – GCP Platform
Site Reliability Engineer (SRE) – GCP Platform

ITC Infotech • Bengaluru

On-site
INR 900,000 - 1,300,000
SRE Developer
SRE Developer

Cloudxtreme • Hyderabad

On-site
INR 1,500,000 - 2,400,000
SRE Lead
SRE Lead

Hdfc Securities • Mumbai

On-site
INR 3,500,000 - 5,500,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Infobell It Solutions • Hyderabad

On-site
INR 2,800,000 - 4,500,000