SRE- Production Support

Cloudxtreme

Hyderabad

On-site

INR 2,400,000 - 3,600,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Cloudxtreme is seeking an experienced SRE to drive reliability across large enterprise applications hosted on cloud platforms. You will lead automation initiatives, collaborate with architects, and implement proactive monitoring and incident response to maintain 24/7 availability.

The role requires strong experience in cloud-based production systems, extensive monitoring tool usage, and seamless collaboration with Dev and Support teams to minimize risk and ensure timely data delivery to clients.

Qualifications

  • Bachelor's degree in science or engineering, with 6+ years enterprise experience.
  • Proven track record leading complex enterprise production applications on cloud platforms (GCP/AWS/Azure).
  • Experience with monitoring tools: Splunk, AppDynamics, Grafana, ThousandEyes, Extrahop, Datadog, Prometheus.
  • Experience with CI/CD pipelines and Agile methodologies.
  • Excellent verbal and written communication.

Responsibilities

  • Practice Site Reliability Engineering (SRE) and solve problems through automation and innovation, augment SRE principles and guidelines.
  • Partner with Architects and SMEs to ensure production resiliency in designs.
  • Identify opportunities to build innovative tools and solve operations problems on large enterprise apps.
  • Real-Time troubleshooting of mission critical workflows and participate in On-call escalations.
  • Oversee Nightly batch runs to avoid SLA breaches and ensure timely data delivery to clients.
  • Collaborate with dev teams during design, build and perform infrastructure upgrades for availability and reliability.
  • Monitor portfolio to identify deficiencies, improve compliance, and enhance risk posture.
  • Partner with Support to rollout telemetry and reduce defect leakage for production delivery.

Job description

Role & responsibilities:

  • PracticeSite ReliabilityEngineering(SRE)andsolve problems through automation and innovation,augmentSRE principles and guidelines.
  • Partner with the Architects and SMEs in AST to ensure implementations are architected and designed from the aspect of production resiliency.
  • Identifyopportunities to build innovative tools and solve unique operations problems on large enterprise and mission critical applications.
  • Real-Time troubleshooting of mission critical application workflowsand participate in Oncall escalations.
  • Oversee Nightly batch runs to make sure no SLA breaches are encountered and Data files to our client firms are delivered on time.
  • Work closely with dev teams during design phase, build and perform infrastructure upgrades to support our applications availability and reliability.
  • Monitor the current-state solution portfolio to identify deficiencies through aging of the technologies used by the application, improve our compliance procedures and enhance our risk posture.
  • Partner within the Support organizations to build and rollout plans for enhanced telemetry and reduce defect leakage for software delivery to production.

Technical Experience and Skill Set

  • Bachelors degree in science or engineering,with6+ years of enterprise experience with listed technical skills.
  • Proven track record leading complex enterprise production applications in GCP based platforms (Vertex AI, Cloud run, BigQuery, Dataflow, Composer, Looker)or in AWS, Azure.
  • Experience in Monitoring tools- Splunk, AppDynamics,Grafana, Thousand Eyes, Extrahop, Datadog, Prometheus.
  • Database experience with Oracle, SQL Server, Mongo DB, PostgreSQL, Aerospike.
  • Experience in High Availability systems, Data Networks, NAS and Networking, IP Routing, leveraging tools to instrument and automate proactively and eventually predictive availability solutions.
  • Experience with Atlassian-Jira, Confluence, Bamboo, Bitbucket, GitHub, familiarity with CICD pipelines and Agile methodologies.
  • Preferred experience with C#,.Net,JAVA and Scripting, and hands on experience with IBM MQ, Kafka, RabbitMQ.
  • Proven capability to provide operational visibility on environmental health to Senior Leadership, Technology and Business partners.
  • Receptive, approachable teammate, with the ability to positively interact with Business partners, Technology teams and Vendor Professional services.
  • Excellent verbal and written communication
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE with Application Support Engineer
SRE with Application Support Engineer

Cloudxtreme • Hyderabad

On-site
INR 1,800,000 - 3,000,000
VS01700 - SRE & Production Reliability Engineer
VS01700 - SRE & Production Reliability Engineer

E4 Software Services Pvt Ltd. • India

On-site
INR 2,000,000 - 4,000,000
Production Support Engineer (6 Years)
Production Support Engineer (6 Years)

Cloudxtreme • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
SRE Lead
SRE Lead

Hdfc Securities • Mumbai

On-site
INR 3,500,000 - 5,500,000
Production Support Lead
Production Support Lead

Cloudxtreme • Hyderabad

On-site
INR 3,200,000 - 6,000,000
SRE Lead
SRE Lead

Acldigital • Ahmedabad District

On-site
INR 1,500,000 - 2,000,000
Application Support Engineer
Application Support Engineer

Cloudxtreme • Hyderabad

On-site
INR 1,800,000 - 3,000,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Site Reliability Engineer (SRE) – GCP Platform
Site Reliability Engineer (SRE) – GCP Platform

ITC Infotech • Bengaluru

On-site
INR 900,000 - 1,300,000