Site Reliability Engineer

SRM Digital

Maharashtra

On-site

INR 3,500,000 - 6,000,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

SRM Digital in Pune, India, seeks an experienced Production Operations Engineer to manage a comprehensive production environment. You will define monitoring strategies, respond to incidents, and drive automation across deployments to improve reliability and velocity.

The role involves supporting CI/CD pipelines, optimizing performance, and coordinating with a global team. Occasional on-call and off-hours work may be required as part of service ownership.

Qualifications

  • Experience with Java/Spring Boot in production
  • Knowledge of Kafka (Axon)
  • Experience with monitoring tools like Splunk or Dynatrace
  • Familiarity with CI/CD platforms and automation tooling
  • Understanding of SRE practices and cloud-native observability

Responsibilities

  • Plan, manage, and oversee all aspects of a Production Environment
  • Define strategies for Application Performance Monitoring and Production optimization
  • Respond to incidents and improve platform based on feedback
  • Automate deployment and support CI/CD pipelines across environments
  • Design and standardize monitoring and alerting mechanisms
  • Engage in lifecycle management from design to refinement
  • Analyze ITSM activities and feedback for resiliency gaps
  • Support pre-live services with capacity planning and launch reviews
  • Lead DevOps automation and best practices
  • Maintain service availability, latency, and health
  • Scale systems through automation and reliability improvements
  • Collaborate with global teams across time zones

Skills

Java / Spring Boot
Kafka
CI/CD Platforms
Cloud-native observability
Remedy / ServiceNow

Tools

Splunk
Dynatrace

Job description

Location: Pune, India(3 days work from office)

Duration: Permanent Role

The Role
  • Plan, manage, and oversee all aspects of a Production Environment
  • Define strategies for Application Performance Monitoring, Optimization in Production environment
  • Respond to Incidents and improvise platform based on feedback and measure the reduction of incidents over time.
  • Support deployment of code into multiple lower environments. Supporting current processes with an emphasis on automating everything as soon as possible.
  • Design, develop and standardize Monitoring and Alerting mechanism for the supported applications.
  • Take a holistic approach to problem solving, by connecting the dots during a production event through the various technology stack that makes up the platform, to optimize meantime to recover.
  • Engage in and improve the whole lifecycle of services—from inception and design, through deployment, operation and refinement.
  • Analyse ITSM activities of the platform and provide feedback loop to development teams on operational gaps or resiliency concerns.
  • Support services before they go live through activities such as system design consulting, capacity planning and launch reviews.
  • Support the application CI/CD pipeline for promoting software into higher environments through validation and operational gating, and lead in DevOps automation and best practices.
  • Maintain services once they are live by measuring and monitoring availability, latency, and overall system health.
  • Scale systems sustainably through mechanisms like automation and evolving systems by pushing for changes that improve reliability and velocity.
  • Work with a global team spread across tech hubs in multiple geographies and time zones.
  • Ability to share knowledge and explain processes and procedures to others.
  • Able to perform on-call duties on a rotational basis.
  • Occasional off hours work required.
Must Have
  • Java / Spring Boot - Knowledge is required
  • Kafka (Axon)
  • Splunk / Dynatrace
  • CI/CD Platforms
  • Cloud-native observability and SRE practices
  • Remedy / servicenow
Operations Skills:
  • Production support leadership
  • Runbook and support model creation
  • Monitoring and alert strategy definition
  • Disaster recovery and resiliency planning
  • Deployment readiness validation
  • Service ownership mindset
  • Continuous improvement and toil reduction
  • Root cause analysis and problem management
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Hyderabad

On-site
INR 1,800,000 - 3,000,000
Site Reliability Engineer
Site Reliability Engineer

Acesoft Labs • Ahmedabad District

Hybrid
INR 400,000 - 700,000
Software Engineer-DevOps
Software Engineer-DevOps

SMC Squared India • Bengaluru

On-site
INR 1,000,000 - 1,500,000
Site Reliability Engineering
Site Reliability Engineering

Maneva Consulting • Pune District

On-site
INR 1,500,000 - 2,400,000
Analyst II, Production Support
Analyst II, Production Support

fis • Pune District

On-site
INR 1,500,000 - 2,300,000
Specialist, Production Services Application Support Analyst II
Specialist, Production Services Application Support Analyst II

BNY Mellon • Pune District

On-site
INR 1,500,000 - 2,000,000
Site Reliability Engineer Lead
Site Reliability Engineer Lead

Hilabs • Pune District

On-site
INR 1,500,000 - 2,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Hdfc Bank • Bengaluru

On-site
INR 2,500,000 - 4,000,000