Site Reliability Engineer (SRE)- Python

Apexon

Bengaluru

On-site

INR 1,100,000 - 1,500,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Apexon is seeking an experienced SRE to join our Bengaluru-based operations. You will focus on operational tasks, incident handling, RCA, and collaboration with L1, Engineering, and Product teams to improve reliability.

You will automate and build tooling using Python, Shell, and JS, monitor systems, maintain runbooks, and train the L1 team on common issues while promoting shift-left practices.

Qualifications

  • 5-6 years of SRE/production support experience.
  • Strong scripting and automation skills with Python and shell.
  • Experience with monitoring, logging and observability tools.
  • Ability to work cross-functionally with Support, Engineering and Product teams.
  • Excellent communication and problem-solving capabilities.

Responsibilities

  • Provide support for client reported user issues.
  • Lead incident responses, perform RCA and provide on-call support.
  • Automation (AI) and tooling using scripting (JS, Shell, Python etc.).
  • Track recurring issues and contribute to Problem Management.
  • Contribute to bug fixes and enhancements.
  • Embed reliability practices and support production releases.
  • Escalate unresolved technical problems to Engineering.
  • Maintain records of support interactions and resolution runbooks.
  • Educate & train L1 team on common issues and shift-left initiative.

Skills

Python
Troubleshooting
Communication
Collaboration
Continuous learning

Tools

Oracle SQL DynamoDB
Cloud AWS
Automation AI
Terraform
MongoDB
Kafka
RESTful API

Job description

Experience: 5-6 years

The SRE team member will primarily handle operational tasks and troubleshooting, steering clear of direct software code changes to maintain efficiency and reduce risk.



  • Deliver responsive and knowledgeable assistance for routine and novel technical challenges

  • Act as a liaison between L1, and Engineering teams, ensuring effective and prompt issue resolution

  • Work with the Product and Engineering teams to drive any improvements


Scope of Work


  • Provide support for client reported user issues

  • Lead Incident responses, perform RCA and provide on-call support

  • Automation (AI) and tooling using scripting (JS, Shell, Python etc.)

  • Track recurring issue, contribute to Problem Management

  • Contribute to bug fixes and enhancements

  • Embed Reliability practices & support production releases

  • Escalate unresolved technical problems to Engineering team

  • Maintain records of support interactions and resolution runbooks

  • Educate & train L1 team on common issue and help in Shift-left initiative


Skill set


  • Technical proficiency: Strong Proficiency in Python, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.), Cloud Concepts / AWS and Automation AI.

  • Good to have Skills: Terraform, MongoDB, Kafka, RESTful API

  • Able to leverage AI tools.

  • Troubleshooting expertise: Strong analytical and problem-solving skills to diagnose issues with experience in monitoring tools (CloudWatch) as well as the knowledge of logging and observability practices.

  • Communication: Excellent written and verbal communication skills to effectively assist users, document support interactions, and collaborate with engineering and product teams.

  • Attention to detail: Ability to maintain accurate records of support cases and resolutions, ensuring compliance with internal policies and protocols.

  • Collaboration: Experience working cross-functionally with Support, Business Operations and Engineering teams to elevate unresolved issues and drive improvements in SRE support processes.

  • Continuous learning: Willingness to stay updated on evolving technologies and security threats to provide expert guidance and support.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer (SRE) – Core IT Infrastructure
Site Reliability Engineer (SRE) – Core IT Infrastructure

TECEZE • Chennai District

On-site
INR 1,000,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Lloyds Technology Centre • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer 2
Site Reliability Engineer 2

GreyOrange • Gurugram District

Hybrid
INR 2,600,000 - 4,600,000
Site Reliability Engineer
Site Reliability Engineer

New Era Technology • Gurugram District

On-site
INR 1,400,000 - 2,000,000
Site Reliability Engineering (SRE)
Site Reliability Engineering (SRE)

Lyzr AI • Bengaluru

Hybrid
INR 1,000,000 - 2,000,000
Site Reliability Engineer
Site Reliability Engineer

Epam Systems • Bengaluru

Hybrid
INR 4,000,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Gemini Solutions • Gurugram District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineering Lead
Site Reliability Engineering Lead

Infosys • Hyderabad

On-site
INR 1,200,000 - 1,800,000