Site Reliability Engineer

Bounteous

Bengaluru

On-site

INR 1,200,000 - 2,000,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A leading company in IT services is seeking a Service Reliability Engineer to ensure operational excellence and system performance. The role involves enhancing incident management, collaborating with engineering teams, and mentoring junior SREs. Candidates should possess strong technical skills in programming, observability tools, and container-based architecture. This full-time position emphasizes a collaborative culture and continuous improvement.

Qualifications

  • Mid-Senior level position focusing on reliability and performance.
  • Experience with automation, incident management, and system design.

Responsibilities

  • Enhance incident management processes and develop prioritization matrices.
  • Collaborate with engineering teams to improve system performance.
  • Mentor other SREs and contribute to their growth.

Skills

Python
Golang
Shell scripting
Ruby
Java
Communication
Problem Solving
Analytical Skills

Tools

Terraform
Ansible
Grafana
Prometheus
Kubernetes
AWS EKS
Docker

Job description

Direct message the job poster from Bounteous

As a Service Reliability Engineer (SRE) you will take a multifaceted approach to ensure technical excellence and operational efficiency within the infrastructure domain. Specializing in reliability, resilience and system performance, you take a lead role in championing the principles of Site Reliability Engineering. By strategically integrating automation, monitoring and incident response, you facilitate the evolution from traditional operations to a more customer-focused and agile approach. Emphasizing shared responsibility and a commitment to continuous improvement, you cultivate a collaborative culture, enabling organizations to meet and exceed their reliability and business objectives.

Job responsibilities

● You will be responsible for understanding requirements or SRE goals in depth from both tech and business perspectives

● You will provide solutions to improve reliability, including identifying and implementing mechanisms and architectures that enable fault tolerance and faster median time to respond and median time to detect

● You will be responsible for enhancing the incident management process, including the development of an incident prioritization matrix, triage, communication, mitigation, post-mortem analysis and implementation of corrective actions

● You will manage client stakeholder expectations and queries during production incidents, providing detailed technical analysis of issues and remediation plans for mitigation and prevention in future, and act as the interface for C-level executives, if or when needed

● You will be a liaison with client engineering teams, build trust and productive relationships with senior client stakeholders and team leads to influence them in making better decisions

● You will be responsible for identifying opportunities for enhancing system performance and reliability in alignment with business SLAs, SLOs, KPIs and objectives, and provide guidance and assistance to SRE teams in implementing the identified improvements

● As an SRE expert, you will collaborate with Thoughtworks application development leads and solution architects, recommending changes in system design and adopting best practices for improved reliability from day one

● You will oversee and mentor other SREs on the team, contributing to their growth and development

Job qualifications

Technical Skills

● You can program with one or more high-level languages such as Python, Golang, Shell scripting, Ruby or Java

● You are familiar with DevOps and GitOps practices, driving the integration of observability automation into CI/CD pipelines, e.g.: GitLab, Jenkins, CircleCI or equivalent

● You have in-depth knowledge of configuration management and Infrastructure as Code (IAC) tools such as Terraform, Ansible, ARM and CloudFormation for provisioning and managing infrastructure

● You have an expertise in observability, logs, tracing and monitoring tools such as Grafana (Loki and Tempo), Prometheus, Graylog, Jaeger, Zipkin, ELK stack or equivalent

● You have a strong understanding of container-based architecture and hands-on experience with orchestration tools such as Kubernetes, AWS EKS, Docker Swarm, Nomad, etc.

● You have in-depth experience in application and infrastructure performance tuning and scaling to handle heavy loads under different scenarios e.g.: Periodic traffic load and tsunami patterns

● You have a good understanding of essential concepts such as quality gates encompassing SLI/SLO/SLA, chaos engineering, golden signals, blameless postmortem methodologies, synthetic monitoring, distributed tracing, end-user monitoring and performance testing

● You have experience with network load balancing, security tech stacks, Transport Layer Security (TLS) and certificate management, and an understanding of standard networking protocols and configurations

Professional Skills

● You have strong communication and articulation skills, and are proficient in English

● You are able to convey resolutions to audiences with varying degrees of technical/business proficiency and bring them to consensus

● You have excellent problem-solving and analytical skills, with a focus on continuous improvement

● You have good listening and presentation skills

● You solve challenging problems and difficult to debug issues with a never give up attitude

● You can collaborate with cross-functional engineering teams to conduct capacity planning and scalability assessments, and design solutions for handling current and future growth

● You have the ability to work under pressure, with composure, during production incidents

● You understand requirements provided by the client on both technical and business aspects, and can break them down for successful implementation

● You’re willing to be part of a rotation- and need-based, 24x7 available team

Seniority level
  • Seniority level
    Mid-Senior level
Employment type
  • Employment type
    Full-time
Job function
  • Job function
    Information Technology
  • Industries
    IT Services and IT Consulting

Referrals increase your chances of interviewing at Bounteous by 2x

Sign in to set job alerts for “Site Reliability Engineer” roles.
Site Reliability Manager, Platforms and Devices, SRE
Bigbasket - Senior Software Development Engineer
Software Engineer - QA Automation (Entry Level)
Senior Site Reliability Engineer- Automation
Sr. Site Reliability Engineer- Spera (ISMP)

We’re unlocking community knowledge in a new way. Experts add insights directly into each article, started with the help of AI.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Lead Site Reliability Engineer – Technical & People Leadership
Sr. Lead Site Reliability Engineer – Technical & People Leadership

Shell Recharge Solutions • Bengaluru

On-site
INR 2,000,000 - 2,500,000
Flexible scheduling
Generous holiday package
Medical benefits
Senior Site Reliability Engineering Manager
Senior Site Reliability Engineering Manager

Seven N Half • Bengaluru

On-site
INR 2,000,000 - 2,500,000
Associate Engineer II - Site Reliability Engineer [T500-17376]
Associate Engineer II - Site Reliability Engineer [T500-17376]

ANSR • Hyderabad

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer (US Timezone)
Site Reliability Engineer (US Timezone)

Skyflow • India

Remote
INR 3,000,000 - 5,500,000
Work from home expense
Excellent Health Insurance Options
Very generous PTO
+2
Site Reliability Engineer
Site Reliability Engineer

NetApp • Bengaluru

Hybrid
INR 800,000 - 1,200,000
Health care benefits
Educational assistance
Volunteer time off
Lead Site Reliability Engineer [T500-19357]
Lead Site Reliability Engineer [T500-19357]

Deutsche Börse • Hyderabad

On-site
INR 1,500,000 - 2,500,000
Site Reliability Engineer
Site Reliability Engineer

Infosys • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Employ • Bengaluru

Hybrid
INR 4,000,000 - 6,000,000
Remote-first
Flexible scheduling
Paid time off
+1
Site Reliability Engineer
Site Reliability Engineer

Persistent Systems • Hyderabad

On-site
INR 600,000 - 1,800,000
Competitive salary
Quarterly promotion cycles
Company-sponsored education and certifications
+3
Site Reliability Engineer
Site Reliability Engineer

LSEG • Bengaluru

On-site
INR 600,000 - 900,000
Competitive salary and benefits
Opportunities for learning and career development
Paid volunteering days
+1