Site Reliability Engineer III

Emburse

Hyderabad

On-site

INR 2,800,000 - 4,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Emburse is seeking an experienced Site Reliability Engineer III to ensure highly available, scalable, and performant systems. The role emphasizes automation, cloud infrastructure, observability, and leading reliability initiatives across distributed environments.

You will mentor junior engineers, drive incident prevention, and align SRE practices with product and engineering roadmaps to improve system resilience.

Qualifications

  • Minimum 6 years of engineering or operations experience focused on reliability and automation.
  • Strong proficiency in Linux-based distributed environments with hands-on experience.
  • Extensive experience with cloud platforms (AWS or Azure) and infrastructure-as-code (Terraform).
  • Excellent scripting skills (Python, Bash, PowerShell); OO programming a plus.
  • Experience building internal tools and automation solutions.
  • Strong English communication and collaboration skills with offshore or distributed teams.
  • Expertise in containerization/orchestration (Docker, Kubernetes) and modern CI/CD pipelines.
  • Experience with observability stacks (Prometheus, Grafana, OpenTelemetry).
  • Background in SaaS or large-scale distributed apps; analytical, proactive, ownership, and mentoring mindset.

Responsibilities

  • Ensure services are designed and operated with 24/7 availability, scalability, and resilience.
  • Proactively identify and implement measures to reduce customer impact.
  • Design, develop, and automate reliable cloud infrastructure and platform services.
  • Apply IaC to manage large-scale distributed systems and automate workflows.
  • Mentor SRE I & II engineers and lead RCA/postmortems for continuous improvement.
  • Lead cross-functional troubleshooting spanning apps, infra, databases, and networks.

Skills

Linux distributed environments
AWS
Azure
Terraform
Python
Bash
PowerShell
Docker
Kubernetes
CI/CD
OpenTelemetry
Prometheus

Education

Bachelor’s degree in Computer Science or STEM field

Tools

Terraform

Job description

The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s systems are highly available, scalable, and performant. This role blends deep technical expertise with strong collaboration and leadership skills to drive operational excellence across distributed systems. The ideal candidate is passionate about automation, cloud infrastructure, observability, and continuous improvement, while mentoring junior engineers and driving reliability culture across the organization.

Essential Functions
Service Reliability & Performance
  • Proactively identify, evaluate, and implement preventative measures to reduce customer impact.
  • Ensure all services are designed and operated with 24/7 availability, scalability, and resilience in mind.
  • Monitor, troubleshoot, and provide visibility to improve site latency, performance, and up time.
Engineering Excellence & Automation
  • Design, develop, and automate reliable cloud infrastructure and platform services.
  • Apply Infrastructure-as-Code (IaC) principles to manage large-scale distributed systems.
  • Write and maintain scripts, tools, and automation frameworks to support operational efficiency.
  • Partner with engineering leadership to develop solutions enabling developer productivity and remove cross functional dependencies.
Collaboration & Process Development
  • Collaborate with Platform Engineering teams on project definitions, requirements, backlog grooming, and planning processes.
  • Align operational goals with product and engineering roadmaps to ensure reliability requirements are met early in the lifecycle.
  • Define non-functional requirements (NFRs) and influence standards for scalability, observability, and fault tolerance.
  • Lead cross-functional troubleshooting of complex issues spanning applications, infrastructure, databases, and networks.
  • Serve as a technical mentor to SRE I and II engineers, guiding them in best practices for reliability, automation, and incident management.
  • Lead root cause analysis and postmortem reviews, driving continuous improvement initiatives.
  • Support offshore and distributed teams, promoting effective collaboration and communication.
  • Participate in design and architecture reviews, offering technical recommendations and documentation for key stakeholders.
Education and Experience

Education: Required: Bachelor’s degree in Computer Science or STEM field

  • Experience: Minimum 6 years of experience in an engineering or operations role with a focus on reliability, scalability, and automation.
Additional Eligibility Qualifications
Required Skills:
  • Strong proficiency in Linux-based distributed environments (up to 70%
  • hands-on work).Deep experience with cloud platforms (AWS or Azure) and Infrastructure-as-Code (Terraform).
  • Excellent scripting skills (Python, Bash, Powershell); object-oriented programming experience is a plus.
  • Demonstrated ability to develop and maintain internal tools and automation solutions.
  • Excellent written and verbal communication skills in English.Strong project management and organizational abilities with a bias for action.
  • Experience collaborating with offshore or globally distributed teams.
  • Expertise in containerization and orchestration technologies (Docker, Kubernetes).
  • Strong understanding of DevOps principles and modern CI/CD pipelines.Experience with observability stacks (Prometheus, Grafana,OpenTelemetry).Familiarity with self-healing systems, and site reliability best practices.
  • Background in SaaS environments or large-scale distributed applications.Analytical thinker with a focus on root-cause problem solving.Self-starter with a strong ownership mentality and accountability.
  • Mentor and collaborator who uplifts teams and promotes learning culture.
  • Committed to operational excellence and continuous improvement.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer II
Site Reliability Engineer II

United States Digital Space LLC • Karnataka

On-site
INR 800,000 - 1,200,000
Site Reliability Engineer
Site Reliability Engineer

InOpTra Digital • Bengaluru

On-site
INR 1,200,000 - 2,000,000
Site Reliability Engineer (SRE) / DevOps Engineer
Site Reliability Engineer (SRE) / DevOps Engineer

New Era Technology • Gurugram District

On-site
INR 1,500,000 - 2,100,000
Senior Associate Site Reliability Engineer
Senior Associate Site Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Hyderabad

On-site
INR 1,400,000 - 2,000,000
Site Reliability Engineering Lead_Truist
Site Reliability Engineering Lead_Truist

Infosys • Bengaluru

On-site
INR 4,000,000 - 7,000,000
Site Reliability Engineer (SRE) – Core IT Infrastructure
Site Reliability Engineer (SRE) – Core IT Infrastructure

TECEZE • Chennai District

On-site
INR 1,000,000 - 2,000,000
SRE Engineer @ Investment Banking | Mumbai
SRE Engineer @ Investment Banking | Mumbai

Net Connect Global • Bengaluru, Mumbai

Hybrid
INR 1,800,000 - 2,400,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Zorba AI • Chennai District

On-site
INR 1,200,000 - 2,400,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

UST • Thiruvananthapuram

On-site
INR 1,200,000 - 1,600,000
Sr. Site Reliability Engineer I
Sr. Site Reliability Engineer I

MetLife • Hyderabad

On-site
INR 1,800,000 - 2,400,000