Senior Consultant - Site Reliability Engineer

Darwinbox Digital Solutions Pvt. Ltd.

Hyderabad

On-site

INR 3,000,000 - 5,200,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Darwinbox Digital Solutions Pvt. Ltd. in India is seeking a Senior Consultant - Site Reliability Engineer to drive reliability, performance, and operational maturity for critical services.

You will lead cross-functional teams, own incident response, and shape SRE practices across observability, automation, and IaC. The role requires deep hands-on leadership, expertise in CI/CD, Terraform/Ansible, and cloud platforms, with mentorship of other engineers and collaboration with Security and Platform

Qualifications

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline preferred; equivalent advanced technical experience and demonstrated SRE leadership may be considered.
  • Relevant certifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes, ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial.

Responsibilities

  • Serve as a principal technical escalation point and SRE consultant for complex, high-impact incidents; lead cross-functional troubleshooting, executive-level technical communication, and service restoration.
  • Troubleshoot across applications, database, API/integration, middleware, cloud, server, network, identity, security, and external dependency layers using logs, metrics, traces, events, and infrastructure telemetry.
  • Lead root cause analysis for significant incidents and drive corrective and preventive actions that reduce recurrence and operational risk.
  • Define, govern, and mature SRE and observability practices, including SLIs/SLOs, error budgets, dashboards, alerting, instrumentation, event correlation, monitoring coverage, and reliability reporting.
  • Lead enterprise reliability, operational-readiness, and supportability assessments; identify gaps in resiliency, automation, observability, documentation, infrastructure, and deployment processes and develop prioritized multi-quarter improvement roadmaps.
  • Lead toil-reduction programs and design reusable automation, self-healing, and auto-remediation patterns that improve engineering productivity and service reliability at scale.
  • Provide technical governance and hands-on leadership for Infrastructure-as-Code, configuration management, GitOps, platform engineering, and CI/CD practices using Terraform, Ansible, Argo CD, Azure DevOps, GitHub, or GitLab.
  • Troubleshoot complex deployment, configuration, pipeline, rollback, and release failures and partner with engineering teams to improve deployment reliability.
  • Support major application upgrades, migrations, platform modernization, patching, environment transitions, and production cutovers.
  • Analyze application and infrastructure performance/capacity trends and recommend scaling, quota, configuration, resiliency, and cost-optimization improvements.
  • Own technical direction for disaster recovery, business continuity engineering, resiliency validation, recovery objectives/procedures, failure testing, and rollback readiness.
  • Partner with Security and engineering teams during critical vulnerabilities or cyber events and support application, infrastructure, authentication, and configuration analysis.
  • Establish SRE standards, reference architectures, playbooks, runbooks, templates, operating procedures, governance mechanisms, and reusable engineering patterns across teams.
  • Evaluate emerging technologies and operating practices in observability, automation, cloud operations, reliability engineering, platform engineering, and AI-assisted operations; lead proof-of-concept evaluations, technical recommendations, and adoption roadmaps.
  • Mentor senior SRE, Production Engineering, DevOps, and Application Support engineers; provide technical coaching in troubleshooting, observability, automation, incident management, root cause analysis, architecture, and reliability practices.
  • Use incident trends, operational metrics, SLO performance, risk indicators, and engineering data to define reliability priorities, influence stakeholders, and drive measurable continuous improvement.

Skills

App & Prod Eng
Observability & Reliability
Cloud & Infra
DevOps & IaC
Automation & Eng
Incident & Problem Mgmt
Performance & Security
Tech Leadership

Education

Bachelor’s degree in CS/IT/Engineering
Cloud/DevOps certifications beneficial

Tools

Terraform
Ansible
Argo CD
Azure DevOps
GitHub
GitLab

Job description

Senior Consultant - Site Reliability Engineer India Full Time Employment

Job Description

Job Description

Senior Consultant – Site Reliability Engineer | HCA Healthcare

GeneralPosition Information

Reports directly to (Title):

Manager– Site Reliability Engineer

Matrix reports to (Title):

As applicable based on functionalalignment

Direct Reports:

PositionSummary:

TheSenior Consultant - Site Reliability Engineering (SRE) is a hands-ontechnical leadership and consulting role responsible for driving thereliability, availability, performance, resilience, and operational maturityof enterprise and business-critical services. The role serves as a principaltechnical escalation point and trusted advisor for complex productionchallenges, shaping SRE strategy and engineering improvements acrossobservability, automation, Infrastructure-as-Code, incident management, reliabilityengineering, performance optimization, and operational readiness. The SeniorConsultant partners with Development, Architecture, SRE, DevOps, Cloud,Database, Network, Infrastructure, Security, and business stakeholders; leadscross-functional initiatives; and mentors senior engineers.

Responsibilities
  • Serve as a principaltechnical escalation point and SRE consultant for complex, high-impact, andbusiness-critical incidents; lead cross-functional troubleshooting,executive-level technical communication, and service restoration.
  • Troubleshoot acrossapplications, database, API/integration, middleware, cloud, server, network,identity, security, and external dependency layers using logs, metrics,traces, events, and infrastructure telemetry.
  • Lead root cause analysisfor significant incidents and drive corrective and preventive actions thatreduce recurrence and operational risk.
  • Define, govern, andmature SRE and observability practices, including SLIs/SLOs, error budgets,dashboards, alerting, instrumentation, event correlation, monitoringcoverage, and reliability reporting.
  • Lead enterprisereliability, operational-readiness, and supportability assessments; identifygaps in resiliency, automation, observability, documentation, infrastructure,and deployment processes and develop prioritized multi-quarter improvementroadmaps.
  • Lead strategictoil-reduction programs and design reusable automation, self-healing, andauto-remediation patterns that improve engineering productivity and servicereliability at scale.
  • Provide technicalgovernance and hands-on leadership for Infrastructure-as-Code, configurationmanagement, GitOps, platform engineering, and CI/CD practices usingtechnologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, orGitLab.
  • Troubleshoot complexdeployment, configuration, pipeline, rollback, and release failures andpartner with engineering teams to improve deployment reliability.
  • Support major applicationupgrades, migrations, platform modernization, patching, environmenttransitions, and production cutovers.
  • Analyze application andinfrastructure performance/capacity trends and recommend scaling, quota,configuration, resiliency, and cost-optimization improvements.
  • Own technical directionfor disaster recovery, business continuity engineering, resiliencyvalidation, recovery objectives/procedures, failure testing, and rollbackreadiness.
  • Partner with Security andengineering teams during critical vulnerabilities or cyber events and supportapplication, infrastructure, authentication, and configuration analysis.
  • Establish SRE standards,reference architectures, playbooks, runbooks, templates, operatingprocedures, governance mechanisms, and reusable engineering patterns acrossteams.
  • Evaluate emergingtechnologies and operating practices in observability, automation, cloudoperations, reliability engineering, platform engineering, and AI-assistedoperations; lead proof-of-concept evaluations, technical recommendations, andadoption roadmaps.
  • Mentor senior SRE,Production Engineering, DevOps, and Application Support engineers; providetechnical coaching in troubleshooting, observability, automation, incidentmanagement, root cause analysis, architecture, and reliability practices.
  • Use incident trends,operational metrics, SLO performance, risk indicators, and engineering datato define reliability priorities, influence stakeholders, and drivemeasurable continuous improvement.

Provide consultative leadership to application and platformteams on reliability architecture, SRE adoption, production readiness, cloudmodernization, and operational risk reduction.

  • Lead reliability reviewswith senior stakeholders, translate technical risk into business impact, anddefine measurable remediation plans, success criteria, and governancecheckpoints.
  • Participate in an on-callor senior production escalation rotation where required.
Education & Experience
  • Bachelor’sdegree in Computer Science, Information Technology, Engineering, or a relateddiscipline preferred; equivalent advanced technical experience anddemonstrated SRE leadership may be considered.
  • Relevantcertifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes,ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial.
Must HaveSkills
  • Application& Production Engineering - Advancedtroubleshooting across complex applications, integrations, dependencies, andproduction environments.
  • Observability& Reliability - Advanced proficiency in Excel andpresentation tools; comfortable working with large data sets, pivots,lookups, and structured trackers.
  • Cloud& Infrastructure - Strong understanding of cloudplatforms, operating systems, databases, networking, infrastructure services,and cross-platform dependencies.
  • DevOps& Infrastructure-as-Code - Hands-on experiencewith CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, andInfrastructure-as-Code practices.
  • Automation& Engineering - Ability to design reusableautomation, reduce operational toil, and develop self-healing or remediationworkflows.
  • Incident& Problem Management - Ability to leadcomplex incident troubleshooting, RCA, corrective actions, and preventiveimprovements.
  • Performance,Capacity & Security - Ability to analyzeperformance/capacity trends and troubleshoot application-security, identity,access, and configuration issues with specialist teams.
  • TechnicalLeadership & Improvement - Expert ability to operate as asenior SRE consultant, mentor experienced engineers, influence architectureand stakeholders, establish enterprise standards, lead transformationroadmaps, and drive measurable reliability improvements.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal SRE
Principal SRE

Lloyds Banking Group • Hyderabad

Hybrid
INR 3,000,000 - 5,500,000
Site Reliability Engineering Lead (Application SRE Lead)
Site Reliability Engineering Lead (Application SRE Lead)

Hirexa Solutions • Bengaluru

Hybrid
INR 3,500,000 - 7,000,000
Site Reliability Engineer
Site Reliability Engineer

Spot Your Leaders & Consulting • Pune District

On-site
INR 2,500,000 - 4,000,000
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

Bahwan CyberTek • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Software Engineer
Software Engineer

PwC • Hyderabad, Bengaluru

Hybrid
INR 2,800,000 - 5,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Infosys • Bengaluru

On-site
INR 900,000 - 1,500,000
SRE Reliability Engineer
SRE Reliability Engineer

NTT DATA BUSINESS SOLUTIONS • Bengaluru

On-site
INR 2,500,000 - 4,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Director – Site Reliability Engineering (SRE)
Director – Site Reliability Engineering (SRE)

Umanist NA • Hyderabad

On-site
INR 4,000,000 - 7,000,000
Senior Site Reliability Engineer (SRE) Engineer
Senior Site Reliability Engineer (SRE) Engineer

Umanist Staffing • Pune District

On-site
INR 2,250,000 - 2,750,000