Senior Consultant - Site Reliability Engineer India Full Time Employment
Job Description
Job Description
Senior Consultant – Site Reliability Engineer | HCA Healthcare
GeneralPosition Information
Reports directly to (Title):
Manager– Site Reliability Engineer
Matrix reports to (Title):
As applicable based on functionalalignment
Direct Reports:
PositionSummary:
TheSenior Consultant - Site Reliability Engineering (SRE) is a hands-ontechnical leadership and consulting role responsible for driving thereliability, availability, performance, resilience, and operational maturityof enterprise and business-critical services. The role serves as a principaltechnical escalation point and trusted advisor for complex productionchallenges, shaping SRE strategy and engineering improvements acrossobservability, automation, Infrastructure-as-Code, incident management, reliabilityengineering, performance optimization, and operational readiness. The SeniorConsultant partners with Development, Architecture, SRE, DevOps, Cloud,Database, Network, Infrastructure, Security, and business stakeholders; leadscross-functional initiatives; and mentors senior engineers.
Responsibilities
- Serve as a principaltechnical escalation point and SRE consultant for complex, high-impact, andbusiness-critical incidents; lead cross-functional troubleshooting,executive-level technical communication, and service restoration.
- Troubleshoot acrossapplications, database, API/integration, middleware, cloud, server, network,identity, security, and external dependency layers using logs, metrics,traces, events, and infrastructure telemetry.
- Lead root cause analysisfor significant incidents and drive corrective and preventive actions thatreduce recurrence and operational risk.
- Define, govern, andmature SRE and observability practices, including SLIs/SLOs, error budgets,dashboards, alerting, instrumentation, event correlation, monitoringcoverage, and reliability reporting.
- Lead enterprisereliability, operational-readiness, and supportability assessments; identifygaps in resiliency, automation, observability, documentation, infrastructure,and deployment processes and develop prioritized multi-quarter improvementroadmaps.
- Lead strategictoil-reduction programs and design reusable automation, self-healing, andauto-remediation patterns that improve engineering productivity and servicereliability at scale.
- Provide technicalgovernance and hands-on leadership for Infrastructure-as-Code, configurationmanagement, GitOps, platform engineering, and CI/CD practices usingtechnologies such as Terraform, Ansible, Argo CD, Azure DevOps, GitHub, orGitLab.
- Troubleshoot complexdeployment, configuration, pipeline, rollback, and release failures andpartner with engineering teams to improve deployment reliability.
- Support major applicationupgrades, migrations, platform modernization, patching, environmenttransitions, and production cutovers.
- Analyze application andinfrastructure performance/capacity trends and recommend scaling, quota,configuration, resiliency, and cost-optimization improvements.
- Own technical directionfor disaster recovery, business continuity engineering, resiliencyvalidation, recovery objectives/procedures, failure testing, and rollbackreadiness.
- Partner with Security andengineering teams during critical vulnerabilities or cyber events and supportapplication, infrastructure, authentication, and configuration analysis.
- Establish SRE standards,reference architectures, playbooks, runbooks, templates, operatingprocedures, governance mechanisms, and reusable engineering patterns acrossteams.
- Evaluate emergingtechnologies and operating practices in observability, automation, cloudoperations, reliability engineering, platform engineering, and AI-assistedoperations; lead proof-of-concept evaluations, technical recommendations, andadoption roadmaps.
- Mentor senior SRE,Production Engineering, DevOps, and Application Support engineers; providetechnical coaching in troubleshooting, observability, automation, incidentmanagement, root cause analysis, architecture, and reliability practices.
- Use incident trends,operational metrics, SLO performance, risk indicators, and engineering datato define reliability priorities, influence stakeholders, and drivemeasurable continuous improvement.
Provide consultative leadership to application and platformteams on reliability architecture, SRE adoption, production readiness, cloudmodernization, and operational risk reduction.
- Lead reliability reviewswith senior stakeholders, translate technical risk into business impact, anddefine measurable remediation plans, success criteria, and governancecheckpoints.
- Participate in an on-callor senior production escalation rotation where required.
Education & Experience
- Bachelor’sdegree in Computer Science, Information Technology, Engineering, or a relateddiscipline preferred; equivalent advanced technical experience anddemonstrated SRE leadership may be considered.
- Relevantcertifications in Cloud, DevOps, SRE, Infrastructure-as-Code, Kubernetes,ITIL, Linux, Microsoft Azure, Google Cloud, or related areas are beneficial.
Must HaveSkills
- Application& Production Engineering - Advancedtroubleshooting across complex applications, integrations, dependencies, andproduction environments.
- Observability& Reliability - Advanced proficiency in Excel andpresentation tools; comfortable working with large data sets, pivots,lookups, and structured trackers.
- Cloud& Infrastructure - Strong understanding of cloudplatforms, operating systems, databases, networking, infrastructure services,and cross-platform dependencies.
- DevOps& Infrastructure-as-Code - Hands-on experiencewith CI/CD, Git/GitOps, Terraform, Ansible, deployment automation, andInfrastructure-as-Code practices.
- Automation& Engineering - Ability to design reusableautomation, reduce operational toil, and develop self-healing or remediationworkflows.
- Incident& Problem Management - Ability to leadcomplex incident troubleshooting, RCA, corrective actions, and preventiveimprovements.
- Performance,Capacity & Security - Ability to analyzeperformance/capacity trends and troubleshoot application-security, identity,access, and configuration issues with specialist teams.
- TechnicalLeadership & Improvement - Expert ability to operate as asenior SRE consultant, mentor experienced engineers, influence architectureand stakeholders, establish enterprise standards, lead transformationroadmaps, and drive measurable reliability improvements.