Sr. Cloud Operations Reliability Engineer (SRE)

NextGen Healthcare

Georgia

On-site

USD 140,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NextGen Healthcare is seeking a Senior Cloud Operations Reliability Engineer to drive operational excellence across cloud platforms and owned services. You will establish observability, define SLOs/SLIs, and lead incident response to improve availability and recovery.

Collaborating with engineering and security teams, you will implement IaC, automation, runbooks, and monitoring dashboards to reduce toil and strengthen production readiness.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field or equivalent experience.
  • 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or related discipline with ownership of production systems.
  • Hands-on experience with Google Cloud Platform (GCP), AWS, or equivalent cloud providers.
  • Expertise in monitoring, observability platforms, alerting, incident response, and root cause analysis.
  • Experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation) and version control.
  • Strong incident management and post-incident review skills with reliability improvements.
  • Experience with Kubernetes, containers, and orchestration.
  • Experience with application performance monitoring (APM) and distributed tracing.
  • Experience mentoring junior engineers or leading operational improvements.

Responsibilities

  • Own service reliability and operational health; set and monitor SLOs/SLIs and drive availability improvements.
  • Lead incident response and post-incident processes, including root cause analysis and remediation.
  • Design reliability-focused automation, tooling, and runbooks; apply IaC to support recovery and consistency.
  • Build observability through comprehensive monitoring, logging, and alerting; establish event correlation and escalation.
  • Analyze performance and capacity; recommend reliability-focused scaling and readiness improvements.
  • Partner with development teams to improve deployment readiness and production supportability.
  • Contribute to disaster recovery planning and operational readiness exercises.
  • Mentor engineers and maintain reliability standards and documentation.
  • Support cloud governance, security initiatives, access control, tagging, and audit readiness.
  • Perform other duties aligned with position objectives.

Skills

Observability
Incident response
IaC
Kubernetes
APM
SRE practices
CI/CD
Automation
Cloud platforms
Mentoring

Education

Bachelor's degree in Computer Science, Engineering, Information Systems, or related field

Tools

Grafana
Prometheus
Datadog
New Relic
Cloud Monitoring
Terraform
Deployment Manager
CloudFormation

Job description

Job Description:

The Senior Cloud Operations Reliability Engineer is responsible for driving operational excellence and strengthening the reliability posture of cloud-based services and supported platforms. This role owns critical reliability initiatives, establishes observability and service health practices, and leads incident response coordination to improve service availability, resiliency, and recovery. The Senior Cloud Operations Reliability Engineer partners with engineering, security, and operations teams to advance reliability practices, support and mature service-level objectives, improve production readiness, and develop reliability-focused automation that reduces operational toil and accelerates incident recovery.

  • Own service reliability and operational health-establish and maintain SLOs/SLIs, design monitoring and alerting strategies, and drive improvements that enhance service availability and performance across cloud platforms.
  • Lead incident response coordination and post-incident processes, including troubleshooting complex production issues, conducting root cause analysis, and driving remediation activities with accountability for timeline and resolution quality.
  • Design and implement reliability-focused automation, operational tooling, and runbooks to reduce manual toil, improve response consistency, and strengthen production readiness and resilience; apply Infrastructure as Code practices where appropriate to support recovery, reliability, and operational consistency.
  • Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures to ensure rapid problem detection and response.
  • Conduct performance and capacity analysis, evaluate utilization trends, identify bottlenecks; provide recommendations for reliability-focused scaling, performance improvement, capacity planning, and operational readiness of cloud-based services.
  • Partner with development and engineering teams to evaluate deployment readiness, support deployment reliability improvements, and implement operational best practices that strengthen service reliability, rollback readiness, and production supportability.
  • Contribute to disaster recovery and business continuity planning, conduct operational readiness exercises, and ensure recovery procedures and documentation reflect current production state and evolving business requirements.
  • Mentor team members and establish reliability standards and practices within Cloud Operations and supported service areas; create and maintain operational documentation, standard operating procedures, and knowledge base materials.
  • Support operational adherence to cloud governance, compliance, and security initiatives; including access control, tagging, logging, and audit readiness, and reliability-related documentation.
  • Perform other duties that support the overall objective of the position.
Education Required:
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.
  • Or, any combination of education and experience which would provide the required qualifications for the position.
Experience Required:
  • 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline with demonstrated ownership of production systems.
  • Extensive hands-on experience supporting production cloud environments using Google Cloud Platform (GCP), AWS, or equivalent cloud service providers.
  • Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
  • Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.) and version control best practices.
  • Strong background in incident management and post-incident review processes; experience driving corrective actions and establishing reliability improvements.
  • Experience with Kubernetes operations, containerization, and orchestration platforms.
  • Experience with application performance monitoring (APM) and distributed tracing.
  • Experience mentoring junior engineers or leading operational improvements initiatives.
License/Certification Required:
  • Google Cloud certifications: Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer.
  • AWS certification: AWS SysOps Administrator or equivalent.
  • Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering (SRE), or ITIL.
Knowledge, Skills & Abilities:
  • Knowledge of: Working knowledge of CI/CD practices, cloud governance, compliance frameworks, disaster recovery, and business continuity planning. Deep technical knowledge of Google Cloud Platform (GCP), AWS, or similar cloud providers; understanding of cloud-native services, networking, security, and compute models.Familiarity with observability tools such as Grafana, Prometheus, Cloud Monitoring, or similar platforms. Security operations, compliance auditing, or audit readiness processes, preferred.
  • Skill in: Hands‑on expertise with monitoring platforms (Datadog, New Relic, Prometheus, Cloud Monitoring, etc.); ability to design effective dashboards, alerts, and health checks. Proficiency in scripting languages (Python, Bash, Go, etc.) to develop automation solutions that reduce manual effort.
  • Ability to: Advanced ability to diagnose complex, multi-layered infrastructure issues and coordinate timely recovery. Ability to translate complex technical findings into actionable recommendations; experience influencing cross‑functional teams on reliability practices.

NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Sr. Cloud Operations Reliability Engineer (SRE)
Sr. Cloud Operations Reliability Engineer (SRE)

NextGen Healthcare Information Systems LLC • Georgia

Hybrid
USD 140,000 - 190,000
Senior Site Reliability Engineer NEX
Senior Site Reliability Engineer NEX

Patterson-UTI • Houston (TX)

On-site
USD 120,000 - 180,000
Manager of Site Reliability Engineering (SRE)
Manager of Site Reliability Engineering (SRE)

Genuine Parts Company • Alabama

On-site
USD 120,000 - 150,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Veloc Inc • Coppell (TX)

On-site
USD 140,000 - 190,000
Site Reliability Engineer
Site Reliability Engineer

JPS Tech Solutions • San Jose (CA)

On-site
USD 130,000 - 160,000
Site Reliability Engineer
Site Reliability Engineer

Compunnel, Inc. • New Jersey

On-site
USD 120,000 - 150,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000
Principal III, SRE
Principal III, SRE

United States Digital Space LLC • United States

Remote
USD 180,000 - 240,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Charlotte (NC)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Discretionary incentive eligibility
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Bank of America • Plano (TX)

On-site
USD 152,000 - 192,000
Industry-leading benefits
Paid time off
Access to resources and support