Director Cloud Operations

WellSpan Health

York (York County)

On-site

USD 180,000 - 230,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

WellSpan Health seeks a Cloud Service Reliability and Operations Leader to guide day-to-day management of cloud and on-premises infrastructure as the organization advances its cloud journey.

You will own incident response, change execution, and collaboration with SRE, DevOps, security, and vendors to maintain reliability, security, and cost awareness across services with standard operating procedures and audit-ready controls.

Qualifications

  • Bachelor's degree in IT, CS, Engineering or related field.
  • Minimum 7+ years in infrastructure/platform operations, with cloud experience.
  • At least 4+ years in people leadership in a matrixed environment.
  • Hands-on cloud experience (AWS/Azure/GCP), ITSM familiarity.
  • Strong incident, problem and change management discipline.
  • Proven ability to lead during outages and critical events.
  • Experience with observability tooling and IaC (Terraform/CFN).
  • Certifications such as AWS/SysOps or ITIL preferred.

Responsibilities

  • Lead day-to-day operations for cloud and on-premises infrastructure.
  • Own incident response, major incident leadership, and RCA.
  • Oversee change execution aligned with ITSM processes.
  • Partner with SRE/DevOps, security, and vendors to ensure reliability.
  • Drive automation, standardization and cost stewardship.
  • Mentor and develop CloudOps staff for 24x7 coverage.

Education

Bachelor's Degree in IT/CS/Engineering
7+ years in infrastructure/platform operations
4+ years people leadership
Cloud operations across AWS/Azure/GCP
Incident & change management discipline
Strong communication under pressure
Cross-functional collaboration
Cost stewardship & FinOps

Tools

Terraform
CloudFormation
CI/CD for infrastructure
Monitoring/Observability tooling
ITIL/ITSM processes

Job description

Leads the day-to-day operational management of the organization's cloud and on-premises environment as it progresses on its cloud journey. Is accountable for cloud service reliability, operational readiness, incident response, change execution, and continual improvement across cloud and on-premises infrastructure and platform services. Partners closely with the Cloud Center of Excellence (CCoE), SRE, DevOps, Network, Security, Enterprise Architecture, application teams, and vendors to ensure cloud services are secure, resilient, cost-aware, and delivered with consistent operational standards.

Duties and Responsibilities:
Cloud Service Reliability and Operations Leadership:
  • Leads daily operations for cloud infrastructure and core services (compute, storage, network, identity integrations, monitoring/logging, backup/DR enablement).
  • Establishes an operations-first culture focused on stability, customer communication, and rapid recovery from outages.
  • Owns operational readiness for new cloud capabilities and migrated workloads, ensuring production support is prepared before go-live.
Incident, Problem, and Major Event Management:
  • Owns cloud-related incident response, escalation, and coordination (including major incident leadership as needed).
  • Ensures clear runbooks, on-call processes, and escalation paths are defined and practiced.
  • Drives root cause analysis, problem management, and corrective action plans to reduce repeat incidents and operational risk.
Change and Release Execution (Cloud Platform):
  • Manages change execution for cloud platform services, aligning with ITSM change processes while enabling speed and reliability.
  • Ensures change planning, risk assessment, approvals, and post-change validation are performed consistently.
  • Improves change success rate through standard change patterns, automation, and pre/post deployment checks.
Monitoring, Observability and Performance:
  • Partners with SRE/Tools teams to implement and mature monitoring, logging, alerting, and dashboards for cloud services and critical workloads.
  • Tracks and reports service health metrics (availability, performance trends, MTTR, incident volume).
Automation and Standardization ("Paved Roads"):
  • Drives automation to reduce manual work and improve repeatability (provisioning, patching, tagging, backup policies, configuration drift detection).
  • Establishes standard operating procedures and supported reference patterns for common cloud services.
  • Collaborates with CCoE and Engineering to build self-service capabilities and standardized service catalogs.
Security, Compliance and Guardrails:
  • Ensures cloud operations align with security policies and controls (least privilege, logging, segmentation, vulnerability remediation support).
  • Partners with Security to operationalize guardrails (policy-as-code where applicable), respond to findings, and improve posture over time.
  • Ensures audit-ready operational evidence (change traceability, access reviews support, logging/retention practices).
FinOps Partnership and Cost Stewardship:
  • Partners with FinOps/Finance to improve cost visibility and control through tagging compliance, right-sizing, scheduling, and elimination of waste.
  • Monitors usage patterns and identifies optimization opportunities. Tracks and reports cost savings/avoidance initiatives.
Migration Support and Cutover Readiness:
  • Supports migration waves by ensuring operational prerequisites are complete (monitoring, backups, DR expectations, access, runbooks, support model).
  • Participates in cutover planning, go/no-go readiness assessments, and hypercare support.
  • Coordinates with vendors/partners and internal teams to resolve cutover issues quickly.
Vendor and Service Provider Management:
  • Manages cloud operations vendors and managed services partners: performance management, SLAs/OLAs, issue escalation, and service reviews.
  • Ensures third-party delivered services meet reliability, security, and customer experience expectations.
People Leadership and Team Development:
  • Hires, coaches, and develops CloudOps staff. Sets clear expectations and builds a culture of ownership and continuous improvement.
  • Ensures skills development aligned to cloud platform needs (training, certifications, mentoring).
  • Builds coverage models that support 24x7 needs where required while maintaining sustainable on-call practices.
Qualifications:
  • Bachelors Degree in IT, Computer Science, Engineering, or related field required.
  • 7+ years of experience in infrastructure/platform operations, cloud operations, or SRE/DevOps-adjacent roles required.
  • 4+ years of people leadership or proven experience leading operational teams in a matrixed environment required.
  • Demonstrated experience operating production environments with strong incident and change management discipline required.
  • Hands-on cloud experience (AWS/Azure/GCP), including networking, identity, security logging, and core platform services; Familiarity with ITSM/ITIL processes (incident/problem/change) and integrating cloud operations into enterprise ITSM workflows; Experience with observability tooling (monitoring/logging/alerting) and on-call operations; Exposure to Infrastructure as Code and automation (e.g., Terraform/CloudFormation, CI/CD for infrastructure) preferred
  • Certifications: AWS SysOps Administrator/Solutions Architect, ITIL Foundation, Security+ or equivalent preferred.
  • Strong communication skills and the ability to lead under pressure during outages and critical events.
  • Strong cross-functional collaboration and vendor management.
  • Automation-first thinking and standardization.
  • Security-conscious operations with audit readiness.
  • Cost awareness and continuous improvement discipline.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Manager-Cloud Operations
Manager-Cloud Operations

WellSpan Health • York

On-site
USD 90,000 - 120,000
Comprehensive health benefits
Retirement savings plan
Paid time off (PTO)
+3
Cloud Operations Architecture Lead
Cloud Operations Architecture Lead

Jobtailor • Portsmouth (NH)

On-site
USD 140,000 - 180,000
Principal Cloud and Production Operations Engineer
Principal Cloud and Production Operations Engineer

Jobtailor • California (MO)

On-site
USD 140,000 - 190,000
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies • Lakewood (CO)

On-site
USD 120,000 - 150,000
Manager of Cloud Platform Operations
Manager of Cloud Platform Operations

Love's Travel Stops • Oklahoma City (OK)

On-site
USD 120,000 - 180,000
Tuition assistance
Paid time off
401(k) 100% match up to 5%
+3
Cloud Solutions Engineer
Cloud Solutions Engineer

Tyler Technologies, Inc. • Plano (TX), Latham (NY), Lubbock (TX), Lakewood (CO)

On-site
USD 93,547 - 150,000
Sr. Cloud Operations Reliability Engineer (SRE)
Sr. Cloud Operations Reliability Engineer (SRE)

NextGen Healthcare • Georgia

On-site
USD 140,000 - 210,000
Cloud Devops Engineer
Cloud Devops Engineer

Gerdau North America • Tampa (FL)

On-site
USD 100,000 - 150,000
Vice President, Head of Cloud Ops and Infrastructure
Vice President, Head of Cloud Ops and Infrastructure

Pailin Group Psc • California (MO)

On-site
USD 200,000 - 300,000
Manager, Cloud Operations & Engineering
Manager, Cloud Operations & Engineering

AgFirst Farm Credit Bank • Columbia (SC)

On-site
USD 100,000 - 130,000