Senior AI-Driven Reliability Engineer

Oracle

Honolulu (HI)

Remote

USD 102,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
Paid time off
Holidays

Job summary

Oracle Health is seeking a seasoned Site Reliability Engineer to own reliability for large-scale, cloud-native healthcare platforms. This hands-on role emphasizes resilient distributed systems, automation, and intelligent ops tooling for AI-powered services at scale.

You will lead cross-team reliability improvements, mentor engineers, and drive observability, deployment safety, and incident response. Strong Kubernetes and Linux troubleshooting skills are essential.

Qualifications

  • 7–10+ years of experience in Site Reliability Engineering, DevOps, Production Engineering, or related infra roles.
  • Experience operating large-scale production systems with high availability requirements.
  • Deep understanding of distributed systems and cloud infrastructure.
  • Hands-on Kubernetes and containerized workloads experience.
  • Strong automation and software engineering skills.
  • Experience improving operational excellence through tooling and rigor.

Responsibilities

  • Lead reliability engineering for large-scale cloud-native platforms.
  • Design and operate highly available distributed systems for AI-driven services.
  • Build automation, self-healing systems, and intelligent tooling.
  • Drive improvements in scalability, observability, deployment safety, and incident response.
  • Lead production investigations and durable long-term fixes.
  • Develop AIOps capabilities including anomaly detection and predictive scaling.
  • Partner with teams to improve architecture, resiliency, and readiness.
  • Mentor engineers and raise operational engineering maturity.

Skills

Kubernetes
Distributed systems
Automation
Observability
Linux troubleshooting
Incident response
Cross-team leadership
CI/CD
Infrastructure as code

Education

Bachelor's degree in a technical field

Tools

CI/CD
Infrastructure as Code
Cloud platforms

Job description

Oracle Health is seeking a seasoned Site Reliability Engineer to own reliability for large-scale, cloud-native healthcare platforms. This hands-on role emphasizes resilient distributed systems, automation, and intelligent ops tooling for AI-powered services at scale.

You will lead cross-team reliability improvements, mentor engineers, and drive observability, deployment safety, and incident response. Strong Kubernetes and Linux troubleshooting skills are essential.

Get your free, confidential resume review.

or drag and drop your file here.