[Remote] Principal Site Reliability Developer- USC Required

Ll Oefentherapie

United States

Remote

USD 86,400 - 199,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Flexible Vacation

Job summary

A leading technology company in the United States is looking for a Site Reliability DevOps Engineer to manage the Clinical AI Assistant platform, ensuring its reliability and performance. The role demands strong experience in Site Reliability Engineering and DevOps, with a focus on automation and orchestration. Candidates should possess excellent skills in container orchestration, infrastructure as code, and scripting. The position offers a competitive salary and comprehensive benefits, including medical, dental, and flexible vacation options.

Qualifications

  • 8+ years of experience in Site Reliability Engineering, DevOps, or related roles.
  • Experience operating large-scale, distributed, production systems.
  • Experience supporting AI/ML or LLM-based systems in production is a plus.

Responsibilities

  • Own the architecture, design, implementation, and production operations.
  • Ensure the reliability and performance of the Clinical AI Assistant platform.
  • Build and operate AIOps-driven capabilities.

Skills

Site Reliability Engineering
DevOps
Container orchestration (Kubernetes, Docker)
Infrastructure as Code (Terraform, Ansible)
Scripting and automation (Bash, Python)
Linux systems expertise

Tools

CI/CD pipelines (Git, Jenkins)
Observability tooling (monitoring, logging)
Major cloud provider (OCI, AWS)

Job description

Overview

Come and join us! Building on our cloud momentum, Oracle has formed a new organization—Oracle Health. This team focuses on product deployment, sustainability, troubleshooting, and product strategy while building a modern, automated healthcare platform. This is a net-new line of business with an entrepreneurial spirit, offering a unique opportunity to help build a world-class engineering organization centered on excellence, innovation, and real-world impact.

As a Site Reliability DevOps Engineer, you will play a critical role in operating and scaling a Clinical AI Assistant platform used by healthcare professionals worldwide. This system is designed to improve the quality, safety, and efficiency of care delivery for billions of patients globally. Your work will directly influence the reliability and performance of AI-driven systems that clinicians depend on in high-stakes environments.

This role goes beyond traditional SRE responsibilities—you will have the opportunity to leverage AI/ML techniques and develop AIOps solutions to proactively manage system reliability, detect anomalies, automate remediation, and continuously improve service performance. You will help define how reliability engineering evolves in the context of intelligent, AI-powered healthcare systems.

You will be responsible for architecture, production operations, capacity planning, performance management, deployment, and release engineering, working across cross-functional teams to deliver highly reliable, scalable, and secure services.

Responsibilities
  • Own the architecture, design, implementation, and production operations of core platform and AI-driven system services

  • Ensure the reliability, availability, and performance of the Clinical AI Assistant platform used in real-world healthcare settings

  • Build and operate AIOps-driven capabilities (e.g., intelligent alerting, anomaly detection, automated remediation, predictive scaling)

  • Continuously improve systems through automation, self-healing mechanisms, and real-time observability

  • Design and develop software to enhance system scalability, efficiency, and resilience

  • Partner with cross-functional teams to prototype and deliver new platform services

  • Lead efforts in capacity planning, demand forecasting, performance tuning, and cost optimization

  • Solve complex distributed systems challenges in cloud-native environments and prevent recurrence through engineering rigor

  • Contribute to platform engineering best practices, including infrastructure as code, CI/CD, and service reliability standards

  • Stay current with emerging technologies in cloud, distributed systems, and AI/ML-driven operations

Key Requirements / Experience

Must-have:

  • Ability to obtain and maintain a federal security clearance (US citizenship required)

  • 8+ years of experience in Site Reliability Engineering, DevOps, or related roles

  • Proven experience operating large-scale, distributed, production systems with high availability requirements

  • Strong experience with container orchestration (Kubernetes, Docker, or similar)

  • Infrastructure as Code expertise (Terraform, Ansible, Chef, Puppet, Packer, etc.)

  • Experience building and operating CI/CD pipelines (Git, Jenkins, GitLab, Rundeck, etc.)

  • Proficiency in scripting and automation (Bash, Python, PowerShell, etc.)

  • Experience with at least one major cloud provider (OCI, AWS, Azure, etc.)

  • Strong Linux systems expertise

  • Experience with observability tooling (monitoring, logging, tracing) and performance optimization

Nice-to-have:

  • Experience supporting or operating AI/ML or LLM-based systems in production

  • Exposure to AIOps, intelligent automation, or ML-driven observability

  • Experience in healthcare or other regulated environments (HIPAA, security, compliance)

  • Background in high-throughput, low-latency systems supporting mission-critical workloads

  • Software engineering experience in Java, Python, C++, or similar languages

Benefits
  1. US: Hiring Range in USD from: $86,400 to $199,500 per annum. May be eligible for bonus and equity.

  2. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

  3. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

  4. Oracle US offers a comprehensive benefits package which includes the following:

    1. Medical, dental, and vision insurance, including expert medical opinion

    2. Short term disability and long term disability

    3. Life insurance and AD&D

    4. Supplemental life insurance (Employee/Spouse/Child)

    5. Health care and dependent care Flexible Spending Accounts

    6. Pre-tax commuter and parking benefits

    7. 401(k) Savings and Investment Plan with company match

    8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

    9. 11 paid holidays

    10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

    11. Paid parental leave

    12. Adoption assistance

    13. Employee Stock Purchase Plan

    14. Financial planning and group legal

    15. Voluntary benefits including auto, homeowner and pet insurance

    The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

[Remote] Senior Site Reliability Developer- USC Required
[Remote] Senior Site Reliability Developer- USC Required

Ll Oefentherapie • United States

Remote
USD 79,000 - 159,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Paid parental leave
+2
[Remote] Senior Site Reliability Developer- USC Required
[Remote] Senior Site Reliability Developer- USC Required

Oracle • United States

On-site
USD 79,000 - 159,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Paid time off
+2
Senior Site Reliability Engineer - Oracle Health (US CITIZEN)
Senior Site Reliability Engineer - Oracle Health (US CITIZEN)

Oracle • United States

On-site
USD 81,000 - 187,000
Medical, dental and vision insurance
401(k) Savings and Investment Plan
Paid time off and holidays
+3
Principal Application Engineer
Principal Application Engineer

Oracle • Frankfort (KY)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
Paid time off
401(k) matching
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Oracle • Nashville (TN)

On-site
USD 84,900 - 209,500
Medical Insurance
Dental Insurance
Vision Insurance
+4
Lead Principal Site Reliability Engineer (SRE)
Lead Principal Site Reliability Engineer (SRE)

Oracle • United States

On-site
USD 105,000 - 264,000
Senior Application Software Engineer- Oracle Health
Senior Application Software Engineer- Oracle Health

Oracle • United States

On-site
USD 87,000 - 187,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Oracle • Reston (VA)

On-site
USD 85,000 - 210,000
Medical, dental, and vision insurance
Paid time off
401(k) Savings Plan
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Oracle • Austin (TX)

On-site
USD 84,000 - 210,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle • Reston (VA)

On-site
USD 81,200 - 187,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
+2