Cloud Reliability Engineer - Scale, Security, Performance

Oracle

United States

On-site

USD 74,000 - 148,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Paid time off
401(k) Savings and Investment Plan

Job summary

Oracle is seeking a Site Reliability Engineer to help build and operate resilient Oracle Cloud Health Data Intelligence Platform services. You will design, deploy, and optimize large-scale distributed systems with a focus on reliability, performance, and security.

The role blends developer instincts with systems expertise and requires clear communication to align with cross-functional teams, customers, and leadership.

Responsibilities

  • Service Ownership – You will be part of the SRE team, whose mission is the shared full stack ownership of a collection of services and/or technology areas, with our Development partners.
  • Ownership Scope – As an SRE, you will understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of the production services you own. In partnership with your Development partners, you will have the responsibility to ensure that services are designed and delivered to be critical with focus on security, resiliency, scale, and performance. SREs are the ultimate authority and are accountable for the end-to-end performance and operability of the services they own.
  • Service Design – As the Oracle Cloud evolves; you will partner with development teams in defining and implementing improvements in service architecture, both current and future. As an SRE, you will be a guide at articulating technical characteristics of your services and the dependencies between services, and guide Development teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. As an SRE, you will support federal project submission process and security compliance for new platforms and system resources.
  • Operations Engineering – You will understand and be able to communicate the scale, capacity, security, performance attributes and requirements of the services you own. You are a domain guide, able to understand and communicate every characteristic of your service stack, such as:
  • degradation and behavior under load of the services and their dependencies
  • end-to-end tuning needs, optimizing resource utilization, as load patterns fluctuate
  • Instrumentation and metrics that clearly describe the service behaviors
  • scaling requirements and patterns
  • resiliency and recoverability, ensuring that backup / restore and disaster recovery capabilities are implemented, tested and maintained
  • Security operations and vulnerability remediation, verifying vulnerabilities are patched or remediated while conforming to corporate and federal security standards and processes.
  • Automation – You will have a clear understanding of automation and orchestration principles, and will be eager to automate, wherever and whenever the possibility arises, while simultaneously eliminating technical debt. Automation must be part of your DNA.
  • Prevention - Once you have authoritatively resolved an issue, you will immediately work on how to more quickly resolve the problem next time, with the goal to eventually prevent the problem happening ever again
  • Technical Experts - As service owner, you are the ultimate partner concern point for complex or critical issues that have not yet been documented as SOPs for Level1 staff. You will usually get called in during major incidents as an SME, when the source of a problem is unclear. You will have the deep understanding of service topology and their dependencies required to solve issues and define mitigations.
  • Broad Interests - SREs are a rare mix of sysadmins and development Engineers, and as such have the ability to understand and explain the affect of product architecture decisions on the ability to run as distributed systems. They are driven by professional curiosity and a desire to a develop deep understanding of the their services and the technologies they depend upon.
  • Represent SRE - Proactive, self-motivated, customer-focused, organized, and a good communicator. SRE can be expected to represent Cloud products and engineering in critical forums.

Job description

Oracle is seeking a Site Reliability Engineer to help build and operate resilient Oracle Cloud Health Data Intelligence Platform services. You will design, deploy, and optimize large-scale distributed systems with a focus on reliability, performance, and security.

The role blends developer instincts with systems expertise and requires clear communication to align with cross-functional teams, customers, and leadership.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Cloud SRE Engineer: Reliability, Scale & Automation
Cloud SRE Engineer: Reliability, Scale & Automation

Oracle • San Juan (PR)

On-site
USD 74,000 - 148,000
Health insurance
Employee stock purchase plan
Paid time off
Security-Cleared SRE — Cloud Reliability Engineer
Security-Cleared SRE — Cloud Reliability Engineer

Oracle • United States

On-site
USD 81,000 - 187,000
Medical, dental, and vision insurance
Paid time off
401(k) Savings and Investment Plan
+3
Senior SRE — Cloud Reliability, Automation & Equity
Senior SRE — Cloud Reliability, Automation & Equity

Oracle • United States

On-site
USD 81,000 - 187,000
Medical, dental and vision insurance
401(k) Savings and Investment Plan
Paid time off and holidays
+3
Senior Cloud Site Reliability Engineer — Automation
Senior Cloud Site Reliability Engineer — Automation

Ll Oefentherapie • Richmond (VA)

On-site
USD 120,000 - 180,000
Senior Site Reliability & Infrastructure Leader
Senior Site Reliability & Infrastructure Leader

Oracle • United States

On-site
USD 121,000 - 265,000
Medical insurance
Dental insurance
Vision insurance
+7
Senior Site Reliability Engineer – Cloud Reliability & Automation
Senior Site Reliability Engineer – Cloud Reliability & Automation

Ll Oefentherapie • Reston (VA)

On-site
USD 85,000 - 210,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+2
Cloud Infra Reliability & Quality Architect (Equity Options)
Cloud Infra Reliability & Quality Architect (Equity Options)

Oracle • Frankfort (KY)

On-site
USD 202,000 - 250,000
Medical, dental, and vision insurance
Short/Long term disability
Life insurance and AD&D
+2
Site Reliability Engineer
Site Reliability Engineer

Ll Oefentherapie • Richmond (VA)

On-site
USD 120,000 - 180,000
Senior Site Reliability Engineering Lead
Senior Site Reliability Engineering Lead

Oracle • Frankfort (KY)

On-site
USD 122,000 - 264,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Vienna (VA)

On-site
USD 96,000 - 265,000
Medical, dental, vision insurance
401(k) with company match
Paid time off and holidays
+1