Site Reliability Engineering Architect

Oracle

Frankfort (KY)

On-site

USD 146,000 - 306,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Paid time off
401(k) with company match
Employee Stock Purchase Plan
Voluntary benefits

Job summary

Oracle is seeking a Reliability & Quality Engineering leader to strengthen design quality and system reliability of large-scale data center electrical distribution architectures. You will evaluate end-to-end power systems, identify failure points, and develop resiliency concepts for fault containment and blast-radius reduction.

The role requires deep reliability and quality engineering experience across mission-critical power infrastructure and data center design platforms.

Qualifications

  • 10 years of experience in reliability engineering, quality engineering, or equivalent high-availability environments.
  • Strong systems-thinking with ability to assess electrical architectures under failure scenarios.
  • Deep understanding of data center electrical distribution concepts and power infrastructure.
  • Experience identifying failure modes, common-mode vulnerabilities, and blast-radius concerns.
  • Proven ability to assess availability impact using structured engineering methods.
  • Relevant product quality and reliability experience across power infrastructure or data center design.

Responsibilities

  • Evaluate end-to-end data center electrical distribution architectures and interfaces.
  • Identify design-level failure points, single points of failure, and protection coordination concerns.
  • Assess performance under credible failure scenarios including faults and transfer events.
  • Develop system-level resiliency concepts to improve fault isolation and blast radius reduction.
  • Translate reliability findings into design standards, product requirements, and acceptance criteria.
  • Apply methods such as FMEA, fault-tree analysis, reliability block diagrams, and root-cause analysis.
  • Evaluate product quality and reliability across power infrastructure and data center platforms.
  • Recommend improvements to specifications, supplier qualification, testing, and design-readiness criteria.
  • Create executive-ready reliability risk assessments communicating availability impact and mitigations.
  • Establish reliability KPIs, defect taxonomies, and governance for construction and handoff.
  • Champion safety, compliance, and operational excellence in high-uptime environments.

Skills

Reliability engineering
Quality engineering
Systems thinking
Cross-functional leadership
Executive communication
FMEA
Root-cause analysis
Power distribution
Data center infrastructure
Safety compliance

Job description

Job Description

Oracle Cloud Infrastructure is designing, standardizing, and operating mission-critical data center electrical infrastructure at extraordinary scale. We are hiring a Reliability & Quality Engineering leader to strengthen the design quality, system reliability, and availability of large-scale data center electrical distribution architectures.

Early in the role, the priority will be evaluating the full data center electrical distribution system, identifying potential failure points, assessing how those risks could impact availability, and developing system-level resiliency concepts such as failure isolation, fault containment, and blast radius reduction. This role requires someone who can look beyond component-level reliability and understand how the overall electrical architecture behaves under credible failure scenarios.

The ideal candidate will bring deep reliability and quality engineering capability across mission-critical power infrastructure, electrical equipment, or standardized data center design products. This person will partner closely with design engineering, construction, commissioning, operations, equipment suppliers, and executive stakeholders to convert reliability analysis into durable design standards, product improvements, risk mitigations, and measurable availability outcomes.

Responsibilities
What you’ll do
  • Evaluate end-to-end data center electrical distribution architectures, including utility or behind-the-meter interfaces, substations, medium-voltage distribution, switchgear, UPS systems, generators, BESS where applicable, protection systems, controls, and downstream power delivery.

  • Identify design-level failure points, single points of failure, common-mode risks, hidden dependencies, protection coordination concerns, and failure modes that could materially impact availability.

  • Assess how electrical systems perform under credible failure scenarios, including equipment faults, transfer events, protection operations, control-system failures, degraded-mode operation, maintenance conditions, and abnormal grid or generation events.

  • Develop system-level resiliency concepts and design recommendations that improve fault isolation, recoverability, maintainability, failure containment, and blast radius reduction.

  • Partner with electrical design engineering, commissioning, operations, construction, supply chain, and equipment vendors to translate reliability findings into design standards, product requirements, test expectations, acceptance criteria, and corrective action plans.

  • Apply quality and reliability engineering methods such as FMEA, fault-tree analysis, reliability block diagrams, root-cause analysis, design-for-reliability reviews, lessons-learned integration, and field-performance trend analysis.

  • Evaluate product quality and reliability across power infrastructure, electrical equipment, standardized electrical products, and repeatable data center design platforms.

  • Recommend improvements to specifications, supplier qualification, equipment testing, design validation, manufacturing quality controls, inspection processes, and supplier feedback loops.

  • Create executive-ready reliability risk assessments that clearly communicate availability impact, technical trade-offs, mitigation options, residual risk, and prioritization across large-scale infrastructure programs.

  • Establish reliability KPIs, defect taxonomies, quality gates, issue review cadence, and design-readiness criteria that improve reliability before construction, commissioning, and operational handoff.

  • Champion safety, compliance, disciplined change management, and operational excellence while improving reliability at the architecture, product, and system levels.

What you’ll bring
  • 10 years of experience in reliability engineering, quality engineering, electrical design assurance, product quality, mission-critical power systems, data center infrastructure, industrial power, utilities, generation, transmission and distribution, or equivalent high-availability environments.

  • Strong systems-thinking capability, with demonstrated ability to evaluate electrical architectures under failure scenarios rather than only assessing individual components or equipment ratings.

  • Deep understanding of data center electrical distribution concepts, including medium-voltage and low-voltage distribution, substations, switchgear, UPS systems, generators, protection schemes, controls, grounding, power quality, redundancy, maintainability, and operational recovery.

  • Experience identifying failure modes, common-mode vulnerabilities, cascading risks, hidden dependencies, and blast-radius concerns across complex electrical systems.

  • Proven ability to assess availability impact and reliability tradeoffs using structured engineering methods such as FMEA, fault-tree analysis, reliability block diagrams, event analysis, root-cause analysis, or similar methodologies.

  • Relevant product quality and reliability experience across power infrastructure, electrical equipment, standardized electrical products, or repeatable data center design platforms.

  • Ability to influence design standards, supplier requirements, equipment qualification expectations, commissioning acceptance criteria, and operational readiness deliverables.

  • Strong cross-functional leadership skills, with experience partnering across design engineering, construction, commissioning, operations, vendors, and executive stakeholders.

  • Excellent communication and executive reporting skills, with the ability to translate complex technical reliability risk into clear decisions, priorities, and implementation plans.

  • Commitment to safety, compliance, disciplined engineering governance, and operational excellence in high-uptime environments.

Disclaimer:

Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only

US: Hiring Range in USD from: $146,300 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle’s differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  1. Medical, dental, and vision insurance, including expert medical opinion

  2. Short term disability and long term disability

  3. Life insurance and AD&D

  4. Supplemental life insurance (Employee/Spouse/Child)

  5. Health care and dependent care Flexible Spending Accounts

  6. Pre-tax commuter and parking benefits

  7. 401(k) Savings and Investment Plan with company match

  8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.

  9. 11 paid holidays

  10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.

  11. Paid parental leave

  12. Adoption assistance

  13. Employee Stock Purchase Plan

  14. Financial planning and group legal

  15. Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level – IC6
About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Architect
Site Reliability Engineering Architect

Oracle • San Juan (PR)

On-site
USD 146,000 - 306,000
Medical insurance
Paid time off
401(k)
Sr. Principal Data Center Electrical Product Development Engineer
Sr. Principal Data Center Electrical Product Development Engineer

Oracle • United States

On-site
USD 146,000 - 306,000
Medical insurance
Dental insurance
401(k) plan
+1
Senior Electrical Reliability Engineer — Critical Power
Senior Electrical Reliability Engineer — Critical Power

Oracle • Nashville (TN)

Hybrid
USD 126,200 - 264,100
Medical, dental, and vision insurance
401(k) with company match
Paid time off and holidays
+1
Facilities Technician – Electrical
Facilities Technician – Electrical

Oracle • United States

On-site
USD 87,000 - 179,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan with company match
Flexible vacation policy
+1
Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Oracle • Frankfort (KY)

On-site
USD 146,000 - 306,000
Architect, Modular Infrastructure & Industrialized Construction
Architect, Modular Infrastructure & Industrialized Construction

Ll Oefentherapie • United States

On-site
USD 161,000 - 339,000
Sr. Principal Data Center Electrical Product Development Engineer
Sr. Principal Data Center Electrical Product Development Engineer

Oracle • Frankfort (KY)

On-site
USD 146,000 - 306,000
Medical, dental, vision insurance
Paid time off
401(k) with company match
+3
Principal Design Quality & Reliability Engineer – OCI Data Center Infrastructure
Principal Design Quality & Reliability Engineer – OCI Data Center Infrastructure

Oracle • United States

On-site
USD 146,300 - 306,400
Medical, dental, and vision insurance
401(k) plan with company match
Paid time off
+1
Sr. Principal Design Quality & Reliability Engineer – OCI Data Center Infrastructure
Sr. Principal Design Quality & Reliability Engineer – OCI Data Center Infrastructure

Oracle • Frankfort (KY)

On-site
USD 202,000 - 250,000
Medical, dental, and vision insurance
Short/Long term disability
Life insurance and AD&D
+2
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Oracle • Nashville (TN)

On-site
USD 84,900 - 209,500
Medical Insurance
Dental Insurance
Vision Insurance
+4