Principal Data Systems Software Engineer - SRE

Oracle Corporation

Seattle (WA)

On-site

USD 140,000 - 190,000

Full time

7 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Global scale projects
Distributed systems exposure
Reliability-focused culture
Career growth opportunities
Flexible work arrangements
Competitive compensation and benefits

Job summary

Oracle Corporation is seeking a Principal Site Reliability Engineer to own operational health, reliability, and continuous improvement of OCI GoldenGate and related services. You will ensure production cloud services run securely, with automation driving efficiency and consistent deployments.

The role requires leading incident response, building observability, and partnering with software teams to scale globally. A TS clearance eligibility and 7+ years of experience are expected.

Qualifications

  • 7+ years production engineering experience (distributed services or cloud platform experience preferred).
  • Automation experience and scripting (bash, Python, etc.).
  • Solid understanding of Software Engineering and Computer Science principles.
  • Solid foundation in Linux Systems Engineering.
  • Solid understanding of Networking, Security and Storage.
  • Experience with Docker.
  • Understanding Infrastructure as Code concept.
  • Experience with Terraform and/or Ansible.
  • Experience building fully automated CI/CD pipelines (Jenkins, Hudson, TeamCity).
  • Ability to formulate meaningful KPIs and business metrics.
  • Experience with cloud services at major providers (AWS, Azure, GCP or OCI) is a big plus.

Responsibilities

  • Operate and support production cloud services that power critical customer migration and data replication workloads.
  • Monitor service health, investigate incidents, and lead troubleshooting efforts for complex production issues.
  • Participate in incident response, root cause analysis, and post-incident reviews with corrective actions.
  • Act as escalation point for complex issues not yet documented as SOPs.
  • Partner with software development teams to improve service architecture and operational readiness.
  • Design and implement automation to reduce operational overhead and improve efficiency.
  • Build and maintain observability solutions: monitoring, alerting, logging, dashboards, metrics.
  • Contribute to capacity planning, performance analysis, disaster recovery readiness, and resilience initiatives.
  • Support services across Oracle Cloud environments; leverage IaC and CI/CD practices.

Skills

Automation
SRE principles
Distributed systems
Networking
Linux systems engineering
Security and storage

Education

Bachelor's degree or equivalent

Tools

Docker
Terraform
Ansible
Jenkins
Hudson
TeamCity

Job description

GoldenGate Service is a new multi-tenant, cloud native service for real-time data integration and replication in heterogeneous IT environments. GoldenGate enables users to replicate and integrate data from different sources, including operational and analytics data, as well as real-time data streaming such as Apache Kafka. This service is expected to scale across thousands of tenants and enable transactional change data capture, data replication, transformations between systems. This cost effective data replication solution is engineered for highest performance and availability to run variety of data integration uses cases of the customers. This is a global team with locations in US, Hungary, Mexico and Oracle’s India Development Centre (IDC) Bangalore.

Our team owns the reliability, availability, operational excellence, and customer experience of these services. We work closely with software development, cloud infrastructure, and customer-facing teams to ensure our services remain secure, resilient, and highly available as adoption continues to grow across commercial and sovereign cloud environments. There is a high degree of autonomy required in this role and we are looking for engineers who want to help drive our roadmap.

This role offers the opportunity to solve complex distributed systems challenges, influence service reliability at scale, and contribute to cloud services used by enterprise customers around the world.

The ideal candidate will be technically strong and get a lot done. Automation is a core tenet for everything you do. You\'ve made life easier and have motivated your team to make both process and service improvements.

Candidates should have broad working knowledge across multiple domains, but we love to see specialization as well. The basics we expect are: Networking, Linux Systems Engineering, Software Engineering/Automation, and Distributed Systems.

Learn more about OCI GoldenGate.

Experience/Qualifications:

  • This position requires that the candidate selected be a US citizen and currently possess and maintain an active Top Secret security clearance with SCI eligibility.
  • 7+ years production engineering experience (distributed services or cloud platform experience preferred)
  • Proven experience with automation and strong experience in a scripting language (bash, python, etc)
  • Solid understanding of Software Engineering and Computer Science principles
  • Solid foundation in Linux Systems Engineering
  • Solid understanding of Networking, Security and Storage
  • Experience with Docker
  • Understanding Infrastructure as Code concept
  • Experience with Terraform and/or Ansible
  • Experience in building fully automated CI/CD pipelines (Jenkins, Hudson, TeamCity, etc)
  • Be able to understand and formulate meaningful KPIs and business metrics
  • Experience in building/managing/operating cloud services at one of the big cloud providers (AWS, Azure, GPC or OCI) is a big plus
Internal Responsibilities

What You'll Do

As a Principal Site Reliability Engineer, you will be responsible for the operational health, reliability, and continuous improvement of OCI GoldenGate and OCI Database Migration Service.

You will:

  • Operate and support production cloud services that power critical customer migration and data replication workloads.
  • Monitor service health, investigate incidents, and lead troubleshooting efforts for complex production issues.
  • Participate in incident response, root cause analysis, and post-incident reviews while driving corrective and preventative actions.
  • Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).
  • Partner with software development teams to improve service architecture, reliability, scalability, and operational readiness.
  • Design and implement automation to reduce operational overhead and improve service efficiency.
  • Build and maintain observability solutions, including monitoring, alerting, logging, dashboards, and operational metrics.
  • Contribute to capacity planning, performance analysis, disaster recovery readiness, and service resilience initiatives.
  • Support services operating across Oracle Cloud commercial and sovereign cloud environments.
  • Leverage Infrastructure-as-Code and CI/CD practices to improve deployment consistency, scalability, and operational efficiency.
  • Drive continuous service improvement through operational reviews, reliability initiatives, and engineering best practices.

This role provides the opportunity to work on large-scale distributed systems supporting enterprise customers worldwide, helping shape the future of Oracle\'s cloud-native data movement platform. This role requires Professional curiosity and a desire to a develop deep understanding of services and technologies.

What We Offer
  • The opportunity to work on cloud services that support enterprise customers at global scale.
  • Exposure to complex distributed systems, cloud-native technologies, and large-scale operational challenges.
  • A collaborative environment where reliability, automation, and engineering excellence are core priorities.
  • Opportunities to influence service architecture, operational strategy, and platform evolution.
  • Career growth through ownership, technical leadership, and collaboration with world-class engineering teams.
  • Competitive compensation, comprehensive benefits, and flexible work arrangements.
External Responsibilities

What You'll Do

As a Principal Site Reliability Engineer, you will be responsible for the operational health, reliability, and continuous improvement of OCI GoldenGate and OCI Database Migration Service.

You will:

  • Operate and support production cloud services that power critical customer migration and data replication workloads.
  • Monitor service health, investigate incidents, and lead troubleshooting efforts for complex production issues.
  • Participate in incident response, root cause analysis, and post-incident reviews while driving corrective and preventative actions.
  • Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).
  • Partner with software development teams to improve service architecture, reliability, scalability, and operational readiness.
  • Design and implement automation to reduce operational overhead and improve service efficiency.
  • Build and maintain observability solutions, including monitoring, alerting, logging, dashboards, and operational metrics.
  • Contribute to capacity planning, performance analysis, disaster recovery readiness, and service resilience initiatives.
  • Support services operating across Oracle Cloud commercial and sovereign cloud environments.
  • Leverage Infrastructure-as-Code and CI/CD practices to improve deployment consistency, scalability, and operational efficiency.
  • Drive continuous service improvement through operational reviews, reliability initiatives, and engineering best practices.

This role provides the opportunity to work on large-scale distributed systems supporting enterprise customers worldwide, helping shape the future of Oracle's cloud-native data movement platform. This role requires Professional curiosity and a desire to a develop deep understanding of services and technologies.

What We Offer

  • The opportunity to work on cloud services that support enterprise customers at global scale.
  • Exposure to complex distributed systems, cloud-native technologies, and large-scale operational challenges.
  • A collaborative environment where reliability, automation, and engineering excellence are core priorities.
  • Opportunities to influence service architecture, operational strategy, and platform evolution.
  • Career growth through ownership, technical leadership, and collaboration with world-class engineering teams.
  • Competitive compensation, comprehensive benefits, and flexible work arrangements.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Data Systems Software Engineer - SRE
Principal Data Systems Software Engineer - SRE

Oracle • Seattle (WA)

On-site
USD 115,000 - 235,000
Competitive compensation
Comprehensive benefits
Flexible work arrangements
Principal Data Systems Software Engineer - SRE
Principal Data Systems Software Engineer - SRE

Oracle • Herndon (VA)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
Paid time off
401(k) with company match
+3
Principal Data Systems Software Engineer - SRE
Principal Data Systems Software Engineer - SRE

Ll Oefentherapie • Seattle (WA), Herndon (VA)

On-site
USD 120,000 - 180,000
Top Secret clearance required
Principal Data Systems Software Engineer
Principal Data Systems Software Engineer

Ll Oefentherapie • Seattle (WA)

On-site
USD 150,000 - 210,000
Principal Software Engineer, Core Infrastructure
Principal Software Engineer, Core Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 150,000 - 210,000
Relocation assistance
Principal Cloud Data Reliability Engineer
Principal Cloud Data Reliability Engineer

Oracle • Seattle (WA)

On-site
USD 115,000 - 235,000
Competitive compensation
Comprehensive benefits
Flexible work arrangements
Principal Cloud Data Reliability Engineer
Principal Cloud Data Reliability Engineer

Oracle Corporation • Seattle (WA)

On-site
USD 140,000 - 190,000
Global scale projects
Distributed systems exposure
Reliability-focused culture
+3
Senior Engineer, Core Infrastructure
Senior Engineer, Core Infrastructure

Ll Oefentherapie • Seattle (WA)

On-site
USD 79,000 - 210,000
Senior Cloud Reliability Engineer – Data Systems
Senior Cloud Reliability Engineer – Data Systems

Oracle • Herndon (VA)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
Paid time off
401(k) with company match
+3
Lead Principal Site Reliability Engineer - (work in Vienna VA location)
Lead Principal Site Reliability Engineer - (work in Vienna VA location)

Oracle • Honolulu (HI)

On-site
USD 96,000 - 264,000
Medical, dental, and vision insurance
Paid time off
401(k) with company match
+4