Principal Core Infrastructure Engineer

Ll Oefentherapie

United States

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Oracle Cloud Infrastructure (OCI) is seeking a Principal Site Reliability Engineer to enhance reliability and performance of storage/database services. You will own live-service issue responses, build automation, and release readiness for production deployments.

You will collaborate across teams to reduce incidents, implement scalable operational solutions, and drive capacity planning and security-focused reliability improvements.

Qualifications

  • 8+ years’ experience in Site Reliability Engineering focusing on storage, networking and database troubleshooting for improving application reliability.
  • Excellent troubleshooting skills for resolving critical production issues in cloud services.
  • Scripting in Linux to automate routine or manual tasks.
  • Experience with CI/CD pipelines and Git tools (GitHub, Bitbucket, Git).
  • Container administration and development using Kubernetes, Docker or similar.
  • Configuration management experience.
  • Programming in Python, Golang, Terraform.
  • Experience with Grafana monitoring.
  • Experience in managing 24x7 high-availability production applications.
  • Excellent verbal and written communication, and organizational skills.
  • Familiar with Agile software development with Jira.

Responsibilities

  • Improve stability, performance, and reliability of OCI storage and database services.
  • Collaborate with multiple development teams to mitigate operational risks.
  • Respond to live service issues and develop automation/tools for reliability.
  • Own release certification by judging production readiness of code.
  • Define and deliver mission-critical, secure, scalable services.

Skills

SRE experience
Cloud storage
Database troubleshooting
Linux scripting
CI/CD pipelines
Jenkins
Git
Kubernetes
Docker
Configuration management
Python
Go (Golang)
Terraform
Grafana
24x7 production support
Agile/Jira

Tools

Jenkins
Kubernetes
Docker
Grafana
Terraform
GitHub
Bitbucket

Job description

Core Infrastructure Engineering within Oracle Cloud Infrastructure (OCI) is seeking a motivated Principal Site Reliability Engineer (SRE) who thrives in a fast-paced, rapidly evolving technology environment. The ideal candidate should have experience supporting cloud-scale, highly distributed storage or database services on major cloud platforms such as OCI, AWS, GCP, or Azure.

You will be focused on improving service reliability, performance and operability of services used by Oracle OCI Tier-0 services and Oracle customers. You will have your hand on the pulse of the services and will play a key role in responding to live service issues. As a hands-on engineer with strong coding skills. You will debug complex production issues, build automation and monitoring tools. You will have the opportunity to create automation and tooling that will allow us to continuously improve our services. You will own the release certification process by determining whether code is ready for production deployment.

In this role, you will be responsible for improving the stability, performance, and reliability of database and storage services which are backbone of OCI. You will collaborate with multiple development teams to identify and resolve cross-functional operational risks by combining engineering expertise, troubleshooting skills, and operational best practices. The role requires a high degree of independence, excellent communication and organizational skills, and a strong commitment to improving customer experience by enhancing service reliability, reducing support tickets, and delivering scalable operational solutions. You will also define and deliver mission-critical services with a strong focus on security, resiliency, scalability, capacity planning, performance management, deployment, and release engineering.

Qualifications
  • 8+ years’ experience in Site Reliability Engineering and in storage, networking and database troubleshooting for improving application reliability, scalability, availability
  • Excellent troubleshooting skills for resolving critical production issues in cloud services
  • Expertise in developing scripts (linux scripting), utilities and tools to automate routine or manual intensive tasks
  • Solid experience with CI/CD pipelines, Jenkins and Version Control tools (GitHub, Bit Bucket, GIT)
  • Container administration and development experience utilizing Kubernetes, Docker or similar
  • Solid experience with Configuration Management tools
  • Programming languages development experience using Python, Golang, Terraform
  • Experience with monitoring tools such as Grafana
  • Experience in managing 24×7 high-availability production applications
  • Excellent organizational, verbal, and written communication skills
  • Good understanding of Agile software development principles including using common tools such as JIRA
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Cloud SRE — Storage & DB Reliability
Principal Cloud SRE — Storage & DB Reliability

Ll Oefentherapie • United States

On-site
USD 150,000 - 190,000
Principal Core Infrastructure Engineer (OCI Object Storage)
Principal Core Infrastructure Engineer (OCI Object Storage)

Ll Oefentherapie • Santa Clara (CA)

Hybrid
USD 170,000 - 250,000
Cleared Site Reliability Engineer - Database
Cleared Site Reliability Engineer - Database

Oracle • Reston (VA)

On-site
USD 79,000 - 159,000
Medical, dental, and vision insurance
401(k) savings plan with company match
Paid parental leave
+1
Cleared Site Reliability Engineer - Database
Cleared Site Reliability Engineer - Database

Ll Oefentherapie • Reston (VA)

On-site
USD 140,000 - 210,000
Cleared Site Reliability Engineer - Database
Cleared Site Reliability Engineer - Database

Ll Oefentherapie • Seattle (WA)

On-site
USD 100,000 - 130,000
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Socket.dev • Nashville (TN)

On-site
USD 180,000 - 240,000
Competitive benefits
Volunteer programs
OCI Senior Core Infrastructure Engineer - Nashville TN
OCI Senior Core Infrastructure Engineer - Nashville TN

Ll Oefentherapie • Nashville (TN)

On-site
USD 120,000 - 180,000
Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Ll Oefentherapie • Seattle (WA)

On-site
USD 89,200 - 209,500
Medical, dental, and vision insurance
401(k) Savings and Investment Plan with company match
Paid time off including flexible vacation and sick leave
Sr. SRE-Oracle DBA
Sr. SRE-Oracle DBA

Oracle • Seattle (WA)

On-site
USD 86,000 - 200,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan with company match
Paid parental leave
Principal Software Engineer, Core Infrastructure
Principal Software Engineer, Core Infrastructure

Ll Oefentherapie • Seattle (WA)

On-site
USD 120,000 - 150,000