Senior SRE Architect | Cloud Reliability & Automation

Ll Oefentherapie

Boise (ID)

On-site

USD 96,000 - 264,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oracle is seeking a Lead Principal Site Reliability Engineer to support a strategic cloud platform and mission-critical applications in a Vienna VA location. You will own reliability, automation, observability, capacity planning, and incident response to ensure high availability and scalable performance.

You will partner with security, development, and operations teams to implement IaC, CI/CD, and best practices for core infrastructure.

Qualifications

  • Bachelor's degree in CS/IT/Engineering or equivalent.
  • 8+ years of Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering experience.
  • Experience supporting production environments with strict availability requirements.
  • Hands-on experience with Oracle Cloud Infrastructure (OCI) or another major cloud provider (AWS, Azure, GCP).
  • Experience with Kubernetes, Docker, and container orchestration platforms.
  • Proficiency in Infrastructure as Code (Terraform preferred).
  • Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
  • Strong scripting skills in Python, Bash, or PowerShell.
  • Experience with Linux system administration.
  • Strong understanding of networking concepts including DRG, DNS, load balancing, routing, APIs, and Endpoints.
  • Experience with Observability tools such as Prometheus, Grafana, ELK/OpenSearch, Splunk, Datadog, New Relic, or OCI Native Observability services.
  • Strong troubleshooting and problem-solving skills.
  • Experience supporting financial services or other highly regulated industries.
  • Knowledge of disaster recovery, backup strategies, and business continuity.

Responsibilities

  • Maintain high availability, reliability, and performance of enterprise applications and cloud infrastructure.
  • Design and implement automation to reduce manual operational effort and improve deployment consistency.
  • Develop Infrastructure as Code (IaC) using Terraform and automation scripts using Python, Bash, or similar languages.
  • Build and maintain CI/CD pipelines to support automated deployments and release management.
  • Monitor applications and infrastructure using observability platforms, including metrics, logs, traces, and alerting.
  • Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Participate in on-call rotations and respond to production incidents with urgency and professionalism.
  • Conduct root cause analysis (RCA) and implement corrective and preventive actions to eliminate recurring issues.
  • Perform capacity planning, performance tuning, and scalability assessments.
  • Support Kubernetes clusters, containerized workloads, and cloud-native applications.
  • Collaborate with development teams to improve application resiliency, fault tolerance, and operational readiness.
  • Partner with security teams to ensure systems comply with enterprise security and regulatory requirements.
  • Create operational runbooks, documentation, and standard operating procedures.
  • Continuously improve platform reliability through automation, monitoring, and operational best practices.

Skills

Site Reliability Engineering (SRE)
Docker
Terraform
Python
Bash
Git
Monitoring & Observability
Root Cause Analysis (RCA)
Automation
Capacity Planning
Performance Tuning
Networking
High Availability
Disaster Recovery
DevOps
Agile Methodologies

Education

Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience

Tools

Kubernetes
Docker
Terraform
Jenkins
GitHub Actions
GitLab CI
Azure DevOps

Job description

Oracle is seeking a Lead Principal Site Reliability Engineer to support a strategic cloud platform and mission-critical applications in a Vienna VA location. You will own reliability, automation, observability, capacity planning, and incident response to ensure high availability and scalable performance.

You will partner with security, development, and operations teams to implement IaC, CI/CD, and best practices for core infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior SRE Lead: Cloud Reliability & Automation
Senior SRE Lead: Cloud Reliability & Automation

Oracle • Honolulu (HI)

On-site
USD 96,000 - 264,000
Medical, dental, and vision insurance
Paid time off
401(k) with company match
+4
Senior SRE Architect: Cloud Reliability & Automation
Senior SRE Architect: Cloud Reliability & Automation

Oracle • Carson City (NV)

On-site
USD 96,000 - 264,000
Health benefits
401(k) with company match
Paid time off and holidays
+2
Senior SRE & Reliability Engineering Manager
Senior SRE & Reliability Engineering Manager

Oracle • Reston (VA)

On-site
USD 121,500 - 264,100
Senior SRE - Cloud Infra, Automation & 24/7 Ops
Senior SRE - Cloud Infra, Automation & 24/7 Ops

Oracle • Reston (VA)

On-site
USD 85,000 - 210,000
Medical, dental, and vision insurance
Paid time off
401(k) Savings Plan
Senior Site Reliability Engineer - Scalable Cloud & Automation
Senior Site Reliability Engineer - Scalable Cloud & Automation

Oracle • North Carolina

On-site
USD 84,000 - 210,000
Senior SRE & Reliability Architect (Cloud & Observability)
Senior SRE & Reliability Architect (Cloud & Observability)

Gen • Tempe (AZ)

On-site
USD 150,000 - 160,000
Senior Principal SRE - 24/7 Cloud Reliability Lead
Senior Principal SRE - 24/7 Cloud Reliability Lead

Ll Oefentherapie • Reston (VA)

On-site
USD 120,000 - 180,000
Lead Principal Site Reliability Engineer - (work in Vienna VA location)
Lead Principal Site Reliability Engineer - (work in Vienna VA location)

Ll Oefentherapie • Boise (ID)

On-site
USD 96,000 - 264,000
Site Reliability Engineering Leader
Site Reliability Engineering Leader

Oracle • Reston (VA)

On-site
USD 140,000 - 180,000
Principal Cloud SRE — Storage & DB Reliability
Principal Cloud SRE — Storage & DB Reliability

Ll Oefentherapie • United States

On-site
USD 150,000 - 190,000