Principal Site Reliability Engineer

Oracle

Bengaluru

On-site

INR 3,000,000 - 4,200,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oracle Cloud Infrastructure (OCI) seeks a Principal Site Reliability Engineer (IC4) to build, operate, and evolve highly available cloud services at global scale. You will lead reliability, automation, observability, and incident excellence while mentoring teams and shaping SRE practices across services.

You will own reliability across distributed systems, drive SLIs/SLOs, and coordinate with software, security, and operations teams to design resilient architectures and scalable solutions.

Qualifications

  • Strong experience operating and troubleshooting large-scale production systems.
  • Deep understanding of distributed systems, availability, scalability, and reliability.
  • Experience designing or operating cloud infrastructure at scale.
  • Proficiency in Linux systems, networking, storage, and compute.
  • Hands-on knowledge of observability, metrics, logging, tracing, and alerting.

Responsibilities

  • Own reliability across complex distributed systems at cloud scale.
  • Develop and enforce SLIs/SLOs, and improve monitoring and incident practices.
  • Lead or contribute to complex production incident response and post-incident analyses.
  • Drive automation and toil reduction through engineering solutions.
  • Mentor engineers and promote SRE best practices across teams.
  • Collaborate with software, architecture, security, and operations groups.

Skills

Large-scale production systems
Distributed systems
Incident management
Observability
Automation
SRE leadership
Cloud infrastructure
Linux & networking
Troubleshooting
Stakeholder communication

Tools

Kubernetes
Containers
Infrastructure-as-Code
Monitoring & tracing tools

Job description

Level - IC4, Principal Site Reliability Engineer

Exp required - 8 to 12 years

Join Oracle Cloud Infrastructure (OCI) as a Principal Site Reliability Engineer (IC4) to build and operate highly available, scalable, and resilient cloud services at global scale. You’ll drive reliability engineering, automation, observability, and incident excellence while solving complex distributed systems challenges and influencing engineering best practices across teams.

About the Role

Oracle Cloud Infrastructure (OCI) is looking for a Principal Site Reliability Engineer (IC4) to help build, operate, and evolve highly available, scalable, and resilient cloud services.

As a Principal SRE, you will take technical ownership of reliability across complex, distributed systems operating at cloud scale. You will work closely with software engineering, architecture, security, and operations teams to influence service design, improve availability and performance, automate operational work, and ensure our services meet their reliability objectives.

This role goes beyond operating production systems. You will identify systemic reliability risks, influence architecture and engineering decisions, lead complex incident investigations, build automation, improve observability, and drive long-term engineering improvements.

You will also serve as a technical leader within the team, mentoring engineers and helping establish strong SRE practices across services.

Key Responsibilities
Reliability Architecture & Capacity Engineering
  • Design and influence architectures for highly available, resilient, scalable, and operationally efficient cloud services.
  • Partner with software development teams during design and implementation to ensure reliability, scalability, observability, security, and operability are built into services from the beginning.
  • Identify architectural and operational risks across multiple services and drive engineering improvements to address them.
  • Define and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), monitoring strategies, and reliability standards.
  • Forecast infrastructure and service capacity based on workload growth, utilization trends, architecture changes, and customer demand.
  • Identify capacity risks and bottlenecks before they impact customers and drive appropriate mitigation plans.
  • Lead technical prototypes and evaluations for new infrastructure, reliability patterns, and operational technologies.
Production Engineering & Service Lifecycle
  • Own and continuously improve the operational health of production services.
  • Analyze service telemetry, operational data, and reliability trends to identify systemic issues and improvement opportunities.
  • Drive improvements across availability, latency, performance, scalability, security, recoverability, and operational efficiency.
  • Establish mechanisms to detect reliability degradation before it becomes customer impacting.
  • Lead complex service lifecycle activities including upgrades, migrations, security updates, disaster recovery, capacity expansion, and decommissioning.
  • Identify recurring operational issues and convert them into engineering problems with sustainable solutions.
Automation & Toil Reduction
  • Identify high-impact opportunities to eliminate repetitive operational work through software engineering and automation.
  • Design and build scalable automation, tooling, and frameworks for deployment, monitoring, diagnostics, mitigation, remediation, and service lifecycle management.
  • Develop automated mechanisms for detecting and recovering from common failure scenarios.
  • Establish engineering standards for operational tooling, ensuring automation is reliable, testable, maintainable, observable, and safe.
  • Measure operational toil and drive initiatives that improve engineering efficiency and reduce manual intervention.
  • Review and improve automation developed by other engineers.
Observability & Performance Engineering
  • Define and improve observability strategies across services using metrics, logs, traces, dashboards, and alerting.
  • Develop meaningful service health indicators that accurately reflect customer experience.
  • Analyze production workloads to identify performance bottlenecks, resource inefficiencies, scaling limitations, and reliability risks.
  • Drive improvements to monitoring and alerting to improve signal quality and reduce operational noise.
  • Use production data and reliability trends to influence architecture, capacity planning, and engineering priorities.
Technical Leadership & Engineering Excellence
  • Provide technical leadership for reliability initiatives spanning multiple services or engineering teams.
  • Influence architecture and design decisions by identifying reliability, scalability, operational, and failure-mode considerations.
  • Lead technical discussions and design reviews for complex infrastructure and reliability challenges.
  • Establish and promote engineering best practices for operating large-scale distributed systems.
  • Mentor SREs and software engineers in troubleshooting, incident management, automation, observability, and reliability engineering.
  • Review designs, operational readiness, automation, and implementation approaches and provide actionable technical feedback.
  • Raise the overall technical and operational maturity of the team.
  • Partner with software engineering, architecture, security, networking, infrastructure, and operations teams to solve complex reliability problems.
  • Clearly communicate service health, operational risks, capacity constraints, incident impact, and reliability priorities to technical and non-technical stakeholders.
  • Anticipate the operational impact of infrastructure, architecture, feature, and tooling changes across multiple services.
  • Drive alignment across teams when reliability improvements require changes across organizational boundaries.
  • Provide clear technical recommendations supported by production data and engineering analysis.
  • Evaluate emerging technologies, engineering approaches, and SRE practices that can improve reliability, scalability, security, or operational efficiency.
  • Identify systemic weaknesses in existing operational processes and drive improvements.
  • Use operational data, incident trends, and engineering metrics to prioritize reliability investments.
  • Contribute reusable tools, patterns, standards, and best practices that benefit teams beyond your immediate area.
  • Stay current with developments in cloud infrastructure, distributed systems, observability, automation, and Site Reliability Engineering.
What We’re Looking For
  • Strong experience operating and troubleshooting large-scale production systems.
  • Strong understanding of distributed systems, high availability, scalability, fault tolerance, and reliability engineering principles.
  • Experience designing or operating cloud infrastructure and services at scale.
  • Strong understanding of Linux systems, networking, storage, compute, and cloud infrastructure concepts.
  • Experience with monitoring, observability, metrics, logging, tracing, and alerting systems.
  • Experience defining or working with SLIs, SLOs, availability targets, and service health metrics.
  • Strong troubleshooting and debugging skills across application and infrastructure layers.
  • Experience leading or significantly contributing to complex production incident response and root cause analysis.
  • Experience identifying and eliminating operational toil through engineering and automation.
  • Ability to influence technical decisions and drive engineering initiatives across teams.
  • Strong written and verbal communication skills.
  • Demonstrated ability to mentor engineers and raise engineering standards within a team.
Preferred Qualifications
  • Experience working with large-scale cloud platforms such as Oracle Cloud Infrastructure (OCI), AWS, Azure, or GCP.
  • Experience with Kubernetes, containers, infrastructure-as-code, and modern deployment technologies.
  • Experience building automation and internal platforms for large-scale infrastructure operations.
  • Experience with capacity planning, performance engineering, disaster recovery, and resilience testing.
  • Experience designing highly available distributed systems and understanding complex failure modes.
  • Experience improving operational readiness and reliability across multiple services.
  • Experience driving reliability initiatives that span multiple engineering teams.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Site Reliability Engineer
Principal Site Reliability Engineer

Oracle India Private Limited • Bengaluru

On-site
INR 3,600,000 - 6,000,000
Principal Site Reliability Engineer
Principal Site Reliability Engineer

United States Digital Space LLC • Karnataka

On-site
INR 3,500,000 - 5,500,000
Senior Cloud Site Reliability Engineer
Senior Cloud Site Reliability Engineer

Augusta Infotech • Bengaluru

Hybrid
INR 1,500,000 - 2,500,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Technologies Pvt. Ltd. • Pune District

On-site
INR 900,000 - 1,400,000
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Sierra Ventures • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle India Private Limited • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Flexible benefits
Medical insurance
Retirement plan
+1
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Arcesium • Hyderabad, Bengaluru

Hybrid
INR 6,000,000 - 9,000,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

AcquireX • Pune District

On-site
INR 1,200,000 - 1,800,000
Health insurance
Flexible working hours
Training opportunities
Lead Site Reliability Engineer
Lead Site Reliability Engineer

Quest Diagnostics • Hyderabad

On-site
INR 4,500,000 - 8,000,000