Get more replies from employers
Send a job-specific resume in minutes.
Oracle India Private Limited is seeking a Principal Site Reliability Engineer (IC4) to own reliability across complex cloud services hosted in OCI. You will influence service design with engineering teams, drive observability, automation, and incident responses at cloud scale.
You will lead cross-functional reliability initiatives, mentor engineers, and raise the engineering maturity of the team while delivering highly available systems and resilient infrastructure.
Oracle Cloud Infrastructure (OCI) is looking for a Principal Site Reliability Engineer (IC4) to help build, operate, and evolve highly available, scalable, and resilient cloud services. As a Principal SRE, you will take technical ownership of reliability across complex, distributed systems operating at cloud scale. You will work closely with software engineering, architecture, security, and operations teams to influence service design, improve availability and performance, automate operational work, and ensure our services meet their reliability objectives. This role goes beyond operating production systems. You will identify systemic reliability risks, influence architecture and engineering decisions, lead complex incident investigations, build automation, improve observability, and drive long-term engineering improvements. You will also serve as a technical leader within the team, mentoring engineers and helping establish strong SRE practices across services.
Design and influence architectures for highly available, resilient, scalable, and operationally efficient cloud services. Partner with software development teams during design and implementation to ensure reliability, scalability, observability, security, and operability are built into services from the beginning. Identify architectural and operational risks across multiple services and drive engineering improvements to address them. Define and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), monitoring strategies, and reliability standards. Forecast infrastructure and service capacity based on workload growth, utilization trends, architecture changes, and customer demand. Identify capacity risks and bottlenecks before they impact customers and drive appropriate mitigation plans. Lead technical prototypes and evaluations for new infrastructure, reliability patterns, and operational technologies.
Own and continuously improve the operational health of production services. Analyze service telemetry, operational data, and reliability trends to identify systemic issues and improvement opportunities. Drive improvements across availability, latency, performance, scalability, security, recoverability, and operational efficiency. Establish mechanisms to detect reliability degradation before it becomes customer impacting. Lead complex service lifecycle activities including upgrades, migrations, security updates, disaster recovery, capacity expansion, and decommissioning. Identify recurring operational issues and convert them into engineering problems with sustainable solutions.
Identify high-impact opportunities to eliminate repetitive operational work through software engineering and automation. Design and build scalable automation, tooling, and frameworks for deployment, monitoring, diagnostics, mitigation, remediation, and service lifecycle management. Develop automated mechanisms for detecting and recovering from common failure scenarios. Establish engineering standards for operational tooling, ensuring automation is reliable, testable, maintainable, observable, and safe. Measure operational toil and drive initiatives that improve engineering efficiency and reduce manual intervention. Review and improve automation developed by other engineers.
Define and improve observability strategies across services using metrics, logs, traces, dashboards, and alerting. Develop meaningful service health indicators that accurately reflect customer experience. Analyze production workloads to identify performance bottlenecks, resource inefficiencies, scaling limitations, and reliability risks. Drive improvements to monitoring and alerting to improve signal quality and reduce operational noise. Use production data and reliability trends to influence architecture, capacity planning, and engineering priorities.
Provide technical leadership for reliability initiatives spanning multiple services or engineering teams. Influence architecture and design decisions by identifying reliability, scalability, operational, and failure-mode considerations. Lead technical discussions and design reviews for complex infrastructure and reliability challenges. Establish and promote engineering best practices for operating large-scale distributed systems. Mentor SREs and software engineers in troubleshooting, incident management, automation, observability, and reliability engineering. Review designs, operational readiness, automation, and implementation approaches and provide actionable technical feedback. Raise the overall technical and operational maturity of the team.
Partner with software engineering, architecture, security, networking, infrastructure, and operations teams to solve complex reliability problems. Clearly communicate service health, operational risks, capacity constraints, incident impact, and reliability priorities to technical and non-technical stakeholders. Anticipate the operational impact of infrastructure, architecture, feature, and tooling changes across multiple services. Drive alignment across teams when reliability improvements require changes across organizational boundaries. Provide clear technical recommendations supported by production data and engineering analysis.
Evaluate emerging technologies, engineering approaches, and SRE practices that can improve reliability, scalability, security, or operational efficiency. Identify systemic weaknesses in existing operational processes and drive improvements. Use operational data, incident trends, and engineering metrics to prioritize reliability investments. Contribute reusable tools, patterns, standards, and best practices that benefit teams beyond your immediate area. Stay current with developments in cloud infrastructure, distributed systems, observability, automation, and Site Reliability Engineering.
About Oracle Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. With AI embedded across our products and services, we help customers turn that promise into a better future for all. At Oracle Cloud Infrastructure, you will have the opportunity to work on cloud services operating at significant scale and solve challenging distributed systems and reliability problems that directly affect our customers. Oracle is committed to creating an inclusive workplace where everyone has the opportunity to contribute, grow, and succeed. Career Level - IC4 Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Experience Level Senior Level