Senior Manager, Site Reliability Engineering

Ll Oefentherapie

Hagåtña (GU)

On-site

USD 122,000 - 264,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oracle is seeking a Senior Manager, Site Reliability Engineering in the United States. The role leads a team of reliability engineers, guiding reliability practices, capacity forecasting, and infrastructure resource planning.

You will coordinate with software development to build scalable systems and oversee incident response and post-incident analyses. Expect proactive automation, robust monitoring, and clear communication of changes and impacts to stakeholders.

Qualifications

  • Minimum 10+ years of experience in a related role.
  • Proven leadership of site reliability or equivalent teams.
  • Fluency in English (written and spoken).

Responsibilities

  • Designs and guides reliable, scalable infrastructures with the software team.
  • Leads forecasting for infrastructure capacity and resource planning.
  • Oversees incident response, root cause analyses, and post-mortems.
  • Drives automation initiatives to improve operational efficiency.
  • Communicates release notes and service impact to stakeholders.
  • Ensures SLAs and SLOs are met across services.
  • Sets priorities and guides cross-functional teams on projects.

Skills

People management
SRE leadership
Capacity planning
Incident response
Automation
Monitoring & metrics
Cross-functional collaboration
Documentation

Job description

Senior Manager, Site Reliability Engineering

United States

Be the First to Apply

  • Job Identification 340248
  • Job Category Product and Research
  • Role People Manager
  • Job Type Regular Employee
  • Does this position require a security clearance? No
  • Years 10+ years
  • Applicants are required to read, write, and speak the following languages English
Job Description

Supports team members in designing and architecting infrastructure and service and shares guidance on practices for reliability and functionality. Provides direction to ensure accurate forecasting and ensure systems have adequate resources, identifying resource gaps. Maintains a collaborative relationship with the software development team to create reliable, scalable infrastructures. Monitors data collection and ensures team members optimize operations and infrastructure reliability. Aids in incident response activities to ensure service reliability. Monitors health and performance reports. Implements standards for identifying and recommending automation. Ensures team members communicate information and articulate the impact of changes. Serves as a senior management point and shares expectations for documentation. Sets expectations for experimenting with new technology, executing improvements, building site reliability knowledge, and providing clear data.

Responsibilities

Key Responsibilities

Capacity Ingestion andManagement:
  • Supportsteam members designing and architecting infrastructure and/or service, sharingguidance on practices and terms for reliability and functionality.
  • Supervisesteam members and provides direction to ensure accurate forecasting of demandsfor infrastructure and response to capacity needs, ensuring systems havesufficient resources to handle current and future workloads and identifyingresource gaps.
  • Maintainsa collaborative relationship with the software development team to developinfrastructures, ensuring features are reliable and scalable according todeployment requirements.
  • Implementsexpectations for identifying opportunities for prototyping and managesprototyping initiatives (e.g., testing new applications or infrastructures,assisting in onboarding) to explore novel approaches.
Incident and ServiceLifecycle Management:
  • Monitorsdata collection, triage, technical analysis, and redirection, ensuring teammembers maintain and optimize operations and infrastructure reliability.
  • Providessupport to team members monitoring services, ensuring they maintain up-to-dateknowledge of performance and document their condition.
  • Leveragesadvanced knowledge to aid team members in performing incident response, rootcause analyses, and/or maintenance on assigned services (e.g., softwareinstalls, version upgrades, security updates, backup and recovery).
  • Monitorscomprehensive health and performance reporting and ensures team members takeappropriate actions based on trends in data.
  • Ensuresteam members adhere to procedures when performing provisioning to supportinfrastructure, applications, and services.
  • Encouragesteam members to experiment with new approaches for and perform decommissioning(e.g., shutting down servers, removing data from databases) to remove objectsthat are no longer needed.
Automation:
  • Implementsstandards for identifying and recommending opportunities for automation andassesses potential benefits to enhance operational efficiency.
  • Takes aproactive role in reviewing and offering feedback on design, automation tools,or scripts, acting as a leader during implementation.
  • Sharesstrategies for conducting testing on automations to ensure they perform taskscorrectly and produce expected results.
Technical Communication andGuidance:
  • Reviewsand provides feedback on release notes and ensures team members communicatecomprehensive information about the scale, capacity, security, performanceattributes, and requirements of services and technology with customers andimmediate and related teams.
  • Proactivelyanticipates and articulates the potential impact of infrastructure, feature,and tool changes, considering their impact across team operations.
  • Servesas a resource to team members on what information to communicate and how tocommunicate.
Troubleshooting andResolution:
  • Serves as a seniormanagement escalation point for incidents and complex issues arising withinOracle services.
  • Monitors the resolution oftechnical issues spanning multiple services, ensuring effective investigationand debugging techniques are leveraged to achieve SLOs (service levelobjectives).
  • Shares expectations fordocumenting incidents performing root cause analyses, guiding team members tocapture essential information for analysis and future reference.
  • Implements guidelines forpost-mortem procedures to prevent incident reoccurrence.
  • Ensures team members adhereto service level agreements (SLAs) made with customers.
Innovation and Improvement:
  • Sets expectations forconducting experiments and evaluating cutting-edge tools and technologies tooptimize infrastructure performance and reliability, taking proactive steps toadhere to security standards.
  • Manages and contributes tothe prioritization of initiatives to improve performance bottlenecks anddeployments, ensuring efficient resource usage, speed, and scalability.
  • Implements standards fordeveloping and maintaining knowledge of site reliability trends and sharingvaluable insights and information with team members, management, and beyond topromote innovative building, testing, deploying, and running services.
  • Leverages analyses and datafrom teams to contribute to business development decisions (e.g., designchanges).
Core Responsibilities
Planning & Execution:
  • Managesmultiple medium- to large-scale projects or initiatives across teams, ensuringtimelines, deliverables, and budgets when applicable are monitored and met.Provides direction to teams on project work, setting priorities, and aligning withbusiness needs. Guides teams on adjusting plans to accommodate resource ortimeline changes.
  • Drivescross-functional partnerships to align expectations and shared objectivesacross multiple teams. Coaches team members to develop strategic relationshipswith business leaders, stakeholders, and external partners to fostercollaboration and long-term success. Promotes inclusivity by actively seekingand listening to diverse perspectives, ensuring others feel heard and respected.
Problem Solving:
  • Providesdirection to multiple teams on addressing complex operational and/or technicalissues as well as providing guidance on analyzing complex data and/orinformation to identify solutions. Reviews and provides insights into unresolvedor critical issues, helping the team to identify potential solutions.
  • Modelsengaging in continuous learning to deepen expertise and stay ahead of industrytrends, integrating best practices into strategic planning. Leverages feedbackto drive personal and team skill improvements. Identifies skill gaps acrossteams, and empowers team members to pursue learning and knowledge sharingopportunities that build their expertise in new areas and coaches them to applylearnings to advance the organization.
  • Drivesteam to collaborate on, develop, and implement ideas to increase the efficiencyand effectiveness of processes, protocols, and workflows within and acrossteams, providing oversight. Guides team to adopt new ideas for alternativeapproaches and methods and encourages feedback for continued improvement.
Performance and Development:
  • Drivesperformance across teams by providing feedback and coaching in alignment withperformance management processes, guidelines, and expectations. Discussesdevelopment goals with team members, shares opportunities to facilitate careerdevelopment, and ensures individual goals are aligned with broaderorganizational goals. Develops and manages talent acquisition pipeline byleading candidate interviews, monitoring promotion eligibility, and/ororchestrating talent resources.
Qualifications

Disclaimer: Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US: Hiring Range in USD from: $121,500 to $264,100 per annum. May be eligible for bonus, equity, and compensation deferral. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity. Oracle US offers a comprehensive benefits package which includes the following:

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - M3

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life‑saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Request a referral from an Oracle employee.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Manager, Site Reliability Engineering
Senior Manager, Site Reliability Engineering

Oracle • Nashville (TN)

On-site
USD 122,000 - 264,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Paid time off and holidays
+1
Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Oracle • Denver (CO)

On-site
USD 180,000 - 260,000
Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Oracle • United States

On-site
USD 146,000 - 306,000
Medical, dental & vision insurance
401(k) with company match
Paid time off & holidays
+1
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Nashville (TN)

On-site
USD 118,000 - 264,000
Medical, dental, and vision insurance
401(k) Matching
Paid time off
+5
Manager, Core Infrastructure Engineering
Manager, Core Infrastructure Engineering

Oracle • Seattle (WA)

On-site
USD 126,000 - 264,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
Manager, Core Infrastructure Engineering
Manager, Core Infrastructure Engineering

Oracle • Columbia (SC)

On-site
USD 126,000 - 264,000
Medical insurance
Dental insurance
Vision insurance
+1
Director, Core Infrastructure Engineering
Director, Core Infrastructure Engineering

Oracle • Nashville (TN)

On-site
USD 169,800 - 355,400
Medical insurance
401(k) with company match
Paid time off and holidays
Director, Platform Software Engineering
Director, Platform Software Engineering

Oracle • Jackson (MS)

On-site
USD 123,000 - 355,000
Health insurance
401(k) Savings and Investment Plan
Paid time off
+5
Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Ll Oefentherapie • United States

On-site
USD 146,000 - 306,000
Medical, dental, and vision insurance
Paid time off
401(k) Savings and Investment Plan
+1
Director, Platform Software Engineering
Director, Platform Software Engineering

Oracle • United States

On-site
USD 123,000 - 355,000
Medical, dental, and vision insurance
Paid time off
401(k) matching
+1