Senior Site Reliability Engineer

Oracle

Bengaluru

On-site

INR 2,500,000 - 4,500,000

Full time

2 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Oracle in Bengaluru is seeking an experienced Site Reliability Engineer to own end-to-end cloud service stacks, focusing on security, scalability, and performance. You will partner with development teams to enhance architecture and automate operations while serving as an escalation point for critical issues.

Responsibilities include incident management, on-call rotations, and building tools for proactive reliability.

Qualifications

  • 8 to 10+ years overall experience in IT industry.
  • Minimum 5 years of experience in Cloud Environment (Compute Knowledge).
  • Strong systems architecture skills.
  • Strong Linux administration (Understanding of different hardware family).
  • Virtualization Technologies.

Responsibilities

  • Work with SRE team on the shared full stack ownership of services and technology areas.
  • Design and deliver mission-critical stack with focus on security, resiliency, scale and performance.
  • Maintain end-to-end performance and operability with SOPs where needed.
  • Guide development teams to engineer premier capabilities for the Oracle Cloud service portfolio.
  • Troubleshoot complex issues across distributed systems and define mitigations.
  • Manage incident response, on-call rotations, and post-incident reviews.
  • Automate repetitive tasks to improve operations efficiency and reduce manual work.

Skills

Cloud Environment
Linux Administration
Scripting
Monitoring Tools
Configuration Tools
CI/CD
Docker/Kubernetes
Networking

Tools

Prometheus/Grafana
New Relic
Elasticsearch
Docker
Kubernetes

Job description

Job Description

Work with product team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the mitigating critical customer incidents, or deployments or testing required to improve security, performance, availability, and scalability of service. Authority for end-to-end performance and operability. Partner with development teams in meeting SLA to unblock customers. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the effect of product architecture decisions on distributed systems. Professional curiosity and a desire to develop deep understanding of services and technologies.

Responsibilities

Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the design and delivery of the mission critical stack, with focus on security, resiliency, scale, and performance. Authority for end-to-end performance and operability. Partner with development teams in defining and implementing improvements in service architecture. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the affect of product architecture decisions on distributed systems. Professional curiosity and a desire to develop deep understanding of services and technologies.

We are looking for a Site Reliability Developer to manage our cloud services. Responsibilities include:

  • Incident Management (NOC/P1 Exposure)
  • Support and troubleshooting complex problem in Staging/Production environments
  • Response and Resolve incidents as per SLA's
  • Organize, Anticipate, Plan and work as On-Call in shifts for multiple services (Open to work in shifts & shows flexibility)
  • Maintain Service High Availability
  • Test and Deploy solutions and automate to replace manual processes
  • Build and maintain Operational tools/procedures
  • Zero downtime deployments and a high availability mindset
  • Define and build innovative solution methodologies and assets around infrastructure, cloud migration and deployment operations at scale.
  • Work with service teams to resolve complex issues that require troubleshooting and knowledge of code.
  • Keep documentation up to date and resolving similar tickets with lower turnaround time and within SLA
  • Ensure production security posture
  • Ensure monitoring is robust and effective
  • Change Management
  • Perform Root Cause Analysis
  • Drive and Execute Operations efficiency programs and automations
  • Design and execution of Tactical and strategical fixes for operational improvements
  • Design, write and test software for better tooling capabilities
  • Reduce Manual workload through targeted automation of repetitive tasks
Required Skills
  • 8 to 10+ years overall experience in IT industry
  • Minimum 5 years of experience in Cloud Environment (Compute Knowledge)
  • Strong systems architecture skills
  • Strong Linux administration (Understanding of different hardware family)
  • Virtualization Technologies
  • Scripting Language (Python /Java /Go)
  • Hands on experience at Monitoring/Instrumentation tools (Prometheus/Grafana, new relic, elastic or equivalent).
  • Experience with maintaining high scale deployments, managing high throughput and IO intensive services.
  • Strong knowledge of system configuration tools such as Chef, Terraform, GIT, Jenkins/Hudson, Artifactory
  • Continuous Integration development/deployment, e.g. Docker, Kubernetes
  • Understanding of Networking, Cloud Computing, Load Balancers, Autoscaling
Qualifications

Career Level - IC3

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle • Hyderabad

On-site
INR 2,400,000 - 4,200,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Oracle India Private Limited • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Flexible benefits
Medical insurance
Retirement plan
+1
Principal Site Reliability Engineer
Principal Site Reliability Engineer

Oracle India Private Limited • Bengaluru

On-site
INR 3,600,000 - 6,000,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Chennai District

On-site
INR 6,000,000 - 11,000,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Dadri

On-site
INR 2,500,000 - 4,200,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Bengaluru

On-site
INR 3,500,000 - 7,000,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Ahmedabad District

On-site
INR 3,500,000 - 7,000,000
Senior Manager, Application Software Engineering
Senior Manager, Application Software Engineering

Oracle • Hyderabad

On-site
INR 3,500,000 - 7,500,000
Flexible medical options
Life insurance
Retirement options
Software Development Manager
Software Development Manager

Oracle • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Competitive benefits
Flexible medical
Life insurance
+2
Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Oracle • Thiruvananthapuram

On-site
INR 1,200,000 - 1,800,000