Senior Site Reliability Engineer

Ll Oefentherapie

Bengaluru

On-site

INR 1,800,000 - 3,200,000

Full time

2 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oracle Bengaluru is seeking an experienced SRE to enhance reliability and scalability of OCI Compute services. You will work with engineering and infrastructure teams to improve operation and incident response.

The role requires 4–8 years of relevant experience, strong scripting skills, and hands-on knowledge of Linux, cloud infra, and monitoring. You will participate in a 12x7 on-call rotation and drive automation initiatives.

Qualifications

  • 4–8 years of experience in SRE, Production Engineering, Cloud Operations, or a similar role.
  • Experience operating and improving highly available production systems.
  • Strong programming or scripting skills in Python, Java, Go, or similar languages.
  • Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change‑management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on technical problems and collaborate effectively with engineering teams.
  • Strong written and verbal communication skills.

Responsibilities

  • Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
  • Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
  • Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
  • Build automation and tooling to reduce operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
  • Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Share technical knowledge and support team members through documentation, reviews, and collaboration.
  • Participate in a 12x7 on-call rotation and support response to customer-impacting incidents.

Skills

SRE Experience
Python/Java/Go
Linux & Cloud
Observability
On-call Rotation
Incident Troubleshooting
CI/CD Pipelines

Tools

OCI
Kubernetes
Infrastructure as Code
CI/CD

Job description

  • Does this position require a security clearance? No
  • Years 3 to 5+ years
  • Additional Info Visa / work permit sponsorship is not available for this position
  • Applicants are required to read, write, and speak the following languages English
Job Responsibilities
  • Improve the reliability, scalability, performance, and operational efficiency of assigned OCI Compute services and components.
  • Investigate and resolve complex production incidents; contribute to mitigation, recovery, RCA, and follow-up actions.
  • Own and improve service-level KPIs, SLOs, dashboards, alerting, deployment validation, and operational procedures for assigned systems.
  • Build automation and tooling to reduce operational toil and improve production safety.
  • Partner with development and infrastructure teams on service architecture, deployment, configuration, and reliability improvements.
  • Use observability, telemetry, event correlation, and AIOps capabilities to improve detection, diagnosis, and incident response.
  • Support upgrades, migrations, patching, capacity planning, performance tuning, security vulnerability management and production rollouts.
  • Troubleshoot distributed-system issues by analyzing service topology, dependencies, configuration, and failure modes.
  • Contribute to incident-management practices, operational readiness, and service ownership improvements.
  • Share technical knowledge and support team members through documentation, reviews, and collaboration.
  • Participate in a 12x7 on-call rotationand support response to customer-impacting incidents.
Mandatory Skills
  • 4–8 yearsof experience in SRE, Production Engineering, Cloud Operations, Systems Engineering, or a similar role.
  • Experience operating and improving highly available production systems.
  • Strong programming or scripting skills in Python, Java, Go, or similar languages.
  • Hands-on experience with Linux, cloud infrastructure, networking, compute, and storage.
  • Experience with production monitoring, alerting, dashboards, logs, metrics, and tracing.
  • Experience owning or improving service SLIs, SLOs, KPIs, and operational procedures.
  • Strong incident troubleshooting, RCA, debugging, and problem-solving skills.
  • Experience with deployment pipelines, release validation, automation, and change‑management practices.
  • Understanding of distributed systems, service dependencies, capacity planning, and performance tuning.
  • Ability to work independently on technical problems and collaborate effectively with engineering teams.
  • Strong written and verbal communication skills.
Preferred Skills
  • Experience with OCIand cloud infrastructure services.
  • Experience with AIOps, anomaly detection, event correlation, predictive alerting, or automated remediation.
  • Experience with Kubernetes, containers, infrastructure-as-code, and CI/CD.
  • Experience with service migrations, fleet maintenance, upgrades, patching, or production rollouts.
  • Experience with architecture reviews, operational-readiness reviews, and post-incident improvements.
  • Experience contributing to technical initiatives, knowledge sharing, code reviews, or operational improvements within the team.
  • Familiarity with security, compliance, and access‑control practices in production environments.
Self-Test Questions
  • Do you have 4–8 yearsof relevant SRE, Production Engineering, Cloud Operations, or Systems Engineering experience?
  • Have you independently operated or improved a production service, system, or infrastructure component?
  • Can you investigate production incidents and contribute to mitigation, recovery, RCA, and follow‑up actions?
  • Do you have hands‑on experience with Linux, cloud infrastructure, distributed systems, networking, compute, or storage?
  • Are you proficient in Python, Java, Go, or a similar language for automation, tooling, and troubleshooting?
  • Have you built or improved automation, deployment validation, CI/CD pipelines, or operational tooling?
  • Do you have experience with monitoring, alerting, logs, metrics, tracing, and service health indicators such as SLOs or KPIs?
  • Can you work independently on assigned technical problems, collaborate with partner teams, and participate in a 12x7 on-call rotation?
About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Request a referral from an Oracle employee.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Ll Oefentherapie • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Senior Advanced Services Engineer
Senior Advanced Services Engineer

Ll Oefentherapie • Bengaluru

On-site
INR 1,200,000 - 1,800,000
Senior Platform Software Engineer
Senior Platform Software Engineer

Oracle • Hyderabad

On-site
INR 1,200,000 - 1,800,000
Competitive benefits
Flexible medical
Life insurance
+2
Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Ll Oefentherapie • Bengaluru

On-site
INR 4,000,000 - 7,500,000
Senior Manager - Cloud Infrastructure Engineering
Senior Manager - Cloud Infrastructure Engineering

Oracle India Private Limited • Bengaluru

On-site
INR 3,500,000 - 6,500,000
Associate Support Engineer - Technical Customer Support
Associate Support Engineer - Technical Customer Support

Oracle India Private Limited • Bengaluru

On-site
INR 400,000 - 600,000
Flexible medical
Life insurance
Retirement options
+1
Senior Network Developer
Senior Network Developer

Oracle • Thiruvananthapuram

On-site
INR 1,200,000 - 1,800,000
Principal Software Engineer - Network Reliability Engineering - AI/ML
Principal Software Engineer - Network Reliability Engineering - AI/ML

Oracle • Bengaluru

On-site
INR 1,800,000 - 2,500,000
Competitive benefits
Flexible medical and life insurance
Retirement options
Lead Principal Advanced Services Engineer
Lead Principal Advanced Services Engineer

Oracle India Private Limited • Bengaluru

On-site
INR 4,000,000 - 6,500,000
Principal Network Developer
Principal Network Developer

Ll Oefentherapie • Bengaluru

On-site
INR 2,000,000 - 3,000,000