Senior Core Infrastructure Engineer - Scale & Reliability

Oracle

Santa Clara (CA)

On-site

USD 79,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Medical, dental, and vision insurance
401(k) with company match
Paid time off
Parental leave

Job summary

Oracle in Santa Clara, CA is seeking a skilled System Design & Architecture engineer to design, implement, and optimize components of distributed systems focused on scalability, resiliency, and operability. The role emphasizes fault-tolerant paths, replication, and automated recovery in large-scale environments.

You will work across design, reliability, performance, and security domains, building automation, runbooks, and telemetry to ensure resilient cloud infrastructure.

Qualifications

  • 3+ years of relevant engineering, architecture, or development/operational experience.
  • Strong experience working on data plane architectures in networking devices.
  • Strong experience with high-concurrency systems.
  • Experience designing, developing and optimizing high performance network solutions using DPDK, C/C++
  • Working experience with Linux OSes/kernels, device drivers, performance testing tools, distributed debugging tools
  • Strong team player with outstanding communication, organization, and interpersonal skills.
  • Comfortable with complex, swiftly evolving software development environments.
  • Ability to learn new technologies quickly and drive, follow, evangelize, and improve cross-team processes.
  • Expert knowledge of cloud infrastructure concepts and technologies.
  • Experience working with geographically distributed teams.
  • Significant work experience in startups or fast-paced enterprise technology development environments.

Responsibilities

  • Implements and contributes to the development for components of distributed systems that support horizontal and vertical scaling.
  • Optimizes code and/or systems for large-scale data processing in large-scale systems.
  • Implements scalability requirements for assigned components and reviews implementation of team members.
  • Leverages components of data plane platforms to handle large-scale data retrieval, storage, and processing.
  • Implements performance and load testing.
  • Collaborates with team to build fault-tolerant components capable of withstanding in-service updates by implementing redundancy, replication, and automatic failover mechanisms.
  • Applies recovery oriented computing principles to design components that effectively handle service disruptions.
  • Implements retry mechanisms, circuit breakers, and timeouts to help handle network unreliability.
  • Implements tests and alarm configurations to proactively detect and address issues/failures.
  • Supports efforts to recover from failures by drafting and executing runbooks and operational procedures.
  • Builds and customizes dashboards, telemetry systems, and alerting mechanisms to monitor component health.
  • Designs and implements functional requirements and testing for assigned features within an existing system.
  • Implements tests scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.
  • Implements standard data replication and synchronization techniques to maintain data integrity and availability.
  • Diagnoses, debugs, and resolves issues in system components to support ongoing operation.
  • Designs and implements automation scripts and tooling used to troubleshoot operational issues.
  • Participates in operational support rotations, assisting in incident responses and root cause investigations.
  • Applies advanced security measures to protect data and applications in multi-tenant environments, including encryption and access controls.
  • Implements remediation plans to continuously improve security.

Skills

High-concurrency systems
Distributed systems
C/C++
DPDK
Cloud infrastructure

Tools

Linux kernel
Device drivers
Performance testing tools
Distributed debugging tools

Job description

Oracle in Santa Clara, CA is seeking a skilled System Design & Architecture engineer to design, implement, and optimize components of distributed systems focused on scalability, resiliency, and operability. The role emphasizes fault-tolerant paths, replication, and automated recovery in large-scale environments.

You will work across design, reliability, performance, and security domains, building automation, runbooks, and telemetry to ensure resilient cloud infrastructure.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Core Infrastructure Engineer - Scalable & Reliable
Senior Core Infrastructure Engineer - Scalable & Reliable

Oracle • Austin (TX)

On-site
USD 79,000 - 210,000
Medical, dental, and vision insurance
Paid time off
401(k) with company match
+2
Senior Distributed Systems Infra Engineer
Senior Distributed Systems Infra Engineer

Oracle • Burlington (MA)

On-site
USD 79,000 - 210,000
Medical, dental, and vision insurance
Disability insurance (short/long term)
Life insurance and AD&D
+5
Senior Core Infra Engineer - Scalable, Fault-Tolerant
Senior Core Infra Engineer - Scalable, Fault-Tolerant

Oracle • Nashville (TN)

On-site
USD 79,000 - 210,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
Senior Cloud Infra Engineer - Scalable, Reliable Systems
Senior Cloud Infra Engineer - Scalable, Reliable Systems

Oracle Corporation • Nashville (TN)

On-site
USD 79,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+11
Senior Platform Engineer, Cloud Core Infrastructure
Senior Platform Engineer, Cloud Core Infrastructure

Oracle • Santa Clara (CA)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off and holidays
+2
Principal Software Engineer, Core Infra: Scale with Equity
Principal Software Engineer, Core Infra: Scale with Equity

Oracle • United States

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Paid time off and holidays
+1
Core Infra Engineer: Scalable, Fault-Tolerant Systems
Core Infra Engineer: Scalable, Fault-Tolerant Systems

Oracle • Seattle (WA)

On-site
USD 180,000 - 240,000
Medical, dental, and vision insurance
Disability insurance
Life insurance
+6
Senior Cloud Infra Engineer: Scale & Reliability
Senior Cloud Infra Engineer: Scale & Reliability

Oracle • Nashville (TN)

On-site
USD 79,000 - 210,000
Senior Cloud Infrastructure Engineer
Senior Cloud Infrastructure Engineer

Ll Oefentherapie • Seattle (WA)

On-site
USD 79,000 - 210,000
Senior Distributed Systems Infra Engineer - Equity Eligible
Senior Distributed Systems Infra Engineer - Equity Eligible

Oracle • Austin (TX)

On-site
USD 79,000 - 210,000
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Life insurance and AD&D
+2