Lead Principal Software Engineer, Core Infrastructure

Oracle

United States

On-site

USD 146,000 - 306,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Medical insurance
Dental insurance
Vision insurance
401(k) with company match
Paid time off
Parental leave

Job summary

Oracle in the United States seeks an IC5-level engineer to mentor teams and architect highly scalable, interdependent distributed systems, such as DNS and L7 proxies, delivering elastic data plane platforms. You will drive performance bottleneck identification and implement robust replication strategies for large-scale workloads.

You will oversee fault-tolerant designs, enhance security controls, and advance automation (IaC) and change management to enable safe patching and rollbacks across

Qualifications

  • 5+ years of experience in large-scale distributed systems.
  • Expertise designing scalable, fault-tolerant architectures.
  • Proficiency with performance tuning and telemetry.
  • Strong collaboration with stakeholders and incident response.

Responsibilities

  • Mentor the team in the architecture and design of scalable distributed systems.
  • Lead bottleneck identification and performance optimization for high throughput.
  • Define scalability requirements with stakeholders and ensure they meet expectations.
  • Design elastic interdependent systems with effective up/down scaling.
  • Develop advanced telemetry, dashboards, and alerting for health.

Skills

System design
Distributed systems
Performance optimization
Telemetry & monitoring
Incident response
Security controls
Automation (IaC)
Patch management

Job description

Job Description

Mentors teams and leads the architecture of highly scalable, interdependent distributed systems. Identifies and removes performance/scalability bottlenecks for hyper‑scale workloads; defines scalability requirements with stakeholders; and designs elastic, high‑impact systems while advancing innovation in data plane platforms. Engineers and oversees fault‑tolerant, in‑service‑upgradable designs; optimizes resilience mechanisms (load‑shedding, throttling, rate‑limiting); and sets SLO‑aligned durability and availability standards across dependent services.

Establishes KPIs and advanced telemetry; applies formal verification for complex features; and develops robust replication/synchronization strategies. Advises and leads resolution of complex production issues, sets operational readiness and SOP standards, and directs incident response and RCAs. Architects advanced security controls, drives remediation and compliance, and delivers enterprise‑level automation (IaC) and change strategies enabling safe, automated patching, updates, and rollbacks.

Responsibilities
System Design & Architecture - System Scalability:
  • Mentor the team in the architecture and design of highly scalable, interdependent distributed systems such as DNS and L7 Proxy, ensuring horizontal and vertical scalability and overall performance, including leveraging distributed state management tools.
  • Lead the identification of performance and scalability bottlenecks and recommend solutions to optimize code and/or systems for large-scale data processing and high-throughput requirements to improve performance for hyper‑scale systems.
  • Lead collaboration with stakeholders to define system scalability requirements, ensuring the defined requirements meet customer expectations.
  • Leverage deep expertise to design high-impact, interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).
  • Drive innovation in the use of data plane platforms.
  • Evaluate whether systems are meeting nonfunctional scalability requirements, and proactively anticipate growing business needs within the business unit.
System Reliability Performance:
  • Define key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running, interdependent systems.
  • Drive the creation and customization of highly complex dashboards, telemetry systems, and alerting mechanisms, proactively ensuring system health and reliability.
System Reliability Design:
  • Design and oversee the implementation of fault‑tolerant, interdependent systems capable of withstanding in‑service updates by implementing sophisticated redundancy, replication, and automatic failover capabilities.
  • Lead the design and implementation of systems that effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
  • Guide the optimization of advanced mechanisms to handle network unreliability, including load‑shedding, throttling, and rate‑limiting.
  • Design interdependent systems that are durable and adhere to service level objectives (SLOs), driving standards for availability and durability of other computing services within the organization.
Correctness / Availability:
  • Maintain expertise in industry standards for verifying correctness and apply existing techniques to interdependent systems. Formally verify complex features (e.g., via TLA+) to ensure system design correctness for various interdependent systems.
  • Develop advanced strategies for data replication and synchronization, ensuring robust data integrity and availability.
Operational Troubleshooting & Incident Management:
  • Advise on efforts to diagnose, debug, and resolve complex issues in active, interdependent systems to support ongoing operation.
  • Develop and implement comprehensive strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
  • Maintain expertise in dependencies, dependents, and owned systems to drive effective troubleshooting and performance.
  • Set standards for operational readiness and standard operating procedures within the department, and hold third‑party partners accountable for meeting those standards.
  • Oversee operational support rotations, providing expert guidance in incident response and leading root cause investigations to prevent future occurrences.
Compliance & Security:
  • Architect advanced security measures to protect data and applications in multi‑tenant environments, and lead initiatives to enhance data and application protection.
  • Guide the execution of comprehensive remediation plans to address identified security vulnerabilities.
  • Ensure cloud infrastructure is in compliance with industry standards and regulations, and guide documentation efforts across projects.
Automation & Change Management:
  • Develop enterprise‑level automation tools and strategies (e.g., Infrastructure as Code (IaC)) and oversee their implementation.
  • Drive alignment of change management plans and organizational initiatives for patching, updating, and rolling back applications, and design interdependent systems to allow for automation of these processes.
Qualifications

Disclaimer:

Certain U.S. based or U.S. customer or client‑facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.

Range and benefit information provided in this posting are specific to the stated locations only.

US: Hiring Range in USD from: $146,300 - $306,400 per year. May be eligible for bonus, equity, and compensation deferral.

Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.

Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.

Oracle US offers a comprehensive benefits package which includes the following:

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance

The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - IC5

About Us

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life‑saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Oracle • Frankfort (KY)

On-site
USD 146,000 - 306,000
Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Oracle • Nashville (TN)

On-site
USD 146,000 - 306,000
Medical, dental, and vision insurance
Paid time off
401(k) plan with match
Lead Principal Core Infrastructure Engineer
Lead Principal Core Infrastructure Engineer

Oracle • United States

On-site
USD 146,000 - 306,000
Medical insurance
Dental insurance
Vision insurance
+4
Senior Engineer, Core Infrastructure
Senior Engineer, Core Infrastructure

Oracle • United States

On-site
USD 79,000 - 210,000
Principal Software Engineer, Core Infrastructure
Principal Software Engineer, Core Infrastructure

Oracle • Austin (TX)

On-site
USD 115,000 - 235,000
Medical benefits
Dental & Vision
401(k) match
+1
Senior Software Engineer, Core Infrastructure
Senior Software Engineer, Core Infrastructure

Oracle • Austin (TX)

On-site
USD 79,000 - 210,000
Medical, dental, and vision insurance
Paid time off
401(k) with company match
+2
Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Oracle • Seattle (WA)

On-site
USD 79,000 - 210,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off and holidays
+1
Lead Principal Systems Software Engineer
Lead Principal Systems Software Engineer

Oracle • Nashville (TN)

On-site
USD 184,000 - 306,000
Medical insurance
Dental insurance
401(k) plan
+1
Director, Core Infrastructure Engineering
Director, Core Infrastructure Engineering

Ll Oefentherapie • Nashville (TN)

On-site
USD 170,000 - 355,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1
Senior Software Engineer, Core Infrastructure
Senior Software Engineer, Core Infrastructure

Oracle • Nashville (TN)

On-site
USD 79,000 - 210,000
Health insurance
401(k) plan
Paid time off
+1