Principal Software Engineer, Core Infrastructure

Oracle Corporation

Nashville (TN)

On-site

USD 150,000 - 210,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Relocation assistance

Job summary

Oracle Cloud Infrastructure (OCI) seeks a Principal Core Infrastructure Engineer to lead the design and evolution of foundational distributed systems behind OCI in Nashville. You will build scalable, elastic, and fault-tolerant services for high-volume data retrieval, storage, and processing, setting reliability and security standards.

The role emphasizes on-site work in Nashville, with relocation assistance potentially available.

Qualifications

  • Bachelor's or master's degree in Computer Science, Computer Engineering, or related field, or equivalent practical experience.
  • 7+ years of professional software-engineering experience on large-scale distributed systems or cloud infrastructure.
  • Strong experience designing and operating highly available, scalable, fault-tolerant distributed systems.
  • Proficiency in Java, C++, C#, or Go.
  • Deep understanding of distributed-systems design, data structures, algorithms, OS, networking, and secure software-development practices.
  • Experience with system-level test automation and production incident response.
  • Demonstrated experience leading or influencing technical architecture and mentoring engineers.
  • Strong problem-solving, communication, and cross-functional collaboration skills.

Responsibilities

  • Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale.
  • Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them.
  • Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.
  • Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design.
  • Apply distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, retries, and timeouts.
  • Establish service-level objectives, KPIs, telemetry, dashboards, and proactive alerting for critical systems.
  • Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness.
  • Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence.
  • Build and improve Infrastructure as Code and operational automation for safe patching, updates, rollbacks, and change management.
  • Apply robust security controls for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.

Skills

Distributed systems
Cloud infrastructure
Java
C++
C#
Go
System-level testing
Incident response
Mentoring
Communication

Education

Bachelor's or Master's degree in CS/Engineering or equivalent
Equivalent practical experience

Tools

Infrastructure as Code

Job description

Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for leading enterprises worldwide. Our cloud offers hyperscale, multi-tenant services deployed across more than 50 regions globally. OCI continues to expand beyond traditional public-cloud boundaries to support dedicated, hybrid, and multicloud solutions, edge computing, and more.

As a Principal Core Infrastructure Engineer, you will lead the design and evolution of foundational distributed systems behind OCI. You will build highly scalable, elastic, and fault-tolerant services for high-volume data retrieval, storage, and processing, and set the technical direction for their reliability, correctness, security, and operational readiness.

This position is office-based and requires onsite presence in Nashville, Tennessee. Relocation assistance may be available in accordance with Oracle's relocation policies.

Internal Responsibilities

What You'll Do

  • Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale.
  • Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them.
  • Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.
  • Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design.
  • Apply sound distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, rate limiting, retries, and timeouts.
  • Establish service-level objectives, key performance indicators, telemetry, dashboards, and proactive alerting for critical systems.
  • Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness.
  • Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence.
  • Build and improve Infrastructure as Code and operational automation that enable safe patching, updates, rollbacks, and change management.
  • Apply robust security controls and remediation practices for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.

What You'll Bring

  • Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
  • 7+ years of professional software-engineering experience, with demonstrated impact on large-scale distributed systems or cloud infrastructure.
  • Strong experience designing and operating highly available, scalable, fault-tolerant distributed systems.
  • Proficiency in one or more object-oriented or systems programming languages, such as Java, C++, C#, or Go.
  • Deep understanding of distributed-systems design, data structures, algorithms, operating systems, networking, and secure software-development practices.
  • Experience with system-level test automation, performance/load testing, reliability engineering, and production incident response.
  • Demonstrated experience leading or influencing technical architecture and mentoring engineers.
  • Strong problem-solving, communication, and cross-functional collaboration skills.

Preferred Qualifications

  • Experience with Oracle Cloud, AWS, Azure, Google Cloud, or other large-scale cloud platforms.
  • Experience with data-plane platforms, distributed storage, microservices, replication, state management, or high-throughput data processing.
  • Experience defining SLOs, building observability systems, and operating services in a 24x7 production environment.
  • Experience with Infrastructure as Code, service automation, security controls, and compliance requirements for cloud infrastructure.
External Responsibilities

What You'll Do

  • Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale.
  • Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them.
  • Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.
  • Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design.
  • Apply sound distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, rate limiting, retries, and timeouts.
  • Establish service-level objectives, key performance indicators, telemetry, dashboards, and proactive alerting for critical systems.
  • Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness.
  • Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence.
  • Build and improve Infrastructure as Code and operational automation that enable safe patching, updates, rollbacks, and change management.
  • Apply robust security controls and remediation practices for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.

What You'll Bring

  • Bachelor's or master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
  • 7+ years of professional software-engineering experience, with demonstrated impact on large-scale distributed systems or cloud infrastructure.
  • Strong experience designing and operating highly available, scalable, fault-tolerant distributed systems.
  • Proficiency in one or more object-oriented or systems programming languages, such as Java, C++, C#, or Go.
  • Deep understanding of distributed-systems design, data structures, algorithms, operating systems, networking, and secure software-development practices.
  • Experience with system-level test automation, performance/load testing, reliability engineering, and production incident response.
  • Demonstrated experience leading or influencing technical architecture and mentoring engineers.
  • Strong problem-solving, communication, and cross-functional collaboration skills.

Preferred Qualifications

  • Experience with Oracle Cloud, AWS, Azure, Google Cloud, or other large-scale cloud platforms.
  • Experience with data-plane platforms, distributed storage, microservices, replication, state management, or high-throughput data processing.
  • Experience defining SLOs, building observability systems, and operating services in a 24x7 production environment.
  • Experience with Infrastructure as Code, service automation, security controls, and compliance requirements for cloud infrastructure.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Platform
Principal Software Engineer, Platform

Oracle Corporation • Nashville (TN)

On-site
USD 150,000 - 210,000
Relocation assistance
Senior Software Engineer, Platform
Senior Software Engineer, Platform

Oracle Corporation • Nashville (TN)

On-site
USD 140,000 - 190,000
Principal Software Engineer, Core Infrastructure - Nashville
Principal Software Engineer, Core Infrastructure - Nashville

Oracle • Nashville (TN)

On-site
USD 115,000 - 235,000
Medical Insurance
Dental Vision
401k Match
+5
Senior Software Engineer, Core Infrastructure - Nashville
Senior Software Engineer, Core Infrastructure - Nashville

Oracle • Nashville (TN)

On-site
USD 79,000 - 210,000
Medical & Dental Insurance
401(k) Plan with company match
Paid time off and holidays
OCI Core Infrastructure Engineer 2 - Nashville Campus
OCI Core Infrastructure Engineer 2 - Nashville Campus

Ll Oefentherapie • Nashville (TN)

On-site
USD 120,000 - 170,000
Principal Core Infrastructure Engineer - AI Infrastructure
Principal Core Infrastructure Engineer - AI Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 180,000 - 275,000
OCI Senior Core Infrastructure Engineer - Nashville TN
OCI Senior Core Infrastructure Engineer - Nashville TN

Ll Oefentherapie • Nashville (TN)

On-site
USD 120,000 - 180,000
OCI Core Infrastructure Engineer 2 - Nashville Campus
OCI Core Infrastructure Engineer 2 - Nashville Campus

Oracle • Nashville (TN)

On-site
USD 70,000 - 166,000
Medical, dental, and vision
401(k) with company match
Paid time off
+1
Principal Platform Software Engineer (OCI - Developer Platform)
Principal Platform Software Engineer (OCI - Developer Platform)

Ll Oefentherapie • Nashville (TN)

On-site
USD 140,000 - 210,000
OCI Senior Core Infrastructure Engineer - Nashville TN
OCI Senior Core Infrastructure Engineer - Nashville TN

Oracle • Nashville (TN)

On-site
USD 79,000 - 210,000
Medical, dental, vision insurance
401(k) with company match
Paid time off & holidays
+1