Principal Core Infrastructure Engineer - AI Infrastructure

Oracle Corporation

Nashville (TN)

On-site

USD 180,000 - 275,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Oracle Corporation’s Cloud Infrastructure team seeks a Principal Member of Technical Staff to own software design and development for major components, focusing on scalable distributed systems and Linux-based platform services. You will drive the architecture, implementation, and operation of critical infrastructure that powers OCI's compute and GPU resources.

As part of a collaborative, agile team, you will work with hardware and software colleagues to provision, secure, and validate server

Qualifications

  • 6-10+ years delivering and operating large scale distributed systems.
  • Strong knowledge of Object Oriented programming such as C++ or Java, and experience with scripting languages such as Python.
  • Experience with tools such as Terraform for Infrastructure as Code.
  • Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures.
  • Strong understanding of databases, NoSQL systems, storage and distributed persistence technologies.
  • Strong troubleshooting and performance tuning skills.

Responsibilities

  • Own the software design and development for major OCI Compute components.
  • Design, operate, and debug stack across Linux, Docker, Java services and Terraform.
  • Collaborate across Compute, Networking, Security, Data Center Engineering, and Hardware Development.
  • Drive reliability, scalability, and high availability for multi-tenant, virtualized infrastructure.

Skills

Distributed systems
Linux development
Systems debugging
C++/Java
Python scripting
Networking
NoSQL databases
Terraform

Tools

Terraform

Job description

As a Principal Member of Technical Staff, you will own the software design and development for major components of Oracle's Cloud Infrastructure. You should be both a rock-solid lead developer, curious problem solver, a distributed systems generalist and/or skilled Linux engineer with Systems triage experiance able to dive deep into any part of the stack and low-level systems to design broad distributed system interactions. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.

This role resides within the Compute AI Infrastructure Bare Metal Provisioning team, which owns the critical infrastructure responsible for automating the full server lifecycle from new platform shape (AMD/Intel/Arm/Nvidia) creation, hardware bring-up to customer-ready instance provisioning and firmware management. The services operate at the intersection of bare metal hardware and full-stack orchestration frameworks, a unique combination where both distributed systems engineers and engineers with background in Linux and firmware are highly valued. The team interfaces directly with components like BMCs, NICs, SmartNICs, ILOMs, GPUs, and custom firmware stacks. The team builds high performance, scalable micro-services and tooling that provision, configure, secure, and validate server platforms across OCI’s massive fleet of Compute and GPU Infrastructure. You will partner closely across other teams in Compute, Networking, Security, Data center Engineering, and Hardware Development to ensure OCI can launch, scale, and maintain new server platforms with minimal operational overhead and high reliability. You will work directly with cutting edge GPU hardware and see the direct impact of your work on the business.

We strive for equity, inclusion, and respect for all. We are committed to the greater good in our products and our actions. We are constantly learning and taking opportunities to grow our careers and ourselves. We challenge each other to stretch beyond our past to build our future.

You are the builder here. You will be part of a team of really smart, motivated, and diverse people and given the autonomy and support to do your best work. It is a dynamic and flexible workplace where you’ll belong and be encouraged.

If you are interested in building large-scale distributed infrastructure for the cloud, want to work on cutting edge GPU infrastructure and the latest Compute systems, have a knack for distributed systems and/or Linux development with Systems experiance then this is your team! Oracle is aggressively investing in the Oracle Cloud to provide the broadest, most comprehensive cloud in the industry.

Internal Responsibilities

Job Responsibilities:
You will own the software design and development for major components of Oracle’s Cloud Infrastructure. You should be both a rock solid developer, driven problem solver and a distributed systems generalist and/or Linux developer with Systems experiance able to dive deep, design, develop, operate, and debug any part of the stack and low level systems such as Linux, Docker, Java web services and Terraform, as well as design broad distributed system interactions. You should have a tenacious attitude to improve the status quo, independently seek out problems to solve and take action to deliver results wherever needed. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.

Qualifications:

  • 6-10+ years experience delivering and operating large scale, highly available distributed systems, Linux development and Systems debugging.

  • Strong knowledge of Object Oriented programming such as C++ or Java, and experience with scripting languages such as Python.

  • Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals.

  • Experience with tools such as Terraform for Infrastructure as Code.

  • Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures.

  • Strong understanding of databases, NoSQL systems, storage and distributed persistence technologies.

  • Strong troubleshooting and performance tuning skills.

  • Experience building multi-tenant, virtualized infrastructure a strong plus.

External Responsibilities

Job Responsibilities:
You will own the software design and development for major components of Oracle’s Cloud Infrastructure. You should be both a rock solid developer, driven problem solver and a distributed systems generalist and/or Linux developer with Systems experiance able to dive deep, design, develop, operate, and debug any part of the stack and low level systems such as Linux, Docker, Java web services and Terraform, as well as design broad distributed system interactions. You should have a tenacious attitude to improve the status quo, independently seek out problems to solve and take action to deliver results wherever needed. You should value simplicity and scale, work comfortably in a collaborative, agile environment, and be excited to learn.

Qualifications:

  • 6-10+ years experience delivering and operating large scale, highly available distributed systems, Linux development and Systems debugging.

  • Strong knowledge of Object Oriented programming such as C++ or Java, and experience with scripting languages such as Python.

  • Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals.

  • Experience with tools such as Terraform for Infrastructure as Code.

  • Working familiarity with networking protocols (TCP/IP, HTTP) and standard network architectures.

  • Strong understanding of databases, NoSQL systems, storage and distributed persistence technologies.

  • Strong troubleshooting and performance tuning skills.

  • Experience building multi-tenant, virtualized infrastructure a strong plus.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Core Infrastructure Engineer - AI Infrastructure
Principal Core Infrastructure Engineer - AI Infrastructure

Ll Oefentherapie • Nashville (TN)

On-site
USD 180,000 - 240,000
Software Developer 5 , AI Infrastructure
Software Developer 5 , AI Infrastructure

Ll Oefentherapie • Seattle (WA)

On-site
USD 120,000 - 160,000
Dynamic and flexible workplace
Lead Principal Software Engineer, Core Infrastructure
Lead Principal Software Engineer, Core Infrastructure

Ll Oefentherapie • Austin (TX)

On-site
USD 135,000 - 306,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off and holidays
+1
Senior Software Engineer, Core Infrastructure
Senior Software Engineer, Core Infrastructure

Emerson • Nashville (TN)

On-site
USD 79,000 - 210,000
Medical, dental, and vision
Disability insurance
Life insurance
+5
Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Oracle • Seattle (WA)

On-site
USD 146,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1
Principal Member of Technical Staff
Principal Member of Technical Staff

Ll Oefentherapie • United States

On-site
USD 96,800 - 223,400
Medical, dental, and vision insurance
401(k) Savings and Investment Plan with company match
Flexible Vacation
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Oracle • Nashville (TN)

On-site
USD 115,000 - 235,000
Medical, dental, and vision insurance
Paid time off
401(k) Savings and Investment Plan
Principal Core Infrastructure Engineer, AI Infrastructure
Principal Core Infrastructure Engineer, AI Infrastructure

Ll Oefentherapie • Nashville (TN)

On-site
USD 115,000 - 235,000
Senior Principal Software Engineer - Network & Storage Acceleration
Senior Principal Software Engineer - Network & Storage Acceleration

Oracle • United States

On-site
USD 135,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Socket.dev • Nashville (TN)

On-site
USD 180,000 - 240,000
Competitive benefits
Volunteer programs