Lead Principal Software Engineer, Core Infrastructure

Oracle Corporation

Seattle (WA)

On-site

USD 180,000 - 240,000

Full time

4 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Oracle Corporation's OCI team seeks a senior IC5 software engineer with deep Kubernetes expertise to lead platform initiatives across the OKE and OCI infrastructure. You will design, build, and operate distributed systems that manage Kubernetes clusters at scale, ensuring reliability, security, and performance.

You will mentor engineers, drive cross-team execution, and apply AI-assisted workflows to accelerate delivery while maintaining production accountability and high code quality.

Qualifications

  • 10+ years of software engineering experience in production systems.
  • Hands-on cloud infrastructure experience on OCI, AWS, Azure, or GCP.
  • Strong Kubernetes expertise including control plane, APIs, scheduling, and networking.

Responsibilities

  • Provide technical leadership for major OKE platform initiatives from architecture through implementation and production operation.
  • Design and build distributed systems to create, update, scale, repair, and operate Kubernetes clusters across OCI regions.
  • Lead technical design reviews, code reviews, and production readiness reviews for complex service changes.

Skills

Kubernetes
Cloud infrastructure
Distributed systems
Go
Java
Linux
Production operations
Mentor engineers

Tools

CSI drivers
CNI
Karpenter

Job description

The Oracle Cloud Infrastructure (OCI) team offers the opportunity to build and operate massive-scale, integrated cloud services in a broadly distributed, multi-tenant cloud environment. OCI builds cloud products for customers who are tackling some of the world's largest technical and business challenges.

Oracle Kubernetes Engine (OKE) is OCI's managed Kubernetes service. OKE enables customers to create, run, scale, secure, and operate Kubernetes clusters on OCI, integrating Kubernetes with OCI compute, networking, storage, identity, observability, security, and automation. The OKE team owns a highly available 24x7 cloud service and is expanding the platform to support larger clusters, higher scale, improved operability, deeper OCI integrations, and increasingly demanding cloud native, AI, and GPU workloads.

We are looking for a senior IC5 software engineer with deep Kubernetes expertise, required cloud infrastructure experience, and a strong distributed systems background. This is a high-impact technical leadership role for an engineer who can define architecture, drive cross-team execution, solve ambiguous production and platform problems, and deliver durable systems that improve both customer experience and operational excellence.

You will work on core OKE platform capabilities including cluster lifecycle management, orchestration, scalability, reliability, performance, automation, observability, security, and integration with OCI infrastructure services. The ideal candidate has hands-on experience designing, building, operating, or deeply debugging production cloud services, infrastructure platforms, or Kubernetes-based systems at meaningful scale.

This role requires advanced Kubernetes experience, including Kubernetes control plane behavior, controllers and operators, scheduling, autoscaling, networking, storage, service discovery, container runtimes, node lifecycle, Kubernetes APIs, and etcd. Experience with Kubernetes networking and storage technologies such as CNI, Cilium, Calico, Flannel, other container networking implementations, CSI drivers, and cloud provider integrations is highly relevant.

OKE is also expanding to support demanding AI and accelerated computing use cases. Experience with AI/ML infrastructure, multi-node GPU clusters, accelerated compute, model training or inference platforms, GPU scheduling, device plugins, Karpenter, cluster autoscaling, CUDA, NCCL, RoCE, InfiniBand, RDMA, SmartNIC/DPU offload, or high-performance AI/HPC networking is a significant plus.

This role also requires an engineer who is ready to use modern agentic engineering practices responsibly. We expect senior engineers to apply AI-assisted and agentic workflows to accelerate design exploration, implementation, testing, debugging, documentation, operational analysis, and developer productivity while maintaining strong ownership, security judgment, code quality, and production accountability.

Internal Responsibilities

As a member of the software engineering division, you will take an active role in defining and evolving standard practices and procedures. You will define specifications for significant new projects and specify, design, develop, troubleshoot, and debug software for OCI's managed Kubernetes service.

Responsibilities include:

  • Provide technical leadership for major OKE platform initiatives from architecture through implementation, launch, and production operation.
  • Design and build distributed systems that create, update, scale, repair, and operate Kubernetes clusters across OCI regions.
  • Improve OKE reliability, scalability, performance, upgrade safety, lifecycle management, observability, automation, and operational tooling.
  • Work deeply with Kubernetes technologies, including control plane components, controllers/operators, scheduling, autoscaling, Kubernetes APIs, container runtimes, node behavior, and etcd.
  • Design, debug, and improve Kubernetes networking and storage integrations, including CNI-based networking, Cilium, Calico, Flannel, other container networking implementations, CSI drivers, and OCI infrastructure integrations.
  • Build automation for cluster validation, health checks, readiness testing, failure detection, remote recovery, and reduction of post-deployment operational issues.
  • Lead technical design reviews, code reviews, incident reviews, and production readiness reviews for complex service changes.
  • Debug difficult production issues across service boundaries, including Kubernetes, Linux, networking, compute, storage, identity, telemetry, and OCI infrastructure dependencies.
  • Apply performance engineering practices including profiling, tracing, latency analysis, throughput optimization, and production diagnostics across distributed systems.
  • Build automation that reduces manual operations, improves fleet health, accelerates diagnosis, and raises the quality bar for OKE engineering.
  • Partner with OCI service teams to deliver end-to-end platform capabilities regardless of organizational boundaries.
  • Apply AI-assisted and agentic engineering workflows to improve engineering velocity, test coverage, debugging, operational analysis, and documentation while ensuring correctness, security, and maintainability.
  • Mentor engineers, influence technical direction, and help establish patterns that scale across the OKE organization.
  • Participate in operating a 24x7 cloud service and use customer feedback, production data, and operational experience to prioritize improvements.

Required qualifications:

  • 10+ years of software engineering experience, or equivalent experience building and operating production software systems.
  • Hands-on cloud infrastructure experience is required, ideally designing, building, operating, or debugging production services or platforms on OCI, AWS, Azure, GCP, or a large-scale private cloud.
  • Strong hands-on Kubernetes expertise is required, including Kubernetes architecture, APIs, control plane behavior, controllers/operators, scheduling, autoscaling, networking, storage, nodes, cluster lifecycle management, or production cluster operations.
  • Advanced Kubernetes knowledge, including CNI, CSI, etcd, service discovery, container runtimes, node lifecycle, and Kubernetes failure modes.
  • Experience with Kubernetes networking technologies such as Cilium, Calico, Flannel, or other CNI implementations.
  • Experience with Kubernetes storage integrations, including CSI drivers or cloud storage integrations.
  • Strong distributed systems fundamentals, including availability, failure handling, performance, scalability, and operational tradeoffs.
  • Experience building highly available infrastructure services, platform services, or cloud native systems used in production.
  • Strong development experience in both Go/Golang and Java is required.
  • Strong Linux, networking, debugging, and production operations skills.
  • Demonstrated ability to lead ambiguous technical projects, influence across teams, and deliver through other engineers without relying on formal authority.
  • Strong communication skills, ownership, judgment, and ability to make pragmatic tradeoffs in production systems.

Preferred qualifications:

  • Experience with AI/ML infrastructure, GPU workloads, multi-node GPU clusters, accelerated compute, model training or inference platforms, GPU scheduling, device plugins, Karpenter, cluster autoscaling, CUDA, NCCL, high-performance networking, or distributed training systems.
  • Experience with eBPF-based networking, Kubernetes network policy, service mesh, ingress, load balancing, overlays/underlays, BGP, VXLAN, SmartNIC/DPU offload, RoCE, InfiniBand, RDMA, or multi-cluster networking.
  • Experience with infrastructure as code and cloud provisioning tools such as Terraform, Packer, cloud-init, IAM, VCN/VPC networking, VPN, FastConnect/Direct Connect, or equivalent cloud primitives.
  • Experience building developer productivity, operational automation, or responsible AI-assisted and agentic engineering workflows.
  • Experience with observability systems, incident response, safe deployment practices, canary analysis, rollback strategies, service health automation, and large fleet operations.
  • Open-source or upstream contribution experience in Kubernetes, cloud native infrastructure, observability, networking, or related systems.
External Responsibilities

As a member of the software engineering division, you will take an active role in defining and evolving standard practices and procedures. You will define specifications for significant new projects and specify, design, develop, troubleshoot, and debug software for OCI's managed Kubernetes service.

Responsibilities include:

  • Provide technical leadership for major OKE platform initiatives from architecture through implementation, launch, and production operation.
  • Design and build distributed systems that create, update, scale, repair, and operate Kubernetes clusters across OCI regions.
  • Improve OKE reliability, scalability, performance, upgrade safety, lifecycle management, observability, automation, and operational tooling.
  • Work deeply with Kubernetes technologies, including control plane components, controllers/operators, scheduling, autoscaling, Kubernetes APIs, container runtimes, node behavior, and etcd.
  • Design, debug, and improve Kubernetes networking and storage integrations, including CNI-based networking, Cilium, Calico, Flannel, other container networking implementations, CSI drivers, and OCI infrastructure integrations.
  • Build automation for cluster validation, health checks, readiness testing, failure detection, remote recovery, and reduction of post-deployment operational issues.
  • Lead technical design reviews, code reviews, incident reviews, and production readiness reviews for complex service changes.
  • Debug difficult production issues across service boundaries, including Kubernetes, Linux, networking, compute, storage, identity, telemetry, and OCI infrastructure dependencies.
  • Apply performance engineering practices including profiling, tracing, latency analysis, throughput optimization, and production diagnostics across distributed systems.
  • Build automation that reduces manual operations, improves fleet health, accelerates diagnosis, and raises the quality bar for OKE engineering.
  • Partner with OCI service teams to deliver end-to-end platform capabilities regardless of organizational boundaries.
  • Apply AI-assisted and agentic engineering workflows to improve engineering velocity, test coverage, debugging, operational analysis, and documentation while ensuring correctness, security, and maintainability.
  • Mentor engineers, influence technical direction, and help establish patterns that scale across the OKE organization.
  • Participate in operating a 24x7 cloud service and use customer feedback, production data, and operational experience to prioritize improvements.

Required qualifications:

  • 10+ years of software engineering experience, or equivalent experience building and operating production software systems.
  • Hands-on cloud infrastructure experience is required, ideally designing, building, operating, or debugging production services or platforms on OCI, AWS, Azure, GCP, or a large-scale private cloud.
  • Strong hands-on Kubernetes expertise is required, including Kubernetes architecture, APIs, control plane behavior, controllers/operators, scheduling, autoscaling, networking, storage, nodes, cluster lifecycle management, or production cluster operations.
  • Advanced Kubernetes knowledge, including CNI, CSI, etcd, service discovery, container runtimes, node lifecycle, and Kubernetes failure modes.
  • Experience with Kubernetes networking technologies such as Cilium, Calico, Flannel, or other CNI implementations.
  • Experience with Kubernetes storage integrations, including CSI drivers or cloud storage integrations.
  • Strong distributed systems fundamentals, including availability, failure handling, performance, scalability, and operational tradeoffs.
  • Experience building highly available infrastructure services, platform services, or cloud native systems used in production.
  • Strong development experience in both Go/Golang and Java is required.
  • Strong Linux, networking, debugging, and production operations skills.
  • Demonstrated ability to lead ambiguous technical projects, influence across teams, and deliver through other engineers without relying on formal authority.
  • Strong communication skills, ownership, judgment, and ability to make pragmatic tradeoffs in production systems.

Preferred qualifications:

  • Experience with AI/ML infrastructure, GPU workloads, multi-node GPU clusters, accelerated compute, model training or inference platforms, GPU scheduling, device plugins, Karpenter, cluster autoscaling, CUDA, NCCL, high-performance networking, or distributed training systems.
  • Experience with eBPF-based networking, Kubernetes network policy, service mesh, ingress, load balancing, overlays/underlays, BGP, VXLAN, SmartNIC/DPU offload, RoCE, InfiniBand, RDMA, or multi-cluster networking.
  • Experience with infrastructure as code and cloud provisioning tools such as Terraform, Packer, cloud-init, IAM, VCN/VPC networking, VPN, FastConnect/Direct Connect, or equivalent cloud primitives.
  • Experience building developer productivity, operational automation, or responsible AI-assisted and agentic engineering workflows.
  • Experience with observability systems, incident response, safe deployment practices, canary analysis, rollback strategies, service health automation, and large fleet operations.
  • Open-source or upstream contribution experience in Kubernetes, cloud native infrastructure, observability, networking, or related systems.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Core Infrastructure
Principal Software Engineer, Core Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 150,000 - 210,000
Relocation assistance
Senior Kubernetes Engineer
Senior Kubernetes Engineer

Ashley Furniture Industries • Tampa (FL)

On-site
USD 150,000 - 190,000
Director, Core Infrastructure Engineering
Director, Core Infrastructure Engineering

Oracle • United States

On-site
USD 170,000 - 355,000
Medical insurance
Dental insurance
Vision insurance
+2
Director, Core Infrastructure Engineering
Director, Core Infrastructure Engineering

Ll Oefentherapie • Seattle (WA)

On-site
USD 170,000 - 355,000
Medical, dental, and vision insurance
Paid time off
401(k) match
+1
Senior Platform Software Engineer-Agentic AI Project
Senior Platform Software Engineer-Agentic AI Project

Oracle Corporation • Nashville (TN)

On-site
USD 120,000 - 180,000
OCI Senior Core Infrastructure Engineer - Nashville TN
OCI Senior Core Infrastructure Engineer - Nashville TN

Ll Oefentherapie • Nashville (TN)

On-site
USD 120,000 - 180,000
Principal Software Engineer
Principal Software Engineer

Oracle Corporation • Nashville (TN)

On-site
USD 140,000 - 210,000
Principal Engineer, Core Infrastructure
Principal Engineer, Core Infrastructure

Ll Oefentherapie • Seattle (WA), San Francisco (CA)

On-site
USD 180,000 - 250,000
Principal Core Infrastructure Engineer - AI Infrastructure
Principal Core Infrastructure Engineer - AI Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 180,000 - 275,000
Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Oracle • Seattle (WA)

On-site
USD 146,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1