AI Infrastructure Engineering Leader — GPU Clusters

Oracle

Seattle (WA)

On-site

USD 121,500 - 306,400

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Short-term and long-term disability
Life insurance and AD&D
401(k) Savings and Investment Plan

Job summary

Oracle Cloud Infrastructure is seeking an experienced Core Infrastructure Engineering Leader to guide a high-performing team responsible for delivering healthy GPU/AI/ML infrastructure with optimal performance. This role drives end-to-end customer execution, including POCs, troubleshooting, automation of provisioning/configuration/monitoring, and post-sale support to streamline operations.

You will design and deploy automated GPU cluster tooling, collaborate with OCI Services and sales, and

Qualifications

  • 7+ years of senior software engineering leadership or related experience.
  • Strong communication and collaboration skills, with the ability to work effectively in cross-functional teams and convey technical concepts to non-technical stakeholders.
  • Demonstrated leadership and people management skills.
  • Proven at building and managing distributed/cloud software engineering solutions.
  • BS or MS degree or equivalent experience relevant to the functional area.
  • Experience using tools like Ansible, Terraform, Python, containerization technologies (e.g., Docker, Kubernetes) and orchestration tools.
  • Solid understanding of networking concepts, security principles, and best practices.
  • Excellent problem-solving skills, with the ability to troubleshoot complex issues and drive resolution in a fast-paced environment.
  • Strong Linux skills with hands-on experience in Oracle Linux/RHEL/CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization.
  • Strong proficiency in at least one of the programming languages such as Python, Rust, Go, Java, or Scala.
  • Proven experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads.

Responsibilities

  • Lead, mentor, and develop a team of Core Infrastructure Engineers responsible for designing, implementing, and maintaining the infrastructure that supports our largest GPU/AI/ML customers.
  • Drive the design, development, testing, validation, and deployment readiness of our automated GPU Cluster deployment tool (like AWS Parallel Cluster, Azure Cycle Cloud) with Slurm and/or Oracle Kubernetes Engine (OKE) to streamline operations and enhance productivity.
  • Build collaborative relationships with OCI Services team, customer and sales team to deliver reliable, scalable infrastructures. Act as a technical liaison between customers, core engineering teams, and support.
  • Work with OCI Strategic customers to grow our business in pre/post sales stages in a technical infra expert role.
  • Take ownership of problems and work to identify solutions. Ability to think through the solution and identify/document potential issues impacting your customers.
  • Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre‑processing techniques.
  • Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime.
  • Document infrastructure designs, configurations, and procedures to facilitate knowledge sharing and ensure maintainability.
  • As a trusted customer advocate, you will help customers/partners understand best practices around advanced GPU solutions, and how to migrate their workloads to the cloud.
  • Educate customers of all sizes on the value proposition of Oracle Cloud and participate in deep architectural discussions to ensure solutions are designed for successful deployment in the cloud.

Skills

Leadership
Communication
People management
Distributed/cloud infra
Infrastructure design
Python
Docker/Kubernetes
Ansible/Terraform
Linux systems
AI/ML/HPC infra
Cross-functional collaboration

Education

BS or MS degree or equivalent

Tools

Ansible
Terraform
Docker
Kubernetes
Python

Job description

Oracle Cloud Infrastructure is seeking an experienced Core Infrastructure Engineering Leader to guide a high-performing team responsible for delivering healthy GPU/AI/ML infrastructure with optimal performance. This role drives end-to-end customer execution, including POCs, troubleshooting, automation of provisioning/configuration/monitoring, and post-sale support to streamline operations.

You will design and deploy automated GPU cluster tooling, collaborate with OCI Services and sales, and

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops

Oracle • United States

On-site
USD 121,000 - 307,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
Head of AI/ML Core Infra & GPU Cluster Ops
Head of AI/ML Core Infra & GPU Cluster Ops

Oracle • Denver (CO)

On-site
USD 122,000 - 306,000
Head of AI/ML Core Infrastructure & GPU Cluster Ops
Head of AI/ML Core Infrastructure & GPU Cluster Ops

Oracle • Salt Lake City (UT)

On-site
USD 122,000 - 306,000
Medical, dental, and vision insurance
Disability insurance
Life insurance
+7
Director of AI/ML Core Infrastructure
Director of AI/ML Core Infrastructure

Oracle • San Juan (PR)

On-site
USD 122,000 - 306,000
Director, AI/ML Core Infrastructure
Director, AI/ML Core Infrastructure

Oracle • Columbia (SC)

On-site
USD 122,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off and holidays
Director, AI Core Infra Engineering
Director, AI Core Infra Engineering

Oracle • Lansing (MI)

On-site
USD 122,000 - 306,000
Medical, dental, and vision insurance
Bonus and equity potential
Paid time off and holidays
+1
Strategic AI/ML GPU Infrastructure Director
Strategic AI/ML GPU Infrastructure Director

Oracle • Atlanta (GA)

Hybrid
USD 122,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1
OCI AI & GPU HPC Infrastructure Architect
OCI AI & GPU HPC Infrastructure Architect

Ll Oefentherapie • United States

On-site
USD 180,000 - 240,000
Senior Cloud GPU Infrastructure Engineer
Senior Cloud GPU Infrastructure Engineer

Oracle • United States

On-site
USD 183,000 - 307,000
Medical Insurance
Disability Insurance
Life Insurance
+5
OCI AI & HPC Infrastructure Architect
OCI AI & HPC Infrastructure Architect

Oracle • United States

On-site
USD 84,000 - 210,000