Director, AI Core Infra Engineering

Oracle

Lansing (MI)

On-site

USD 122,000 - 306,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical, dental, and vision insurance
Bonus and equity potential
Paid time off and holidays
Comprehensive benefits package

Job summary

Oracle is seeking a seasoned Core Infrastructure Engineering Leader to guide a high-performing team responsible for delivering healthy, high‑performance GPU/AI/ML infrastructure. You will own end‑to‑end technical customer execution, from POCs to post‑sale support, and build automation tools to streamline provisioning and monitoring.

You will partner with OCI Services and sales to provide reliable, scalable infrastructure, tune performance, and mentor engineers.

Qualifications

  • 7 years of senior software engineering leadership or related experience.
  • Strong communication and collaboration skills, with cross-functional teamwork.
  • Demonstrated leadership and people management skills.
  • Proven experience building distributed/cloud software engineering solutions.
  • Experience with automation and cloud tooling (Ansible, Terraform, Python, Docker, Kubernetes).
  • BS or MS degree or equivalent experience.

Responsibilities

  • Lead, mentor, and develop a team of Core Infrastructure Engineers responsible for designing, implementing, and maintaining the infrastructure that supports our largest GPU/AI/ML customers.
  • Drive the design, development, testing, validation, and deployment readiness of our automated GPU Cluster deployment tool with Slurm and/or Oracle Kubernetes Engine (OKE).
  • Build collaborative relationships with OCI Services team, customer and sales team to deliver reliable, scalable infrastructures. Act as a technical liaison between customers, core engineering teams, and support.
  • Work with OCI Strategic customers to grow our business in pre/post sales stages in a technical infra expert role.
  • Take ownership of problems and work to identify solutions. Ability to think through the solution and identify/document potential issues impacting your customers.
  • Optimize infrastructure performance by tuning parameters, optimizing resource utilization, and implementing caching and data pre-processing techniques.
  • Troubleshoot infrastructure performance, scalability, and reliability issues and implement solutions to mitigate risks and minimize downtime.
  • Document infrastructure designs, configurations, and procedures to facilitate knowledge sharing and ensure maintainability.
  • As a trusted customer advocate, you will help customers/partners understand best practices around advanced GPU solutions, and how to migrate their workloads to the cloud.
  • Educate customers of all sizes on the value proposition of Oracle Cloud and participate in deep architectural discussions to ensure solutions are designed for successful deployment in the cloud.

Skills

Leadership
Cloud infrastructure
Cross-functional communication
People management
Problem solving

Education

BS or MS degree in a relevant field

Tools

Ansible
Terraform
Python
Docker
Kubernetes
Networking concepts

Job description

Oracle is seeking a seasoned Core Infrastructure Engineering Leader to guide a high-performing team responsible for delivering healthy, high‑performance GPU/AI/ML infrastructure. You will own end‑to‑end technical customer execution, from POCs to post‑sale support, and build automation tools to streamline provisioning and monitoring.

You will partner with OCI Services and sales to provide reliable, scalable infrastructure, tune performance, and mentor engineers.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Director, AI/ML Core Infrastructure
Director, AI/ML Core Infrastructure

Oracle • Columbia (SC)

On-site
USD 122,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off and holidays
Director of AI/ML Core Infrastructure
Director of AI/ML Core Infrastructure

Oracle • San Juan (PR)

On-site
USD 122,000 - 306,000
Head of AI/ML Core Infra & GPU Cluster Ops
Head of AI/ML Core Infra & GPU Cluster Ops

Oracle • Denver (CO)

On-site
USD 122,000 - 306,000
Head of AI/ML Core Infrastructure & GPU Cluster Ops
Head of AI/ML Core Infrastructure & GPU Cluster Ops

Oracle • Salt Lake City (UT)

On-site
USD 122,000 - 306,000
Medical, dental, and vision insurance
Disability insurance
Life insurance
+7
Strategic AI/ML GPU Infrastructure Director
Strategic AI/ML GPU Infrastructure Director

Oracle • Atlanta (GA)

Hybrid
USD 122,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1
AI Infrastructure Engineering Leader — GPU Clusters
AI Infrastructure Engineering Leader — GPU Clusters

Oracle • Seattle (WA)

On-site
USD 121,000 - 307,000
Medical, dental, and vision insurance
Short-term and long-term disability
Life insurance and AD&D
+1
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops

Oracle • United States

On-site
USD 121,000 - 307,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
OCI AI & GPU HPC Infrastructure Architect
OCI AI & GPU HPC Infrastructure Architect

Ll Oefentherapie • United States

On-site
USD 180,000 - 240,000
Senior AI/ML Infra Architect & Cloud Engineer
Senior AI/ML Infra Architect & Cloud Engineer

Oracle • San Juan (PR)

On-site
USD 85,000 - 210,000
Medical, dental, vision insurance
401(k) with company match
Paid time off
Senior AI/ML Infra Architect
Senior AI/ML Infra Architect

Oracle • Lansing (MI)

On-site
USD 85,000 - 210,000