Manager, Core Infrastructure Engineering

Oracle Corporation

Nashville (TN)

On-site

USD 150,000 - 190,000

Full time

5 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Oracle Corporation seeks a Manager to lead a class of early‑career Infrastructure Engineers for OCI AI/ML workloads. You will mentor engineers, shape onboarding for GPU clusters, and collaborate across OCI teams, customers, and sales to deliver reliable cloud infrastructure. Strong leadership and cloud ops experience are essential.

The role emphasizes hands-on guidance, problem solving, and scalable processes to enable engineers to independently support AI/ML workloads on OCI.

Qualifications

  • This role requires leading or mentoring engineers in a cloud infra setting.
  • Experience designing and operating distributed, customer-facing cloud infrastructure.
  • Strong communication and collaboration skills for coaching juniors.

Responsibilities

  • Lead, mentor, and develop a class of early-career engineers with clear growth plans.
  • Design onboarding programs for GPU cluster deployment and related tooling.
  • Provide coaching, feedback, and stretch assignments to build judgement.
  • Build cross-functional partnerships with OCI Services, customers, and sales.
  • Document designs, runbooks, and procedures for repeatable practices.

Skills

Leadership
Mentoring
Communication
Cloud infrastructure
Distributed systems
On-call / troubleshooting
Programming
Linux administration
Automation tooling
Networking & security
Stakeholder management

Education

BS or MS degree or equivalent experience

Tools

Ansible
Terraform
Python
Docker
Kubernetes

Job description

Job Description

Oracle Cloud Infrastructure (OCI) is building some of the world's largest and most advanced GPU clusters to power the next generation of AI. The Strategic Customer Engineering (SCE) Core Infrastructure team—also known as the AI/ML Forward Deployed Infrastructure Engineering team—provides white‑glove engineering and operational support to OCI's most strategic AI/ML infrastructure customers.

As a trusted partner to our customers, we play a critical role in designing, deploying, operating, and optimizing the infrastructure that powers some of the largest and most demanding GPU and AI/ML environments in the world. Our team works closely with customers and internal engineering organizations to ensure exceptional reliability, performance, and scalability for mission‑critical AI workloads.

We are seeking a Manager to lead and develop a class of early‑career Infrastructure Engineers. This leader will create the conditions for new professionals to build sound technical judgment, strong customer habits, and the practical skills required to support strategic AI/ML infrastructure customers on OCI. The manager will combine people leadership with hands‑on operational guidance, helping the class grow into effective engineers who can independently contribute to customer execution, automation, troubleshooting, and infrastructure performance.

#LI-ES2

Internal Responsibilities

Lead, mentor, and develop a class of early‑career Engineers, setting clear expectations for technical growth, customer engagement, quality, and ownership.

Design and run an onboarding and development program that gives early‑career professionals progressive experience with GPU cluster deployment, Slurm and/or Oracle Kubernetes Engine (OKE), automation, monitoring, multicloud, database and operational readiness.

Provide regular coaching, feedback, career guidance, and practical learning opportunities. Pair team members with appropriate mentors and create stretch assignments that build confidence and sound engineering judgment.

Guide the team in building collaborative relationships with OCI Services, customers, sales, and core engineering teams. Model clear technical communication and help team members become effective, trusted customer partners.

Establish a safe, accountable environment where early‑career professionals learn to take ownership of problems, investigate issues thoroughly, document risks, and escape with context when needed.

Coach team members on infrastructure performance fundamentals, including parameter tuning, resource utilization, caching, and data pre‑processing techniques for AI/ML workloads.

Oversee the team's approach to troubleshooting performance, scalability, and reliability issues. Review complex cases, remove barriers, and ensure lessons learned are incorporated into team practices.

Ensure the team documents infrastructure designs, configurations, runbooks, and procedures so knowledge is shared, work is maintainable, and new professionals can learn from repeatable practices.

Help team members develop into credible customer advocates who can explain advanced GPU solution best practices and support customers and partners as they migrate workloads to the cloud.

Partner with senior technical leaders to calibrate the class's development progress, staffing needs, and readiness for increasingly independent customer‑facing and engineering responsibilities.

Qualifications:

7+ years of relevant infrastructure, cloud, platform, or software engineering experience, including experience leading or mentoring engineers.

Strong communication and collaboration skills, with the ability to coach early‑career professionals and translate technical concepts for customers, partners, and non‑technical stakeholders.

Demonstrated people‑management, mentoring, or talent‑development skills, with a thoughtful and consistent approach to feedback and performance management.

Experience building, operating, or supporting distributed and cloud infrastructure solutions in a customer‑focused environment.

Demonstrated ability to:

  • Develop short-, medium-, and long‑term development plans that connect individual growth to team and business objectives.
  • Work across functional areas with senior leaders and technical partners to secure opportunities, support, and clear expectations for an early‑career class.
  • Influence through coaching, sound judgment, and clear communication in situations involving competing priorities or sensitive feedback.

BS or MS degree, or equivalent experience, relevant to the functional area.

Working knowledge of infrastructure automation and orchestration tools such as Ansible, Terraform, Python, Docker, Kubernetes, and related tooling.

Solid understanding of networking concepts, security principles, and operational best practices.

Excellent problem‑solving skills, with the ability to guide others through structured troubleshooting and drive timely resolution in a fast‑paced environment.

Strong Linux skills with hands‑on experience in Oracle Linux/RHEL/CentOS, Ubuntu, or Debian, including system administration, package management, shell scripting, and performance optimization.

Proficiency in at least one programming language, such as Python, Rust, Go, Java, or Scala, and the ability to use code reviews and pairing to develop others.

Experience designing, implementing, managing, or supporting infrastructure for AI/ML, HPC, GPU, or similarly complex cloud workloads.

Core Responsibilities

Planning & Execution:

Plan and guide the class's work across onboarding, customer support, automation, and development goals. Set clear priorities, assign work that matches each person's readiness, monitor delivery, and adjust learning plans as business needs change.

Collaboration & Partnership:

Build productive cross‑functional partnerships that give early‑career professionals access to the expertise and context they need. Model inclusive collaboration, actively seek diverse perspectives, and help the team communicate expectations clearly with stakeholders and customers.

Problem Solving:

Coach the team to analyze operational and technical issues using sound problem‑solving practices. Review complex or ambiguous cases, help identify root causes, and turn lessons learned into guidance that prevents recurrence.

Continuous Learning:

Create structured opportunities for team members to build expertise through training, paired work, technical reviews, and hands‑on customer scenarios. Track individual skill growth, identify gaps early, and reinforce a culture of knowledge sharing and continuous learning.

Continuous Improvement:

Enable the team to improve processes, protocols, runbooks, and workflows. Gather feedback from early‑career professionals and partners, prioritize practical improvements, and help the team learn how to turn ideas into durable operational practices.

Performance and Development:

Drive consistent performance and development through timely, specific feedback, regular coaching, and clear growth expectations. Support hiring, onboarding, talent reviews, and promotion readiness in partnership with leadership and HR, while aligning individual development plans to organizational needs.

External Responsibilities

Lead, mentor, and develop a class of early‑career Engineers, setting clear expectations for technical growth, customer engagement, quality, and ownership.

Design and run an onboarding and development program that gives early‑career professionals progressive experience with GPU cluster deployment, Slurm and/or Oracle Kubernetes Engine (OKE), automation, monitoring, multicloud, database and operational readiness.

Provide regular coaching, feedback, career guidance, and practical learning opportunities. Pair team members with appropriate mentors and create stretch assignments that build confidence and sound engineering judgment.

Guide the team in building collaborative relationships with OCI Services, customers, sales, and core engineering teams. Model clear technical communication and help team members become effective, trusted customer partners.

Establish a safe, accountable environment where early‑career professionals learn to take ownership of problems, investigate issues thoroughly, document risks, and escape with context when needed.

Coach team members on infrastructure performance fundamentals, including parameter tuning, resource utilization, caching, and data pre‑processing techniques for AI/ML workloads.

Oversee the team's approach to troubleshooting performance, scalability, and reliability issues. Review complex cases, remove barriers, and ensure lessons learned are incorporated into team practices.

Ensure the team documents infrastructure designs, configurations, runbooks, and procedures so knowledge is shared, work is maintainable, and new professionals can learn from repeatable practices.

Help team members develop into credible customer advocates who can explain advanced GPU solution best practices and support customers and partners as they migrate workloads to the cloud.

Partner with senior technical leaders to calibrate the class's development progress, staffing needs, and readiness for increasingly independent customer‑facing and engineering responsibilities.

Qualifications:

7+ years of relevant infrastructure, cloud, platform, or software engineering experience, including experience leading or mentoring engineers.

Strong communication and collaboration skills, with the ability to coach early‑career professionals and translate technical concepts for customers, partners, and non‑technical stakeholders.

Demonstrated people‑management, mentoring, or talent‑development skills, with a thoughtful and consistent approach to feedback and performance management.

Experience building, operating, or supporting distributed and cloud infrastructure solutions in a customer‑focused environment.

Demonstrated ability to:

  • Develop short-, medium-, and long‑term development plans that connect individual growth to team and business objectives.
  • Work across functional areas with senior leaders and technical partners to secure opportunities, support, and clear expectations for an early‑career class.
  • Influence through coaching, sound judgment, and clear communication in situations involving competing priorities or sensitive feedback.

BS or MS degree, or equivalent experience, relevant to the functional area.

Working knowledge of infrastructure automation and orchestration tools such as Ansible, Terraform, Python, Docker, Kubernetes, and related tooling.

Solid understanding of networking concepts, security principles, and operational best practices.

Excellent problem‑solving skills, with the ability to guide others through structured troubleshooting and drive timely resolution in a fast‑paced environment.

Strong Linux skills with hands‑on experience in Oracle Linux/RHEL/CentOS, Ubuntu, or Debian, including system administration, package management, shell scripting, and performance optimization.

Proficiency in at least one programming language, such as Python, Rust, Go, Java, or Scala, and the ability to use code reviews and pairing to develop others.

Experience designing, implementing, managing, or supporting infrastructure for AI/ML, HPC, GPU, or similarly complex cloud workloads.

Core Responsibilities

Planning & Execution:

Plan and guide the class's work across onboarding, customer support, automation, and development goals. Set clear priorities, assign work that matches each person's readiness, monitor delivery, and adjust learning plans as business needs change.

Collaboration & Partnership:

Build productive cross‑functional partnerships that give early‑career professionals access to the expertise and context they need. Model inclusive collaboration, actively seek diverse perspectives, and help the team communicate expectations clearly with stakeholders and customers.

Problem Solving:

Coach the team to analyze operational and technical issues using sound problem‑solving practices. Review complex or ambiguous cases, help identify root causes, and turn lessons learned into guidance that prevents recurrence.

Continuous Learning:

Create structured opportunities for team members to build expertise through training, paired work, technical reviews, and hands‑on customer scenarios. Track individual skill growth, identify gaps early, and reinforce a culture of knowledge sharing and continuous learning.

Continuous Improvement:

Enable the team to improve processes, protocols, runbooks, and workflows. Gather feedback from early‑career professionals and partners, prioritize practical improvements, and help the team learn how to turn ideas into durable operational practices.

Performance and Development:

Drive consistent performance and development through timely, specific feedback, regular coaching, and clear growth expectations. Support hiring, onboarding, talent reviews, and promotion readiness in partnership with leadership and HR, while aligning individual development plans to organizational needs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Software Engineer, Core Infrastructure
Principal Software Engineer, Core Infrastructure

Oracle Corporation • Reston (VA)

On-site
USD 130,000 - 190,000
Senior Core Infrastructure Engineer, AI Infrastructure
Senior Core Infrastructure Engineer, AI Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 150,000 - 190,000
Lead Principal Software Engineer, Core Infrastructure
Lead Principal Software Engineer, Core Infrastructure

Oracle Corporation • Seattle (WA)

On-site
USD 180,000 - 240,000
Senior Logistics & Transportation Manager (ONSITE: Abilene, TX)
Senior Logistics & Transportation Manager (ONSITE: Abilene, TX)

Oracle Corporation • Abilene (TX)

On-site
USD 110,000 - 140,000
Senior Manager, Core Infrastructure Engineering
Senior Manager, Core Infrastructure Engineering

Oracle • Seattle (WA)

On-site
USD 146,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1
Lead Principal Technical Program Manager
Lead Principal Technical Program Manager

Oracle Corporation • Nashville (TN)

On-site
USD 160,000 - 210,000
Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Oracle Corporation • Nashville (TN)

On-site
USD 150,000 - 190,000
Principal Software Engineer, Core Infrastructure
Principal Software Engineer, Core Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 150,000 - 210,000
Relocation assistance
Principal Core Infrastructure Engineer
Principal Core Infrastructure Engineer

Oracle Corporation • Nashville (TN)

On-site
USD 120,000 - 160,000
Senior Core Infrastructure Engineer (Nashville, TN)
Senior Core Infrastructure Engineer (Nashville, TN)

Oracle Corporation • Nashville (TN)

On-site
USD 120,000 - 180,000