Principal Systems Engineer

Oracle Corporation

Nashville (TN)

Remote

USD 140,000 - 200,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Oracle Corporation seeks a Principal Systems Engineer to lead AI Infra Operations for OCI GPU infrastructure. You will drive automation, tooling, monitoring, and reliability across regions, collaborating with software, hardware, and operations teams to keep GPU fleets highly available.

Responsibilities include incident response, root-cause analysis, and optimizing AI2 Ops processes. You will mentor junior engineers and contribute to runbooks, ensuring safe rollout and continuous improvement of

Qualifications

  • 8+ years in software operations or infrastructure automation.
  • Expert Linux admin experience (Ubuntu/Oracle Linux) in large-scale environments.
  • Strong understanding of distributed systems and service communication patterns.
  • Experience with provisioning, validation, repair workflows, and fleet recovery.
  • Strong problem-solving and troubleshooting abilities.
  • Excellent communication and teamwork.
  • Experience with observability tooling (metrics, logging, dashboards, alerting).
  • Experience with AI agents and tooling.
  • Experience leading on-call operations and incident response.
  • Bachelor’s degree in CS/Engineering or related field.

Responsibilities

  • Build and maintain automation and tooling for OCI GPU infrastructure across regions.
  • Lead collaboration with software, hardware, and operations teams to maintain a highly available GPU fleet.
  • Develop monitoring, alerting, and diagnostics using Grafana for GPU fleet health and utilization.
  • Serve as senior escalation point for complex GPU host/repair issues.
  • Participate in incident response and root-cause analyses to remove blockers.
  • Improve AI2 Ops processes, GPU fleet automation, and region build readiness.
  • Participate in on-call rotations and support critical infrastructure issues.
  • Document procedures, automation workflows, and runbooks.
  • Build and improve AI agents with safe rollout and monitoring.
  • Mentor junior engineers in operational best practices.

Skills

Python
Bash
Linux administration
Distributed systems
On-call operations
Incident response
Observability
Team collaboration
Bachelor's degree
GPU infrastructure
Grafana

Education

Bachelor's degree in CS/Engineering or related

Tools

Grafana

Job description

We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Systems Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate closely with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.

Internal Responsibilities
  • Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.

  • Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.

  • Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.

  • Serve as the senior escalation point for complex GPU host and repair issues

  • Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.

  • Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.

  • Participate in on-call rotations and provide support for critical infrastructure issues.

  • Document operational procedures, automation workflows, troubleshooting guides, and runbooks.

  • Build and improve AI agents, ensuring safe rollout, execution and monitoring.

  • Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.

Required Qualifications:

  • 8+ years of software operations or infrastructure automation experience with strong proficiency in Python and Bash.

  • Expert Linux administration experience, particularly Ubuntu and Oracle Linux, in large-scale production environments.

  • Strong understanding of distributed systems, including peer-to-peer, node-to-node, and service-to-service communication patterns.

  • Strong data-center and host-lifecycle experience, including provisioning, validation, repair workflows, hardware replacement, and fleet recovery.

  • Strong problem-solving and troubleshooting skills.

  • Excellent communication and teamwork skills.

  • Experience with observability tooling, including metrics, logging, dashboards, and alerting.

  • Experience with AI agents and tooling

  • Experience leading on-call operations and incident response.

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Preferred Skills:

  • Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.

  • Experience operating or automating GPU, compute, or other large-scale cloud infrastructure.

External Responsibilities
  • Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.

  • Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.

  • Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.

  • Serve as the senior escalation point for complex GPU host and repair issues

  • Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.

  • Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.

  • Participate in on-call rotations and provide support for critical infrastructure issues.

  • Document operational procedures, automation workflows, troubleshooting guides, and runbooks.

  • Build and improve AI agents, ensuring safe rollout, execution and monitoring.

  • Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.

Required Qualifications:

  • 8+ years of software operations or infrastructure automation experience with strong proficiency in Python and Bash.

  • Expert Linux administration experience, particularly Ubuntu and Oracle Linux, in large-scale production environments.

  • Strong understanding of distributed systems, including peer-to-peer, node-to-node, and service-to-service communication patterns.

  • Strong data-center and host-lifecycle experience, including provisioning, validation, repair workflows, hardware replacement, and fleet recovery.

  • Strong problem-solving and troubleshooting skills.

  • Excellent communication and teamwork skills.

  • Experience with observability tooling, including metrics, logging, dashboards, and alerting.

  • Experience with AI agents and tooling

  • Experience leading on-call operations and incident response.

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Preferred Skills:

  • Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.

  • Experience operating or automating GPU, compute, or other large-scale cloud infrastructure.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Systems Engineer
Principal Systems Engineer

Oracle • Nashville (TN)

On-site
USD 180,000 - 240,000
Senior Systems Engineer
Senior Systems Engineer

Oracle Corporation • Nashville (TN)

On-site
USD 120,000 - 180,000
Principal GPU Infra Automation Architect
Principal GPU Infra Automation Architect

Oracle • Nashville (TN)

On-site
USD 180,000 - 240,000
Senior GPU Infra Automation Architect
Senior GPU Infra Automation Architect

Oracle Corporation • Nashville (TN)

Remote
USD 140,000 - 200,000
Principal Core Infrastructure Engineer - AI Infrastructure
Principal Core Infrastructure Engineer - AI Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 180,000 - 275,000
Principal Software Developer, AI Infrastructure
Principal Software Developer, AI Infrastructure

Ll Oefentherapie • Austin (TX)

On-site
USD 120,000 - 160,000
Software Developer 5 , AI Infrastructure
Software Developer 5 , AI Infrastructure

Ll Oefentherapie • Seattle (WA)

On-site
USD 120,000 - 160,000
Dynamic and flexible workplace
Senior Core Infrastructure Engineer, AI Infrastructure
Senior Core Infrastructure Engineer, AI Infrastructure

Ll Oefentherapie • Nashville (TN)

On-site
USD 130,000 - 190,000
GPU/CPU Systems Engineer
GPU/CPU Systems Engineer

Oracle • Seattle (WA)

On-site
USD 135,200 - 306,400
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Flexible vacation and paid time off
+1
GPU/CPU Systems Engineer
GPU/CPU Systems Engineer

Oracle • San Francisco (CA)

On-site
USD 135,200 - 306,400
Medical, dental, and vision insurance
401(k) Savings Plan with company match
Flexible Vacation
+1