Senior GPU Infra & Automation Engineer

Oracle Corporation

Nashville (TN)

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Oracle Corporation is seeking a Senior Systems Engineer to join the AI Infra Operations team, supporting GPU infrastructure in OCI. You will help develop and maintain automation and tooling for GPU fleets across regions, ensuring high availability and robust observability.

You will troubleshoot production issues on Linux hosts and GPU systems, own incident response, and drive improvements in runbooks, monitoring, and deployment pipelines.

Qualifications

  • 5+ years of software operations, systems administration or infrastructure automation experience.
  • Proficiency in Python and Bash.
  • Strong Linux administration experience.
  • Good understanding of host-level networking and connectivity troubleshooting.
  • Understanding of distributed systems communication patterns.
  • Experience with data center or large-scale infrastructure operations.
  • Troubleshooting across GPU hosts, OS, hardware, networking and services.
  • Automation using shell scripting and operational tooling.
  • Strong ownership, incident handling and reliability mindset.
  • Hands-on experience with monitoring, deployment, configuration and support tooling.

Responsibilities

  • Troubleshoot and resolve production issues involving Linux hosts, GPU systems, and host-network interactions.
  • Support GPU host operations across provisioning, validation, incident response, and readiness activities.
  • Own incident response and resolution for assigned services and escalate as needed.
  • Identify recurring failure patterns and contribute improvements to runbooks, tooling, monitoring, and procedures.
  • Collaborate with partner teams to resolve hardware, OS, networking, and service dependency issues.
  • Contribute automation and scripting to reduce manual effort and improve reliability.
  • Participate in on-call rotations and support critical infrastructure issues.
  • Provide data-driven analysis to address problems rather than symptoms.

Skills

Python
Bash
Linux
Networking
Distributed systems
GPU infrastructure
Automation scripting
Incident management
Monitoring

Job description

Oracle Corporation is seeking a Senior Systems Engineer to join the AI Infra Operations team, supporting GPU infrastructure in OCI. You will help develop and maintain automation and tooling for GPU fleets across regions, ensuring high availability and robust observability.

You will troubleshoot production issues on Linux hosts and GPU systems, own incident response, and drive improvements in runbooks, monitoring, and deployment pipelines.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal GPU Infra Automation Architect
Principal GPU Infra Automation Architect

Oracle • Nashville (TN)

On-site
USD 180,000 - 240,000
Senior GPU Infra Engineer for Scalable AI Platforms
Senior GPU Infra Engineer for Scalable AI Platforms

Oracle • Austin (TX)

On-site
USD 89,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Senior AI Infra & GPU Compute Engineer
Senior AI Infra & GPU Compute Engineer

Oracle • United States

On-site
USD 89,000 - 210,000
Medical insurance
Dental insurance
Vision insurance
+2
Senior GPU Infra Engineer for Scalable AI Platform
Senior GPU Infra Engineer for Scalable AI Platform

Ll Oefentherapie • Nashville (TN)

On-site
USD 130,000 - 190,000
Senior Systems Engineer
Senior Systems Engineer

Oracle Corporation • Nashville (TN)

On-site
USD 120,000 - 180,000
Senior AI Infra & GPU Platform Engineer
Senior AI Infra & GPU Platform Engineer

Oracle • Nashville (TN)

On-site
USD 89,000 - 210,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
Senior Cloud GPU Infrastructure Engineer
Senior Cloud GPU Infrastructure Engineer

Oracle • United States

On-site
USD 183,000 - 307,000
Medical Insurance
Disability Insurance
Life Insurance
+5
Principal AI Infra Engineer - GPU & HPC Scale
Principal AI Infra Engineer - GPU & HPC Scale

Ll Oefentherapie • Nashville (TN)

On-site
USD 115,000 - 235,000
Principal Systems Engineer
Principal Systems Engineer

Oracle • Nashville (TN)

On-site
USD 180,000 - 240,000
Senior Director, AI Infrastructure Software
Senior Director, AI Infrastructure Software

Oracle • Santa Clara (CA)

On-site
USD 194,000 - 414,000
Medical/Dental/Vision
Disability Insurance
Life Insurance
+7