Senior Systems Engineer

Oracle Corporation

Nashville (TN)

On-site

USD 120,000 - 180,000

Full time

2 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Oracle Corporation is seeking a Senior Systems Engineer to join the AI Infra Operations team, supporting GPU infrastructure in OCI. You will help develop and maintain automation and tooling for GPU fleets across regions, ensuring high availability and robust observability.

You will troubleshoot production issues on Linux hosts and GPU systems, own incident response, and drive improvements in runbooks, monitoring, and deployment pipelines.

Qualifications

  • 5+ years of software operations, systems administration or infrastructure automation experience.
  • Proficiency in Python and Bash.
  • Strong Linux administration experience.
  • Good understanding of host-level networking and connectivity troubleshooting.
  • Understanding of distributed systems communication patterns.
  • Experience with data center or large-scale infrastructure operations.
  • Troubleshooting across GPU hosts, OS, hardware, networking and services.
  • Automation using shell scripting and operational tooling.
  • Strong ownership, incident handling and reliability mindset.
  • Hands-on experience with monitoring, deployment, configuration and support tooling.

Responsibilities

  • Troubleshoot and resolve production issues involving Linux hosts, GPU systems, and host-network interactions.
  • Support GPU host operations across provisioning, validation, incident response, and readiness activities.
  • Own incident response and resolution for assigned services and escalate as needed.
  • Identify recurring failure patterns and contribute improvements to runbooks, tooling, monitoring, and procedures.
  • Collaborate with partner teams to resolve hardware, OS, networking, and service dependency issues.
  • Contribute automation and scripting to reduce manual effort and improve reliability.
  • Participate in on-call rotations and support critical infrastructure issues.
  • Provide data-driven analysis to address problems rather than symptoms.

Skills

Python
Bash
Linux
Networking
Distributed systems
GPU infrastructure
Automation scripting
Incident management
Monitoring

Job description

We are seeking a Senior Systems Engineer to join the AI Infra Operations team supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will assist in the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously working on and improving GPU operations.

Internal Responsibilities
  • Independently troubleshoot and resolve production issues involving Linux hosts, GPU systems, and host-network interactions.
  • Support GPU host operations across provisioning, validation, incident response, and operational readiness activities.
  • Own incident response and resolution for assigned services while escalating appropriately when problems cross team or domain boundaries.
  • Identify recurring failure patterns and contribute improvements to runbooks, tooling, monitoring, and operational procedures.
  • Work with partner teams to resolve issues that span hardware, operating systems, networking, and service dependencies.
  • Contribute automation and scripting improvements that reduce manual effort and improve operator effectiveness.
  • Participate in on-call rotations and provide support for critical infrastructure issues.
  • Provide data-driven analysis to address problems instead of symptoms.

Required Qualifications:

  • 5+ years of software operations, systems administration or infrastructure automation experience with proficiency in Python and Bash.
  • Strong Linux administration experience.
  • Good understanding of host-level networking fundamentals, including TCP/IP, routing, DNS, interface behavior, and connectivity troubleshooting.
  • Understanding of peer-to-peer, node-to-node, and service-to-service communication patterns in distributed systems.
  • Experience supporting data center operations or large-scale infrastructure environments.
  • Demonstrated troubleshooting experience across GPU hosts, operating systems, hardware, networking, drivers, and dependent services in production.
  • Automation experience using shell scripting and operational tooling to reduce manual effort and improve reliability.
  • Strong operational ownership, incident handling, and problem-solving skills in production environments.
  • Hands-on experience with monitoring, deployment, configuration, and support tooling used in day-to-day operations.

Preferred Skills:

  • Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.
External Responsibilities
  • Independently troubleshoot and resolve production issues involving Linux hosts, GPU systems, and host-network interactions.
  • Support GPU host operations across provisioning, validation, incident response, and operational readiness activities.
  • Own incident response and resolution for assigned services while escalating appropriately when problems cross team or domain boundaries.
  • Identify recurring failure patterns and contribute improvements to runbooks, tooling, monitoring, and operational procedures.
  • Work with partner teams to resolve issues that span hardware, operating systems, networking, and service dependencies.
  • Contribute automation and scripting improvements that reduce manual effort and improve operator effectiveness.
  • Participate in on-call rotations and provide support for critical infrastructure issues.
  • Provide data-driven analysis to address problems instead of symptoms.

Required Qualifications:

  • 5+ years of software operations, systems administration or infrastructure automation experience with proficiency in Python and Bash.
  • Strong Linux administration experience.
  • Good understanding of host-level networking fundamentals, including TCP/IP, routing, DNS, interface behavior, and connectivity troubleshooting.
  • Understanding of peer-to-peer, node-to-node, and service-to-service communication patterns in distributed systems.
  • Experience supporting data center operations or large-scale infrastructure environments.
  • Demonstrated troubleshooting experience across GPU hosts, operating systems, hardware, networking, drivers, and dependent services in production.
  • Automation experience using shell scripting and operational tooling to reduce manual effort and improve reliability.
  • Strong operational ownership, incident handling, and problem-solving skills in production environments.
  • Hands-on experience with monitoring, deployment, configuration, and support tooling used in day-to-day operations.

Preferred Skills:

  • Hands-on experience with GPU infrastructure, including NVIDIA and AMD based systems, and familiarity with GPU shapes or instance families.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Principal Systems Engineer
Principal Systems Engineer

Oracle • Nashville (TN)

On-site
USD 180,000 - 240,000
Senior GPU Infra & Automation Engineer
Senior GPU Infra & Automation Engineer

Oracle Corporation • Nashville (TN)

On-site
USD 120,000 - 180,000
Principal Core Infrastructure Engineer - AI Infrastructure
Principal Core Infrastructure Engineer - AI Infrastructure

Oracle Corporation • Nashville (TN)

On-site
USD 180,000 - 275,000
Senior Core Infrastructure Engineer, AI Infrastructure
Senior Core Infrastructure Engineer, AI Infrastructure

Ll Oefentherapie • Nashville (TN)

On-site
USD 130,000 - 190,000
Principal GPU Infra Automation Architect
Principal GPU Infra Automation Architect

Oracle • Nashville (TN)

On-site
USD 180,000 - 240,000
Principal Software Developer, AI Infrastructure
Principal Software Developer, AI Infrastructure

Ll Oefentherapie • Austin (TX)

On-site
USD 120,000 - 160,000
Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Oracle Corporation • Nashville (TN)

On-site
USD 150,000 - 190,000
Artificial Intelligence Engineer
Artificial Intelligence Engineer

Calance • Costa Mesa (CA)

Hybrid
USD 180,000 - 240,000
Software Developer 5 , AI Infrastructure
Software Developer 5 , AI Infrastructure

Ll Oefentherapie • Seattle (WA)

On-site
USD 120,000 - 160,000
Dynamic and flexible workplace
GPU Systems Infrastructure Engineer
GPU Systems Infrastructure Engineer

Blue Signal Search • Fremont (CA)

On-site
USD 120,000 - 170,000