Senior GPU Infrastructure Automation Engineer

Oracle

Región Centro

Presencial

MXN 600.000 - 900.000

Jornada completa

Hace 6 días
Sé de los primeros/as/es en solicitar esta vacante

Recibe más respuestas de empleadores

Envía un currículum específico para el puesto de trabajo en cuestión de minutos.

Descripción de la vacante

Oracle in Mexico is seeking a highly capable Developer to join the AI2 Ops team, focusing on GPU infrastructure automation across multiple OCI regions and building robust tooling for provisioning, monitoring, and reliability.

You will collaborate with engineering, product, and operations to improve GPU operations, runbooks, and incident response while ensuring high availability and efficient capacity utilization.

Formación

  • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
  • 3+ years of software development or infrastructure automation with Python and Bash.
  • Strong Linux experience including troubleshooting and system administration.
  • Hands-on experience with infrastructure-as-code and Terraform and RESTful APIs.
  • Strong problem-solving and troubleshooting skills.
  • Excellent communication and teamwork.

Responsabilidades

  • Design, build, and maintain software, automation, and operational tooling for OCI GPU infrastructure across multiple geographic regions.
  • Automate GPU infrastructure provisioning, configuration, validation, and deployment using Python, Bash, Terraform, and related tooling.
  • Collaborate with software engineers, hardware teams, and operations partners to build scalable, reliable, and highly available GPU platform services.
  • Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using Grafana.
  • Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.
  • Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.
  • Participate in on-call rotations and provide support for critical infrastructure issues.
  • Document operational procedures, automation workflows, troubleshooting guides, and runbooks.

Conocimientos

Python
Bash
Linux
REST APIs
Problem solving
Communication

Educación

Bachelor's degree in Computer Science

Herramientas

Terraform
Docker
Kubernetes
Grafana
CI/CD tooling

Descripción del empleo

Oracle in Mexico is seeking a highly capable Developer to join the AI2 Ops team, focusing on GPU infrastructure automation across multiple OCI regions and building robust tooling for provisioning, monitoring, and reliability.

You will collaborate with engineering, product, and operations to improve GPU operations, runbooks, and incident response while ensuring high availability and efficient capacity utilization.

Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior Core Infrastructure Engineer
Senior Core Infrastructure Engineer

Oracle • Región Centro

Presencial
MXN 600.000 - 900.000
Senior GPU Server Platform Engineer - Onsite Mexico
Senior GPU Server Platform Engineer - Onsite Mexico

Oracle • Región Centro

Presencial
MXN 1.746.000 - 2.271.000
Senior Product Engineer – GPU Server Platforms (Ciudad Juarez, México)
Senior Product Engineer – GPU Server Platforms (Ciudad Juarez, México)

Oracle • Región Centro

Presencial
MXN 1.746.000 - 2.271.000
Senior Product Engineer – GPU Server Platforms (Ciudad Juarez, México)
Senior Product Engineer – GPU Server Platforms (Ciudad Juarez, México)

Ll Oefentherapie • Región Centro

Presencial
MXN 600.000 - 800.000
Senior Site Reliability Engineer — Scale & Automate
Senior Site Reliability Engineer — Scale & Automate

Oracle • Región Centro

Presencial
MXN 600.000 - 900.000
Data Center Operations Engineer III — GPU Infra
Data Center Operations Engineer III — GPU Infra

Lambda • Santiago de Querétaro

Presencial
MXN 531.000 - 708.000
Health, dental, and vision coverage
Wellness stipend
Commuter stipend
+1
Senior Data Center Support Services Technician - Location Queretaro
Senior Data Center Support Services Technician - Location Queretaro

Oracle • Santiago de Querétaro

Presencial
MXN 350.000 - 520.000
Senior DevOps Engineer - Data Platform & Fleet Analytics
Senior DevOps Engineer - Data Platform & Fleet Analytics

Intel • Región Centro

Presencial
MXN 700.000 - 1.100.000
Hybrid Release & AI Engineer: Automate Deployments
Hybrid Release & AI Engineer: Automate Deployments

3M HEALTHCARE • Ciudad de México

Híbrido
MXN 420.000 - 700.000
Health & Wellness
Flexible time off
Continuous learning
+3
Oracle DBA Engineer (Hybrid Guadalajara) RMAN & ODA Expert
Oracle DBA Engineer (Hybrid Guadalajara) RMAN & ODA Expert

HCLTech • Región Centro

Híbrido
MXN 600.000 - 1.000.000
Major medical insurance
Minor medical insurance
Saving fund
+5