Platform ML Infra Engineer: GPU, Kubernetes & MLOps

Oracle

United States

On-site

USD 92,500 - 209,500

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health insurance
401(k) match
Paid time off
Equity
Disability insurance

Job summary

Oracle is seeking a Platform Software Engineer to build the execution layer behind OCI’s AI and GPU growth. You’ll work across Oracle data platforms, GPU scheduling, orchestration pipelines, and MLOps tooling — turning customer demand into production‑grade systems that scale across our cluster fleet.

This is a builder role on a small, high‑leverage team. The role emphasizes hands‑on distributed systems work with Kubernetes, Python, and modern ML infra tooling.

Qualifications

  • 3–6 years of software engineering experience, ideally in distributed systems, data platforms, or ML infrastructure.
  • Strong Python; working knowledge of Go or Rust is a plus.
  • Hands-on experience with Kubernetes, container runtimes, and at least one workflow engine (Argo, Airflow, Prefect, Temporal).
  • Familiarity with GPU workloads, CUDA toolchains, or inference serving frameworks (vLLM, TensorRT-LLM, Triton).
  • Experience with LLM agent frameworks (LangGraph, CrewAI, or similar) is a strong plus.

Responsibilities

  • Integrate with Oracle data platforms to move training and inference data across pipelines.
  • Build and extend GPU schedulers and capacity-aware placement logic for diverse fleet.
  • Develop orchestration pipelines for training, fine-tuning, and inference workloads using Kubernetes, Argo, and Slurm.
  • Ship MLOps tooling — observability, automated triage, cost and SLO agents — to reduce operator load.
  • Collaborate with PMs, SAs, and customer-facing teams to convert field requirements into reusable components.

Skills

Python
Go
Rust
Kubernetes
Argo
Airflow
Prefect
Temporal
CUDA
LangGraph
CrewAI

Tools

Docker
Slurm
TensorRT-LLM

Job description

Oracle is seeking a Platform Software Engineer to build the execution layer behind OCI’s AI and GPU growth. You’ll work across Oracle data platforms, GPU scheduling, orchestration pipelines, and MLOps tooling — turning customer demand into production‑grade systems that scale across our cluster fleet.

This is a builder role on a small, high‑leverage team. The role emphasizes hands‑on distributed systems work with Kubernetes, Python, and modern ML infra tooling.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Software Engineer - AI/GPU Orchestration & MLOps
Platform Software Engineer - AI/GPU Orchestration & MLOps

Ll Oefentherapie • United States

On-site
USD 140,000 - 210,000
Platform Software Engineer
Platform Software Engineer

Ll Oefentherapie • United States

On-site
USD 140,000 - 210,000
Lead ML Infrastructure Architect – GenAI & GPU Clusters
Lead ML Infrastructure Architect – GenAI & GPU Clusters

Ll Oefentherapie • United States

On-site
USD 130,000 - 170,000
Principal ML Systems Engineer - GenAI Infra & GPU
Principal ML Systems Engineer - GenAI Infra & GPU

Oracle • United States

On-site
USD 114,600 - 234,600
Medical, dental, and vision insurance
401(k) savings plan
Paid time off and sick leave
+1
Strategic AI/ML GPU Infrastructure Director
Strategic AI/ML GPU Infrastructure Director

Oracle • Atlanta (GA)

Hybrid
USD 122,000 - 306,000
Medical, dental, and vision insurance
401(k) with company match
Paid time off
+1
Head of AI/ML Core Infrastructure & GPU Cluster Ops
Head of AI/ML Core Infrastructure & GPU Cluster Ops

Oracle • Salt Lake City (UT)

On-site
USD 122,000 - 306,000
Medical, dental, and vision insurance
Disability insurance
Life insurance
+7
Head of AI/ML Core Infra & GPU Cluster Ops
Head of AI/ML Core Infra & GPU Cluster Ops

Oracle • Denver (CO)

On-site
USD 122,000 - 306,000
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops

Oracle • United States

On-site
USD 121,000 - 307,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
Senior Cloud GPU Infrastructure Engineer
Senior Cloud GPU Infrastructure Engineer

Oracle • United States

On-site
USD 183,000 - 307,000
Medical Insurance
Disability Insurance
Life Insurance
+5