Senior Cloud Platform Engineer - GPU Inference Orchestration

Neura Market

Hinoba-an

On-site

PHP 2,322,000 - 3,980,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Together AI is seeking a software engineer to build a Kubernetes-native control plane that provisions and runs our GPU inference fleet. You will design a manifest-driven API where the inference team declares needs, while controllers handle reconciliation, provider/runtime selection, and lifecycle management under the hood.

You will also build systems for efficiency, including defragmentation and rebalancing, and scheduling/bin-packing improvements to maximize GPU utilization without sacrificing

Qualifications

  • Strong software engineering background with modern languages (Go/Python/Rust) and production-grade software practices.
  • Experience implementing durable workflows with tools like Temporal or Cadence and building control planes.
  • Hands-on with Kubernetes controllers/operators or similar reconciliation loops to model state over time.
  • Experience designing event-driven systems using message queues or streams (Kafka, NATS, SQS).
  • A product mindset; building internal platforms or APIs used by other teams.

Responsibilities

  • Build the provisioning state machine for full lifecycle management of hardware and GPU drivers.
  • Create a self-service control plane API for one-click provisioning, scaling, and teardown of clusters.
  • Automate self-healing: detect, drain, repair, and reintroduce healthy capacity automatically.
  • Ensure reliability: idempotency, retries, rollback, drift detection in the deployment pipeline.
  • Collaborate with the ML/inference platform team to encode topology and scheduling requirements.
  • Treat infrastructure code as software: strong typing, tests, reviews, CI/CD.

Skills

Go
Python
Rust
Temporal
Kubernetes
Event-driven

Job description

Together AI is seeking a software engineer to build a Kubernetes-native control plane that provisions and runs our GPU inference fleet. You will design a manifest-driven API where the inference team declares needs, while controllers handle reconciliation, provider/runtime selection, and lifecycle management under the hood.

You will also build systems for efficiency, including defragmentation and rebalancing, and scheduling/bin-packing improvements to maximize GPU utilization without sacrificing

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Lead Solution Architect
Lead Solution Architect

Yotta • Hinoba-an

On-site
PHP 9,404,000 - 13,166,000
Senior GPU Systems Architect: Multi-GPU AI Data Center
Senior GPU Systems Architect: Multi-GPU AI Data Center

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 7,524,000 - 11,912,000
Hybrid work model
Platform AI Engineer
Platform AI Engineer

Hewlett Packard Enterprise Development LP • Hinoba-an

On-site
PHP 1,500,000 - 2,700,000
Staff Platform Engineer: Scalable AI Infra & Kubernetes
Staff Platform Engineer: Scalable AI Infra & Kubernetes

Ansys • España

On-site
PHP 7,528,000 - 10,038,000
Senior Software Engineer — Infra Agent Systems Remote India Together AI India
Senior Software Engineer — Infra Agent Systems Remote India Together AI India

Neura Market • Hinoba-an

Remote
INR 3,000,000 - 5,400,000
GPU System Architect for Scalable AI Data Centers
GPU System Architect for Scalable AI Data Centers

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 1,200,000 - 1,800,000
GPU Architect
GPU Architect

NVIDIA Corporation • Hinoba-an

Hybrid
PHP 1,200,000 - 1,800,000
Senior GPU System/Fabrics Architect
Senior GPU System/Fabrics Architect

NVIDIA Corporation • Hinoba-an

On-site
PHP 7,524,000 - 11,912,000
Hybrid work model
Platform AI Engineer: Kubernetes & Cloud Platform Lead
Platform AI Engineer: Kubernetes & Cloud Platform Lead

Hewlett Packard Enterprise Development LP • Hinoba-an

On-site
PHP 1,500,000 - 2,700,000
Senior Cloud Reliability Engineer
Senior Cloud Reliability Engineer

Infios BR Ltda. • Mexico

Hybrid
PHP 7,509,000 - 11,890,000