Senior AI Scheduler Engineer: Kubernetes & GPUs

Firmus

Singapore

On-site

SGD 180,000 - 240,000

Full time

5 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Firmus Technologies in Singapore is seeking a Senior AI/Platform Engineer to lead the Kubernetes-native scheduler development for large-scale AI factories. You will design, implement, and operate the Model-to-Grid workload management stack, enabling topology-aware placement, fault-tolerant execution, and high GPU utilization.

Collaborate with AI, inference, platform, and security teams to deliver scalable, multi-tenant scheduling with strong observability and CI/CD pipelines, contributing to a

Qualifications

  • 5+ years in DevOps, platform engineering, distributed systems, cloud infrastructure, or similar roles.
  • Deep Kubernetes expertise including control planes, scheduling, controllers, and multi-tenancy.
  • Experience building or extending workload schedulers or orchestration systems (Kubernetes Scheduler Framework, etc.).
  • Strong Go programming with Python for automation and tooling.
  • Experience with GPU workloads, scheduling, and topology-aware placement.
  • Proficiency with observability, performance analysis, and SLOs, using Prometheus, Grafana, OpenTelemetry, and related tools.
  • CI/CD, GitOps, IaC and safe production deployments (GitHub Actions, GitLab CI, Argo CD, Terraform, Helm).
  • Security-conscious multi-tenant platform engineering and incident response.

Responsibilities

  • Design, build, operate and improve the Kubernetes-native scheduling platform for AI workloads.
  • Develop CRDs, controllers, admission webhooks, plugins, APIs and automation for workload submission and management.
  • Define scheduling policies for training, fine-tuning, inference, benchmarking, and batch workloads.
  • Implement topology-aware placement considering GPU locality, NVLink/NVSwitch topology, and storage locality.
  • Develop AI-factory resource-aware scheduling aligned with capacity, health, power, thermal, and maintenance signals.
  • Integrate scheduler with Kubernetes, GPU plugins, network and storage services, and observability systems.
  • Build scheduler templates, model recipes, and self-service workflows for workload submission and troubleshooting.
  • Establish CI/CD, GitOps, and operator experiences for scheduler components.

Skills

Kubernetes
Go
Python
CI/CD
GitOps
Distributed systems
GPU workloads
Topology-aware scheduling
Observability
Multi-tenant

Tools

Slurm
Slinky
Kueue
KAI
Volcano
YuniKorn
Run:AI
GitHub Actions
GitLab CI
Argo CD
Flux
Terraform
Helm
Kustomize
OPA
Kyverno

Job description

Firmus Technologies in Singapore is seeking a Senior AI/Platform Engineer to lead the Kubernetes-native scheduler development for large-scale AI factories. You will design, implement, and operate the Model-to-Grid workload management stack, enabling topology-aware placement, fault-tolerant execution, and high GPU utilization.

Collaborate with AI, inference, platform, and security teams to deliver scalable, multi-tenant scheduling with strong observability and CI/CD pipelines, contributing to a

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer — Kubernetes Scheduler
Senior AI Platform Engineer — Kubernetes Scheduler

Firmus Technologies • Singapore

On-site
SGD 120,000 - 170,000
Senior AI Engineer (Kubernetes & Customised Scheduler)
Senior AI Engineer (Kubernetes & Customised Scheduler)

Firmus Technologies • Singapore

On-site
SGD 120,000 - 170,000
Senior AI Engineer (Kubernetes & Customised Scheduler)
Senior AI Engineer (Kubernetes & Customised Scheduler)

Firmus • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Infra Architect: GPU Clusters & HPC Ops
Senior AI Infra Architect: GPU Clusters & HPC Ops

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra DevOps Engineer - GPU/Cloud HPC
AI Infra DevOps Engineer - GPU/Cloud HPC

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 57,000 - 77,000
Senior AI Platform Engineer — Kubernetes & GPU Infra
Senior AI Platform Engineer — Kubernetes & GPU Infra

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Infra Engineer — GPUs, HPC & Kubernetes
Senior AI Infra Engineer — GPUs, HPC & Kubernetes

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Compute & HPC Infrastructure Engineer
Senior AI Compute & HPC Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: GPU Clusters & HPC Networking
AI Infra Engineer: GPU Clusters & HPC Networking

The Supreme HR Advisory Pte Ltd • Singapore

On-site
SGD 56,000 - 78,000