Senior AI Scheduling & Orchestration Engineer

Bitdeer

Singapore

On-site

SGD 180,000 - 240,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Bitdeer in Singapore is seeking a senior distributed systems engineer to design and optimize its Kubernetes-based scheduling stack for multi-node AI workloads and HPC environments. You will lead the orchestration of GPU resources and work with storage teams to support production-scale workloads.

We value 6+ years of experience, strong communication, and hands-on expertise with PyTorch, MPI, Terraform, and Terraform-driven automation. Join us to shape the future of AI cloud services.

Qualifications

  • 6+ years of distributed systems engineering experience.

Responsibilities

  • Design and implement advanced batch scheduling architectures using Kubernetes-based frameworks to support multi-node scheduling.
  • Develop cluster-wide admission control and queueing mechanisms to manage high-volume AI workloads.
  • Leverage DRA and custom scheduler plugins to efficiently allocate GPU resources.
  • Architect topology-aware pod placement for low-latency communication across NVLink/InfiniBand fabrics.
  • Implement automated GPU sharing technologies and multi-tenancy isolation policies.
  • Collaborate with GPU Systems and Storage teams to ensure tight integration with bare-metal hardware and storage I/O.
  • Drive reliability and scalability of the scheduling stack, resolving contention and deadlocks in HPC.
  • Mentor junior engineers and conduct design reviews to maintain architectural excellence.

Skills

Kubernetes scheduling
Distributed systems engineering
AI workload patterns
GPU/ HPC orchestration
Terraform
Go-based Operators
PyTorch
MPI

Education

Bachelor’s or Master’s in CS/EE

Tools

Kubernetes
Terraform
Go
PyTorch
MPI
Ray

Job description

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

What you will be responsible for:
  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively.
  • Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics.
  • Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization.
  • Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
  • Drive the reliability and scalability of the scheduling stack, resolving resource contention and deadlock scenarios in large-scale HPC environments.
  • Mentor junior engineers and conduct design reviews to maintain architectural excellence in our orchestration layer.
How you will stand out:
  • Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
  • Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI).
  • Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments.
  • Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference.
  • Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
  • Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals.
  • Ability to work in a high-velocity engineering environment and translate complex, ambiguous requirements into concrete, scalable engineering solutions.
What you will experience working with us:
  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Scheduling & Orchestration Engineer
Senior AI Scheduling & Orchestration Engineer

Bitdeer Technologies Group • Singapore

On-site
SGD 180,000 - 240,000
Welfare benefits
Training & mentoring
Startup spirit
Senior AI Scheduling & Orchestration Engineer
Senior AI Scheduling & Orchestration Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer Group • Singapore

On-site
SGD 120,000 - 180,000
Senior AI Platform Engineer
Senior AI Platform Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 260,000
Attractive welfare benefits
Career development opportunities
Hybrid/onsite options
Senior Kubernetes Control Plane Engineer
Senior Kubernetes Control Plane Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 180,000 - 240,000
Senior AI Observability & Telemetry Engineer
Senior AI Observability & Telemetry Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Senior GPU Systems & Fabric Engineer
Senior GPU Systems & Fabric Engineer

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 150,000 - 190,000
Sr. GPU Cloud K8S Expert (SRE SME)
Sr. GPU Cloud K8S Expert (SRE SME)

Bitdeer • Singapore

On-site
SGD 150,000 - 230,000
Sr. GPU Cloud K8S Expert (SRE SME)
Sr. GPU Cloud K8S Expert (SRE SME)

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 180,000
Sr. GPU Cloud K8S Expert (SRE SME)
Sr. GPU Cloud K8S Expert (SRE SME)

Bitdeer (NASDAQ: BTDR) • Singapore

On-site
SGD 120,000 - 180,000
Attractive welfare benefits
Mentoring and professional development