DevOps Engineer, GPUaaS

Singtel

Singapore

On-site

Confidential

Full time

3 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Singtel is seeking an DevOps Engineer for its GPU-as-a-Service (GPUaaS) within the RE:AI division. You will design and operate large GPU clusters for AI/ML workloads, automate provisioning, and build CI/CD pipelines for AI models.

You will monitor performance and troubleshoot GPU frameworks, drivers, and networking in a multi-tenant cloud environment. You will collaborate with software and admin teams, enhance automation, and support users of GPU-accelerated systems across on‑prem and cloud

Qualifications

  • Bachelor’s degree in CS/Engineering or related field.
  • Strong Linux system administration skills.
  • Experience with Jenkins, Kubernetes, Ansible, Terraform.
  • CI/CD, automation and monitoring principles.

Responsibilities

  • Design, deploy and support large-scale GPU clusters for AI/ML workloads.
  • Automate provisioning of GPU resources on‑prem and cloud.
  • Implement CI/CD pipelines for AI models and GPU apps.
  • Monitor cluster health, performance and availability.
  • Troubleshoot Slurm, Kubernetes, GPU drivers and CUDA.
  • Ensure security for multi-tenant GPUaaS environments.
  • Collaborate with software and admin teams to streamline workflows.
  • Provide user support for GPU-accelerated systems.

Skills

Linux administration
Kubernetes
Jenkins
Ansible
Terraform
Python
Bash scripting
CI/CD
Monitoring
Zabbix
Prometheus
TensorFlow
PyTorch
NVIDIA GPUs
Docker
MPI
RDMA
NCCL
Slurm
InfiniBand
Networking

Education

Bachelor’s degree in Computer Science/Engineering

Tools

Docker
Kubernetes
Prometheus
Zabbix
NVIDIA CUDA
Slurm

Job description

About Singtel Digital InfraCo – RE:AI

Singtel Digital InfraCo’s RE:AI division is building Asia’s most advanced and sustainable AI infrastructure ecosystem. RE:AI enables enterprises, research institutions, and digital-native businesses to accelerate innovation through responsible, high-performance AI compute and connectivity solutions.

About Singtel Digital InfraCo – RE:AI

Singtel Digital InfraCo’s RE:AI division is building Asia’s most advanced and sustainable AI infrastructure ecosystem. RE:AI enables enterprises, research institutions, and digital-native businesses to accelerate innovation through responsible, high-performance AI compute and connectivity solutions.

Be a Part of Something BIG!

As an DevOps Engineer for SingTel’s GPU-as-a-Service (GPUaaS), you will help in implementing processes and integration of operations to advance customer’s AI and HPC capabilities. You will be exposed to both physical data center implementation and software solutions in a Singtel GPU-as-a-Service (GPUaaS). This position requires a forward-thinking individual who thrives in dynamic environments and is committed to driving continuous improvement in GPU for AI and HPC environments. This role is suitable for professionals looking to develop their expertise in DevOps and AI/HPC cloud platforms.

Make an impact by
  • Design, deploy and support large-scale, distribute GPU clusters for AI and ML workloads.
  • Manage and automate provisioning of GPU resources in both on-prem and cloud platforms.
  • Design, implement and manage CI/CD pipelines for AI models and GPU-accelerated applications.
  • Monitor cluster usage, health, performance and availability.
  • Improve infrastructure provisioning, management, and monitoring through automation.
  • Troubleshoot compute resource system level issues such as Slurm, Kubernetes, GPU drivers, CUDA, IB networking.
  • Optimize system parameters (e.g., OS, drivers, networking, library) for AI workload performance.
  • Conduct GPU cluster benchmark and keeping up with the latest advancements in GPU technology.
  • Set up monitoring and logging for GPU resources using Zabbix, Prometheus, NVIDIA DCGM and other tools.
  • Implement security best-practices for multi-tenant GPU-as-a-Service (GPUaaS) environment.
  • Collaborate with software and administrator to to streamline workflows and improve collaboration.
  • Providing technical support and guidance to users of GPU-accelerated systems.
  • Work with senior DevOps engineer to identify bottlenecks and improve development and operational processes for AI and HPC GPU cloud.
  • Learning to solve problems in high-performance distributed computation for AI and HPC GPU cloud computing.
  • Participate in rotational or scheduled shift work as required to support platform operations.
Skills For Success
  • Bachelor’s degree in Computer Science/Engineering, Information Technology, Systems Engineering, or a related field.
  • Strong Linux system administration skills in Ubuntu/CentOS/Rocky Linux, etc.
  • Experience with DevOps tools such as Jenkins, Kubernetes, Ansible and Terraform.
  • Solid understanding of DevOps practices, including CI/CD, automation, and monitoring.
  • Proficiency in scripting languages (e.g., Python, Bash).
  • Experience in implementing monitoring solutions such as Zabbix, Prometheus.
  • Familiarity with AI frameworks such as TensorFlow, PyTorch.
  • Understanding of cloud architectures (IaaS, PaaS), GPU architecture and NVIDIA GPUs.
  • Strong verbal, written, and presentation skills in English.
  • Team player with experience in cross-functional coordination.
  • Strong technical problem solving and analytical skills for system optimization.
Desirable Qualifications
  • Understanding of how collective communications (MPI, RDMA, and NCCL) works, as well as an understanding of GPU specific acceleration works on GPU cluster.
  • Knowledge of DevOps/ML Ops technologies in GPU cluster such as Docker/containers, Kubernetes, data center deployments
  • Familiarity with Slurm or other HPC workload managers to manage GPU clusters.
  • Understanding of AI & HPC networking technologies such as InfiniBand, RoCE, DPUs.
  • System-level experience specifically GPU-based systems (NVIDIA GPU and SDKs)
  • Understanding how AI and HPC workloads interact with both GPU HW and SW infrastructure.
Your Career Growth Starts Here.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

DevOps Engineer, GPUaaS
DevOps Engineer, GPUaaS

Singtel Group • Singapore

On-site
SGD 120,000 - 170,000
DevOps Engineer, GPUaaS
DevOps Engineer, GPUaaS

Singapore Telecommunications Limited • Singapore

On-site
SGD 120,000 - 180,000
Data Centre Operations Engineer
Data Centre Operations Engineer

Facade Today • Singapore

On-site
SGD 44,000 - 66,000
Health and wellness benefits
Training and development programs
Internal mobility opportunities
Network Engineer, GPUaaS
Network Engineer, GPUaaS

Singtel • Singapore

On-site
SGD 60,000 - 100,000
Full suite of health and wellness benefits
Ongoing training and development programs
Internal mobility opportunities
Data Centre Operations Engineer
Data Centre Operations Engineer

Singtel • Singapore

On-site
Confidential
Health and wellness benefits
Training and development programs
Internal mobility opportunities
Project Director, GPUaaS
Project Director, GPUaaS

Singapore Telecommunications Limited • Singapore

On-site
SGD 180,000 - 260,000
Flexible work arrangements
Health and wellness benefits
Training and development programs
+1
Network Engineer, GPUaaS
Network Engineer, GPUaaS

Singtel Group • Singapore

On-site
SGD 70,000 - 90,000
Full suite of health and wellness benefits
Ongoing training and development programs
Internal mobility opportunities
Project Director, GPUaaS (Singapore, Singapore)
Project Director, GPUaaS (Singapore, Singapore)

wimatec MATTES GmbH • Singapore

On-site
SGD 120,000 - 160,000
Flexible work arrangements
Full suite of health and wellness benefits
Ongoing training and development programs
+1
GPUaaS DevOps Engineer for AI & HPC
GPUaaS DevOps Engineer for AI & HPC

Singtel • Singapore

On-site
Confidential
GPUaaS Project Director: AI Infrastructure Leader
GPUaaS Project Director: AI Infrastructure Leader

Singtel • Singapore

On-site
SGD 100,000 - 130,000
Flexible work arrangements
Full suite of health and wellness benefits
Ongoing training and development programs
+1