Senior Linux Infra Engineer: Bare Metal & AI GPU Storage

Uvation

Singapore

Remote

SGD 150,000 - 260,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Uvation seeks a Senior Linux Infrastructure Engineer to design, deploy, and operate large‑scale Linux‑based infrastructure powering enterprise workloads and AI/ML environments.

Role emphasizes hands‑on BMaaS, GPU infrastructure, high‑performance storage, and data centre operations with strong hardware‑level expertise. Collaboration with an established DevOps team is expected while focusing on infrastructure engineering rather than CI/CD pipelines.

Qualifications

  • Linux & Bare Metal Infrastructure Expert‑level Linux administration (Ubuntu required; RedHat and SUSE preferred)
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience operating Bare Metal as a Service (BMaaS) platforms and large‑scale infrastructure environments
  • Strong understanding of server hardware, including: BIOS/UEFI RAID controllers Firmware management iLO/iDRAC/IPMI NICs and SmartNICs HBA cards Hardware diagnostics and troubleshooting
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale
  • AI Factory & GPU Infrastructure Experience deploying and managing GPU‑accelerated infrastructure for AI/ML workloads
  • Understanding of NVIDIA GPU technologies including: A100, H10200, H200, B20 Hequivalent GPU platforms NVIDIA DGX and OEM GPU servers
  • GPU provisioning and lifecycle management
  • GPU monitoring and performance optimization
  • Knowledge of AI Factory architecture and infrastructure requirements
  • Experience supporting GPU clusters, AI training environments, and high‑performance computing (HPC) workloads
  • Understanding of: GPU resource allocation and scheduling Multi‑GPU systems GPU networking requirements
  • High‑bandwidth, low‑latency infrastructure design
  • Familiarity with NVIDIA ecosystem technologies such as: CUDA, NCCL, GPU Direct Storage, NVIDIA Fabric Manager, NVIDIA Base Command (preferred)
  • Enterprise Storage & Data Platforms
  • Advanced Linux storage administration: LVM, XFS, EXT4, NFS, iSCSI, Fibre Channel SAN, Multipath IO
  • Hands‑on Ceph expertise: Cluster architecture, MON, OSD, MDS, RBD, CephFS, RGW
  • Capacity planning, Performance tuning, Failure recovery
  • Experience with high‑performance AI storage platforms such as: WEKA, VAST Data, Dell PowerScale, Pure Storage FlashBlade, NetApp
  • NVMe‑over‑Fabric (NVMe‑oF), RDMA, GPU Direct Storage
  • Parallel file systems, AI data pipelines
  • Networking & Infrastructure
  • Bonding, VLANs, Routing, MTU optimization
  • High‑performance data centre networking: 100G/200G/400G Ethernet, RoCE
  • spine‑leaf architectures
  • Familiarity with NVIDIA Spectrum‑X, Mellanox/NVIDIA ConnectX adapters
  • Layer 2/3 infrastructure design and troubleshooting
  • Operations & Reliability
  • High availability, clustering, disaster recovery
  • Bash and Python scripting for automation
  • Runbooks and infrastructure standards
  • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
  • DCIM tools
  • IPAM solutions
  • AWS, Azure, or hybrid cloud exposure
  • We Are Not Looking For CI/CD pipeline engineers
  • Terraform/GitOps focus not primary

Responsibilities

  • Design, deploy, operate, and troubleshoot large‑scale Linux‑based infrastructure
  • Manage BMaaS platforms and bare metal lifecycle from hardware to firmware
  • Provision GPU infrastructure for AI/ML workloads and HPC environments
  • Optimize GPU provisioning, monitoring, and performance for AI factories
  • Hands‑on hardware diagnostics, BIOS/UEFI, RAID, iLO/iDRAC/IPMI, NICs, HBA, and SmartNICs
  • Architect scalable storage with Ceph and AI storage platforms
  • Plan capacity, ensure high availability, and disaster recovery readiness
  • Maintain high‑bandwidth, low‑latency networking for data centers
  • Collaborate with AI/ML teams to support GPU clusters and AI workflows
  • Document runbooks, standards, and operational practices
  • Track and implement NVIDIA ecosystem technologies and NVIDIA Base Command

Skills

Linux administration
Bare metal deployment
BMaaS platforms
GPU infrastructure
AI Factory environments
HPC workloads
GPU provisioning
GPU monitoring
GPU lifecycle management
Networking for data centers
NVIDIA GPU technologies
NVIDIA ecosystem tools
Linux performance tuning

Tools

Ceph
NVIDIA DGX
NVIDIA Base Command
NVIDIA Fabric Manager

Job description

Uvation seeks a Senior Linux Infrastructure Engineer to design, deploy, and operate large‑scale Linux‑based infrastructure powering enterprise workloads and AI/ML environments.

Role emphasizes hands‑on BMaaS, GPU infrastructure, high‑performance storage, and data centre operations with strong hardware‑level expertise. Collaboration with an established DevOps team is expected while focusing on infrastructure engineering rather than CI/CD pipelines.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer — GPU Cluster & ML Platform
Senior AI Infra Engineer — GPU Cluster & ML Platform

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infra Engineer (ML Platform, AI Native production, Algorithm, cutting-edge technology, multinational company)
AI Infra Engineer (ML Platform, AI Native production, Algorithm, cutting-edge technology, multinational company)

DADACONSULTANTS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Infra Architect: GPU Clusters & HPC
AI Infra Architect: GPU Clusters & HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Senior AI Infra Engineer: Virtualisation & HPC Storage
Senior AI Infra Engineer: Virtualisation & HPC Storage

Firmus Technologies • Singapore

On-site
SGD 232,378 - 335,657
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
AI Systems Infra Engineer - Multi-GPU HPC
AI Systems Infra Engineer - Multi-GPU HPC

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert
AI Infrastructure Engineer - GPU HPC & Kubernetes Expert

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
AI Infra Engineer: GPU HPC Clusters & Orchestration
AI Infra Engineer: GPU HPC Clusters & Orchestration

THE SUPREME HR ADVISORY PTE. LTD. • Singapore

On-site
SGD 56,000 - 78,000
Remote APAC InfraOps Engineer, GPU Infrastructure
Remote APAC InfraOps Engineer, GPU Infrastructure

Lightning AI • Singapore

On-site
SGD 165,000 - 205,000
Health coverage
Equity RSUs
401(k) matching
+8
Senior AI Storage Architect for GPU-Accelerated Cloud
Senior AI Storage Architect for GPU-Accelerated Cloud

Bitdeer Technologies Group • Singapore

On-site
SGD 120,000 - 160,000
Welfare benefits
Training & mentoring
Development opportunities