Site Reliability Engineer — AI HPC Infrastructure

Firmus Technologies

Singapore

On-site

SGD 120,000 - 160,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Firmus Technologies is seeking a Site Reliability Engineer to support the daily ops of our AI-accelerated HPC infrastructure. You will work with Field Service Engineers and HPC teams to ensure stability and performance of GPU servers, storage, and networking in secure environments.

Ideal candidates have a Bachelor's degree in a related field and 5+ years in field service, with strong Linux, scripting, Kubernetes, Slurm, and observability tools to diagnose and resolve complex issues.

Qualifications

  • Bachelor's degree in computer engineering, computer science, or a related technical field.
  • 5+ years of field service technical experience.
  • Strong understanding of server hardware, firmware lifecycle, Linux, and troubleshooting.
  • Experience with scripting (Bash, Python).
  • Familiarity with Slurm, Kubernetes and observability tools.

Responsibilities

  • Deploy, configure and maintain GPU servers and storage in secure environments.
  • Perform hardware diagnostics, systems checks and firmware updates.
  • Collaborate to deploy customer environments (bare-metal, HPC clusters, Kubernetes, Slurm).
  • Serve as first line of engineering support for onsite issues (hardware, network, software).
  • Troubleshoot incidents and escalate issues for improvements.
  • Participate in on-call rotation for 24/7 availability.
  • Document incident details and resolutions to improve future problem-solving.
  • Maintain up-to-date documentation for the team.

Skills

Linux
Bash
Python
Kubernetes
Slurm
NVIDIA BCM
Prometheus
Grafana
ELK
Troubleshooting
Communication

Education

Bachelor's degree in CS/CE

Tools

Configuration management
CI/CD tools
Workload manager
Cluster software

Job description

Firmus Technologies is seeking a Site Reliability Engineer to support the daily ops of our AI-accelerated HPC infrastructure. You will work with Field Service Engineers and HPC teams to ensure stability and performance of GPU servers, storage, and networking in secure environments.

Ideal candidates have a Bachelor's degree in a related field and 5+ years in field service, with strong Linux, scripting, Kubernetes, Slurm, and observability tools to diagnose and resolve complex issues.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Engineer: Virtualisation & HPC Storage
Senior AI Infra Engineer: Virtualisation & HPC Storage

Firmus Technologies • Singapore

On-site
SGD 232,378 - 335,657
Site Reliability Engineer
Site Reliability Engineer

Firmus Technologies • Singapore

On-site
SGD 120,000 - 160,000
Senior AI Infrastructure Engineer (Virtualisation)
Senior AI Infrastructure Engineer (Virtualisation)

Firmus Technologies • Singapore

On-site
SGD 232,378 - 335,657
Hardware Engineer
Hardware Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Senior Infrastructure Engineer - HPC
Senior Infrastructure Engineer - HPC

Kerry Consulting • Singapore

On-site
SGD 90,000 - 140,000
Senior Platform Engineer
Senior Platform Engineer

Firmus Technologies • Singapore

On-site
SGD 140,000 - 260,000
Server Engineer(AI Cluster)
Server Engineer(AI Cluster)

PaleBlueDot AI • Singapore

On-site
SGD 120,000 - 170,000
Systems Engineer (HPC)
Systems Engineer (HPC)

FUJITSU ASIA PTE LTD • Singapore

On-site
SGD 90,000 - 150,000
Senior AI/HPC Systems Engineer
Senior AI/HPC Systems Engineer

NVIDIA Gruppe • Singapore

On-site
SGD 120,000 - 180,000
System Engineer
System Engineer

RUNSUN SERVICE PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000