Site Reliability Engineer

Firmus Technologies

Singapore

On-site

SGD 120,000 - 160,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Firmus Technologies is seeking a Site Reliability Engineer to support the daily ops of our AI-accelerated HPC infrastructure. You will work with Field Service Engineers and HPC teams to ensure stability and performance of GPU servers, storage, and networking in secure environments.

Ideal candidates have a Bachelor's degree in a related field and 5+ years in field service, with strong Linux, scripting, Kubernetes, Slurm, and observability tools to diagnose and resolve complex issues.

Qualifications

  • Bachelor's degree in computer engineering, computer science, or a related technical field.
  • 5+ years of field service technical experience.
  • Strong understanding of server hardware, firmware lifecycle, Linux, and troubleshooting.
  • Experience with scripting (Bash, Python).
  • Familiarity with Slurm, Kubernetes and observability tools.

Responsibilities

  • Deploy, configure and maintain GPU servers and storage in secure environments.
  • Perform hardware diagnostics, systems checks and firmware updates.
  • Collaborate to deploy customer environments (bare-metal, HPC clusters, Kubernetes, Slurm).
  • Serve as first line of engineering support for onsite issues (hardware, network, software).
  • Troubleshoot incidents and escalate issues for improvements.
  • Participate in on-call rotation for 24/7 availability.
  • Document incident details and resolutions to improve future problem-solving.
  • Maintain up-to-date documentation for the team.

Skills

Linux
Bash
Python
Kubernetes
Slurm
NVIDIA BCM
Prometheus
Grafana
ELK
Troubleshooting
Communication

Education

Bachelor's degree in CS/CE

Tools

Configuration management
CI/CD tools
Workload manager
Cluster software

Job description

Firmus Technologies

Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs. Our mission is to create the most energy-efficient AI infrastructure, combining cutting edge technology with a steadfast commitment to sustainability.



Firmus Technologies

Firmus Technologies is a global leader pioneering the solution to AI’s energy challenge, founded in Australia in 2019 by a visionary team of entrepreneurs. Our mission is to create the most energy-efficient AI infrastructure, combining cutting edge technology with a steadfast commitment to sustainability. Through ground-breaking research and development, we invented a verticalized AI Factory - a new class of digital infrastructure that replaces traditional data centres. Built on new approaches to liquid cooling, energy management, water use and modular construction methodology, the Firmus AI Factory delivers low-cost AI tokens across Asia-Pacific.



Firmus AI Cloud

We provide customers with access to energy savings via our large-scale GPU cloud, Firmus AI Cloud. Rated Silver in The GPU Cloud ClusterMAX™ Rating System, our cloud empowers developers, enterprise, education and government users to train AI models with unmatched efficiency and cost savings. With an ever-growing list of services and applications, we are committed to building a cloud experience for our customers that is market-leading, proprietary and built to scale.



Why you’ll love working here


  • A fast-paced and dynamic environment working with next-gen technology. You’ll be operating at the intersection of sustainability and artificial intelligence – helping to transform an industry.

  • Working with and access to colleagues who are true innovators and leaders in their field.

  • As an emerging company, we work as a close-knit team. Work with the founders, grow a strong network, and witness the impact you make first-hand as we democratise AI tools for everyone – more sustainably, and more affordably.

  • We believe that people from diverse backgrounds come together to do their best work, be their authentic selves, and build great things. We are proud to be an equal opportunity employer.



Role Summary

Firmus Technologies is seeking a skilled Site Reliability Engineer to join our Operations team, supporting the daily operations and maintenance of our AI-accelerated High-Performance Computing (HPC) infrastructure. This role will work closely with Field Service Engineers, HPC and Network Engineering teams, and assist the Global Operations Centre (GOC). This is a unique opportunity to contribute directly to the stability and growth of cutting-edge AI infrastructure.



Key Responsibilities


  • Support in the deployment, configuration, and maintenance of various high-end GPU servers, storage servers, networking equipment and software components in highly secure environments.

  • Perform hardware diagnostics, systems functionality and firmware updates as required.

  • Collaborate with engineering teams to assist the tailored customer environments deployment (eg: bare-metal systems, HPC Clusters, Kubernetes, Slurm etc).

  • Serve as first line of engineering support for onsite operational issues, including troubleshooting hardware, network and software problems, and firmware compliance.

  • Troubleshoot incidents, elevate critical issues and provide feedback to appropriate teams for improvements.

  • Participate in an on-call rotation to ensure 24/7 availability and responsiveness to critical issues.

  • Provide technical support to the GOC Support Specialist team in troubleshooting compute infrastructure related problems.

  • Document incident details, resolutions, and lessons learned to enhance future problem-solving.

  • Maintain clear, accurate, and up-to-date documentation to promote effective knowledge sharing across the team.

  • Communicate effectively with GOC, HPC Engineers, internal teams, stakeholders, and end-users to ensure alignment on issue resolution.

  • Take part in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning.



Skills And Experience


  • Bachelor’s degree in computer engineering, computer science, or a related technical field.

  • 5+ years of experience in field service technical areas.

  • Strong understanding of server hardware technology, firmware lifecycle, Linux environments and troubleshooting hardware problems, with adherence to physical and system-level security standards.

  • Experience with scripting languages (eg: Bash, Python)

  • Familiarity with using configuration management, CICD tools, workload manager and cluster softwares (eg: Slurm, Kubernetes, Nvidia BCM) and Observability tools (eg: Prometheus, Grafana, ELK, etc)

  • Excellent problem-solving and analytical skills.

  • Ability to work independently and as part of a team.

  • Strong communication skills, both written and verbal.



Location & Reporting


  • Based in: Singapore

  • Reporting to: Senior Operations Manager



Employment Basis

Full-time



Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.



Join us in our mission to revolutionize the AI industry through sustainable practices and cutting-edge engineering.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer (Virtualisation)
Senior AI Infrastructure Engineer (Virtualisation)

Firmus Technologies • Singapore

On-site
SGD 232,378 - 335,657
Network Engineer
Network Engineer

Firmus Technologies • Singapore

On-site
SGD 120,000 - 180,000
Strategic Writer
Strategic Writer

Firmus Technologies • Singapore

Hybrid
SGD 90,000 - 130,000
Principal Solutions Architect, AI Data Infrastructure
Principal Solutions Architect, AI Data Infrastructure

Firmus Technologies • Singapore

On-site
SGD 120,000 - 160,000
Senior Platform Engineer
Senior Platform Engineer

Firmus Technologies • Singapore

On-site
SGD 140,000 - 260,000
Procurement Executive
Procurement Executive

Firmus Technologies • Singapore

On-site
SGD 60,000 - 90,000
System Engineer for a Cloud-like HPC Infrastructure
System Engineer for a Cloud-like HPC Infrastructure

Singapore-ETH Centre • Singapore

Hybrid
SGD 90,000 - 130,000
Flexible hybrid work (2 days from home
Healthcare insurance
NS mark certification
+1
Senior Mechanical Engineer
Senior Mechanical Engineer

firmus technologies • Singapore

On-site
SGD 80,000 - 110,000
System Engineer for a Cloud-like HPC Infrastructure
System Engineer for a Cloud-like HPC Infrastructure

Immigration Policy Lab • Singapore

Hybrid
SGD 120,000 - 180,000
Flexible hybrid work
25 days annual leave
Dental benefits
+2
Site Reliability Engineer — AI HPC Infrastructure
Site Reliability Engineer — AI HPC Infrastructure

Firmus Technologies • Singapore

On-site
SGD 120,000 - 160,000