Site Reliability Engineer: AI Infra & HPC

Firmus Technologies

City of Melbourne

On-site

AUD 120,000 - 180,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Firmus Technologies in Melbourne is seeking a skilled Site Reliability Engineer, AI Infrastructure, to join our Operations team. You will support daily operations and maintenance of our AI-accelerated HPC infrastructure, collaborating with Field Service Engineers and HPC/Network teams.

This role offers on-site work in Australia with opportunities to influence 24/7 reliability, security standards, and deployment in Kubernetes and Slurm environments, while advancing your career in a fast-growing

Qualifications

  • Bachelor's degree in a technical field required.
  • 5+ years in field service technical areas.
  • Strong understanding of server hardware, Linux, and security standards.
  • Experience with scripting languages (Bash, Python).
  • Familiarity with configuration management, CI/CD, workload managers and cluster software (Slurm, Kubernetes, Nvidia BCM) and observability tools (Prometheus, Grafana, ELK).
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and in a team, with strong communication.

Responsibilities

  • Support deployment, configuration, and maintenance of high-end GPU servers and related software in secure environments.
  • Perform hardware diagnostics, system functionality checks and firmware updates as required.
  • Collaborate with engineering to tailor customer environments (bare-metal, HPC, Kubernetes, Slurm).
  • Serve as first-line engineering support for onsite issues including hardware, network and software problems.
  • Troubleshoot incidents, escalate critical issues, and provide feedback for improvements.
  • Participate in on-call rotation for 24/7 availability.
  • Document incident details, resolutions, and lessons learned for knowledge sharing.
  • Maintain up-to-date documentation and communicate with stakeholders.

Skills

Server hardware
Linux admin
Bash & Python
Config management
CI/CD tools
Slurm
Kubernetes
Nvidia BCM
Observability
Problem solving
Communication

Education

Bachelor’s degree in computer engineering, computer science, or related technical field

Tools

Kubernetes
Slurm
Nvidia BCM
Prometheus
Grafana
ELK

Job description

Firmus Technologies in Melbourne is seeking a skilled Site Reliability Engineer, AI Infrastructure, to join our Operations team. You will support daily operations and maintenance of our AI-accelerated HPC infrastructure, collaborating with Field Service Engineers and HPC/Network teams.

This role offers on-site work in Australia with opportunities to influence 24/7 reliability, security standards, and deployment in Kubernetes and Slurm environments, while advancing your career in a fast-growing

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Site Reliability Engineer, AI Infrastructure
Site Reliability Engineer, AI Infrastructure

Firmus Technologies • City of Melbourne

On-site
AUD 120,000 - 180,000
Senior AI Platform Reliability Engineer
Senior AI Platform Reliability Engineer

Matchbox • City of Melbourne

Hybrid
AUD 180,000 - 260,000
Senior Kubernetes Engineer: AI Infra at Scale
Senior Kubernetes Engineer: AI Infra at Scale

Firmus Technologies Pty Ltd. • City of Melbourne

On-site
AUD 180,000 - 240,000
Senior AI Infra Engineer — Kubernetes for GPU Clusters
Senior AI Infra Engineer — Kubernetes for GPU Clusters

Matchbox • Sydney

Hybrid
AUD 180,000 - 240,000
Senior Platform Reliability Engineer - AI FactoryOS, 24/7
Senior Platform Reliability Engineer - AI FactoryOS, 24/7

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

Firmus Technologies Pty Ltd. • City of Melbourne

On-site
AUD 180,000 - 240,000
Senior AI Infrastructure Engineer, Kubernetes
Senior AI Infrastructure Engineer, Kubernetes

Matchbox • Sydney

Hybrid
AUD 180,000 - 240,000
Data Centre Field Service Technician
Data Centre Field Service Technician

Matchbox • City of Melbourne

Hybrid
AUD 75,000 - 110,000
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Field HPC Infrastructure Engineer
Field HPC Infrastructure Engineer

Nityo Infotech • City of Melbourne

On-site
AUD 90,000 - 120,000