Site Reliability Engineer, AI Infrastructure

Firmus

City of Melbourne

On-site

AUD 120,000 - 160,000

Full time

2 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Firmus Technologies in Melbourne, Australia, is seeking a Site Reliability Engineer for AI Infrastructure to join the Operations team. You will support daily operations of the AI-accelerated HPC infrastructure, collaborating with Field Service Engineers, HPC and Network teams, and the GOC.

This is a hands-on role focused on reliability, performance, and scalable control of cutting-edge AI compute resources.

Qualifications

  • Bachelor’s degree in a technical field is required.
  • Minimum 5 years in field service technical roles with hardware focus.
  • Strong Linux troubleshooting and security standard adherence.
  • Proficient scripting in Bash and Python, with automation experience.
  • Experience with Slurm, Kubernetes, and observability tools.

Responsibilities

  • Deploy, configure, and maintain AI infrastructure and GPU servers.
  • Perform hardware diagnostics, firmware updates, and security checks.
  • Collaborate with HPC/NVA engineers for tailored deployments.
  • Provide first-line engineering support for onsite issues and on-call rotation.
  • Document incidents, resolutions, and best practices for knowledge sharing.
  • Maintain up-to-date documentation and communicate with stakeholders.

Skills

Field service exp
Linux environments
Scripting Bash/Python
Security standards
Problem solving
Communication skills

Education

Bachelor's degree in computer engineering/CS

Tools

Slurm
Kubernetes
NVIDIA BCM
Prometheus
Grafana
ELK

Job description

Site Reliability Engineer, AI Infrastructure

Melbourne, Australia

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

Whyyou’lllove working here

As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.

We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work.

Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.

Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.

What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in
rather than drawing from them.

Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.

ROLE SUMMARY

Firmus Technologies is seeking a skilled Site Reliability Engineer, AI Infrastructure to join our Operations team, supporting the daily operations and maintenance of our AI-accelerated high-performance computing (HPC) infrastructure. This role will work closely with Field Service Engineers, HPC and Network Engineering teams, and assist the Global Operations Centre (GOC). This is a unique opportunity to contribute directly to the stability and growth of cutting-edge AI infrastructure.

KEY RESPONSIBILITIES

  • Support in the deployment, configuration, and maintenance of various high-end GPU servers, storage servers, networkingequipmentand software components in highly secure environments.
  • Perform hardware diagnostics, systems functionality and firmware updates asrequired.
  • Collaborate with engineering teams toassistin tailored customer environments deployment (eg: bare-metal systems, HPC Clusters, Kubernetes,Slurmetc).
  • Serve as first line of engineering support for onsite operational issues, including troubleshooting hardware,networkand software problems, and firmware compliance.
  • Troubleshoot incidents, elevate criticalissuesand provide feedback toappropriate teamsfor improvements.
  • Participate in an on-call rotation to ensure 24/7 availability and responsiveness to critical issues.
  • Provide technical support to the GOC Support Specialist team in troubleshootingcompute infrastructurerelated problems.
  • Document incident details, resolutions, and lessons learned to enhance future problem-solving.
  • Maintain clear,accurate, and up-to-date documentation to promote effective knowledge sharing across the team.
  • Communicate effectively with GOC, HPC Engineers, internal teams, stakeholders, and end-users to ensure alignment on issue resolution.
  • Take part in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning.

SKILLS AND EXPERIENCE

  • Bachelor’s degree in computer engineering, computer science, or a related technical field.
  • 5+ years of experience in field service technical areas.
  • Strong understanding of server hardware technology, firmware lifecycle, Linux environments and troubleshooting hardware problems, with adherence to physical and system-level security standards.
  • Experience with scripting languages ( eg : Bash, Python)
  • Familiarity with using configuration management, CICD tools, workload manager and cluster softwares ( eg : Slurm , Kubernetes, Nvidia BCM) and Observability tools ( eg : Prometheus, Grafana, ELK, etc)
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and as part of a team.
  • Strong communication skills, both written and verbal.

Location

This role is based in Brooklyn, Melbourne, Australia.

Employment Basis

Full-time

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Matchbox • City of Launceston

On-site
AUD 110,000 - 170,000
Site Reliability Engineer
Site Reliability Engineer

Firmus Technologies • City of Launceston

On-site
AUD 90,000 - 120,000
Data Center Field Service Technician
Data Center Field Service Technician

Firmus Technologies • City of Launceston

On-site
AUD 65,000 - 90,000
Principal Ai Infrastructure Engineer, Kubernetes
Principal Ai Infrastructure Engineer, Kubernetes

Uncover • Sydney

On-site
AUD 180,000 - 240,000
Senior Site Reliability Engineer – AI HPC Infra
Senior Site Reliability Engineer – AI HPC Infra

Firmus Technologies • City of Launceston

On-site
AUD 90,000 - 120,000
Compliance & GRC Lead for AI Infrastructure
Compliance & GRC Lead for AI Infrastructure

Firmus Technologies • Sydney

On-site
AUD 140,000 - 190,000
Senior Power Systems SCADA and Controls Engineer
Senior Power Systems SCADA and Controls Engineer

Firmus Technologies • City of Launceston

On-site
AUD 130,000 - 170,000
Senior HPC Infrastructure Engineer
Senior HPC Infrastructure Engineer

Matchbox • Council of the City of Sydney

On-site
AUD 180,000 - 240,000
Compliance Specialist
Compliance Specialist

Firmus Technologies • Sydney

On-site
AUD 140,000 - 190,000
Mechanical Project Engineer
Mechanical Project Engineer

Firmus Technologies • City of Launceston

On-site
AUD 90,000 - 135,000