Infrastructure Automation Engineer, AI Cluster Commissioning

Matchbox

City of Melbourne

On-site

AUD 140,000 - 210,000

Full time

3 days ago
Be an early applicant
Application generator

Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Firmus Technologies is seeking an Infrastructure Automation Engineer to join our Commissioning team in Australia. The role focuses on developing, maintaining, and improving software tools for commissioning from power-on through bring-up, testing, and hand-over to operations.

You will work with open-source and closed-source tools, cloud services, and SaaS platforms to deliver production-ready infrastructure.

Qualifications

  • Bachelor's degree in computer science, engineering, or related field.
  • 5+ years in software automation, tooling, or infrastructure automation for Linux/HPC/AI environments.
  • Strong scripting skills (bash, Python) and experience with IaC.
  • Experience configuring Linux servers, networks, and storage.
  • Familiarity with AI/HPC tooling (Slurm, NVIDIA DCGM, NCCL, Kubernetes).

Responsibilities

  • Develop custom automation tools for system bring-up and testing.
  • Implement monitoring, diagnostics, and health-checks for early detection of issues.
  • Automate network, storage, and communication tooling for efficient commissioning.
  • Create documentation to support operations and future projects.
  • Develop CI/CD pipelines for tooling and configuration changes.

Skills

Academic background
Automation
Linux
Scripting
REST APIs

Education

Bachelor's degree in computer science or engineering

Tools

Redfish
IPMI
DCGM
Prometheus
Grafana

Job description

Infrastructure Automation Engineer, AI Cluster Commissioning

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure - the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

Why Firmus?

As an NVIDIA Cloud and Engineering partner in Asia Pacific, you will gain skills, experience, and exposure across the AI industry and be part of shaping what this industry looks like for decades to come.

We are founder-led, not a big corporate. Decisions happen fast, our leaders are accessible, and there's minimum bureaucracy between you and the work. Ownership comes early. Whatever your role, you will have a direct line to outcomes, helping shape how the business grows as we scale nationally across a long-term, large-scale roadmap.

Work alongside founders and experts in AI infrastructure, energy systems and next-generation compute.

What we build here has impact beyond the business. Our AI Factories are designed to operate as assets to the energy grid to actively strengthen the communities and regions they operate in rather than drawing from them.

Considering applying? You don't need a perfect background to join our team. If you're driven and curious, there's a path for you. We back our people to grow into new domains and take on challenges beyond their previous experience.

Role Summary

Firmus Technologies is seeking a skilled Infrastructure Automation Engineer to join our Commissioning team. The position will play a crucial role in the development, maintenance, and improvement of the software tools involved in the commissioning process from initial power-on through integration, phased bring-up, component and system testing, benchmarking, and acceptance testing to hand-over to operations. These tools are used by internal teams, partners, and vendors working together to transform physically built infrastructure into a production-ready platform. This role involves custom tool development, integration with a wide variety of other closed-source and open-source tools, cloud services, and SaaS platforms. We are deploying leading edge AI Factory technology and the tools supporting the commissioning process are critical for successful on-time delivery. This role offers an exciting opportunity to work at the forefront of AI technology and contribute to the growth of AI infrastructure.

Key Responsibilities
Hardware bring-up

In collaboration with internal teams, partners, and vendors, develop custom tools and automate processes for initial system bring-up, hardware configuration, firmware updates, system inventory, testing, and issue tracking. Create accompanying documentation to support the team throughout the commissioning process, for the operations team once the system goes into production, and to accelerate future projects.

System and status monitoring

Work with vendors to implement, configure, and adapt system and status monitoring tools suitable for implementation early in the commissioning phase. Implement health-checks and diagnostics tools which assist remediation teams in finding root cause and resolving issues. Adapt these tools for new platforms, new features, new metrics, and prepare for integration into production systems.

Network tools

Commissioning of the network also requires suitable tools and processes, which need to be automated and streamlined. In close collaboration with the networking team, vendors, and partners implement suitable tools and an efficient process from initial power-on and OS deployment to firmware updates and network configuration. In addition, develop and implement suitable diagnostics and monitoring tools which are critical to resolve issues from the physical layer through to the application layer.

Storage

Develop tooling and automation frameworks supporting deployment, configuration, benchmarking, and validation of storage platforms.

Communication tools

Efficient and effective communication between various teams involved in the commissioning process needs to be supported by suitable tools. Integrate messaging tools, inventory databases, component trackers, issue trackers, time and resource trackers, reporting platforms, dashboards, and operational analytics to avoid duplication and support a fast, reliable, and repeatable commissioning process.

Testing and benchmarking

Develop test and benchmarking tools, customise them for the specific project and customer requirements, and execute the tests to confirm the successful completion of the commissioning process.

Automation and documentation

Leverage modern Infrastructure as Code (IaC) tools to simulate, automate and accelerate the commissioning process. Implement CI/CD pipelines for tool and configuration changes to improve speed, consistency, and auditability. Automate repeatable tasks, document the tools and processes, and prepare them for hand-over to operations teams.

Security

In all tools and throughout the deployment process implement best security practices including security updates and patches, secure handling of credentials, keys, and secrets, secure system setup, network and storage configurations.

Success Measures
  • Reduced commissioning time through automation.
  • High reliability and adoption of commissioning tools.
  • Successful automation of repeatable deployment activities.
  • Effective monitoring and diagnostic coverage.
  • Accurate documentation and operational handover.
  • Reduction in manual effort and deployment errors.
Skills & Experience
  • Bachelor's degree in computer science, engineering, or a related technical field.
  • 5+ years of experience developing software, automation platforms, operational tooling, or infrastructure automation solutions for large-scale Linux, cloud, HPC, or AI environments.
  • Solid understanding of advanced hardware technologies, particularly high performance GPU infrastructure, high end network environments, and high performance storage solutions.
  • Experience in Linux systems including server, network, and storage configuration.
  • Experience developing production software tools and APIs, consuming and integrating REST APIs, applying modern software development practices, building automated testing frameworks.
  • Experience using scripting, automation, and IaC tools including bash, Python, Ansible.
  • Experience using infrastructure tools and APIs including Redfish, IPMI, DCGM, Prometheus, Grafana, NetBox, etc.
  • Familiarity with AI and HPC tooling such as Slurm, NVIDIA DCGM, NCCL, Kubernetes, and distributed system validation frameworks.
  • Excellent problem-solving and analytical skills.
  • Ability to work independently and as part of a team.
  • Strong communication skills, both written and verbal.
  • Willingness to undertake international and/or domestic travel for on-site deployments and commissioning as required.
Location & Reporting

This role is based in Australia or Singapore with regular visits to current and future project sites in Australia and SE Asia.

Report to: Head of AI Cluster Commissioning

Employment Basis: Full-time

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infrastructure Automation Engineer, AI Cluster CommissioningNew
Infrastructure Automation Engineer, AI Cluster CommissioningNew

Firmus Technologies Pty Ltd. • City of Melbourne

On-site
AUD 120,000 - 180,000
Infrastructure Automation Engineer, AI Cluster Commissioning
Infrastructure Automation Engineer, AI Cluster Commissioning

Firmus • City of Melbourne

On-site
AUD 120,000 - 180,000
Infrastructure Automation Engineer, AI Cluster Commissioning
Infrastructure Automation Engineer, AI Cluster Commissioning

Engg • City of Melbourne

On-site
AUD 100,000 - 140,000
Storage Engineer, AI Cluster CommissioningNew
Storage Engineer, AI Cluster CommissioningNew

Firmus Technologies Pty Ltd. • City of Melbourne

On-site
AUD 120,000 - 180,000
Network Engineer, AI Cluster CommissioningNew
Network Engineer, AI Cluster CommissioningNew

Firmus Technologies Pty Ltd. • City of Melbourne

On-site
AUD 150,000 - 210,000
Site Reliability Engineer, AI Infrastructure
Site Reliability Engineer, AI Infrastructure

Firmus Technologies • City of Melbourne

On-site
AUD 120,000 - 180,000
Network Engineer, AI Cluster Commissioning
Network Engineer, AI Cluster Commissioning

Firmus • City of Melbourne

On-site
AUD 120,000 - 180,000
Senior Platform Reliability Engineer (Fabric and Interconnect)
Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies Pty Ltd. • Sydney

On-site
AUD 150,000 - 190,000
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Senior Platform Reliability Engineer (Fabric and Interconnect)
Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus • Sydney

On-site
AUD 140,000 - 190,000