Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies

Sydney

On-site

AUD 140,000 - 210,000

Full time

45 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Firmus Technologies in Australia is seeking a Senior Platform Reliability Engineer (Fabric and Interconnect) to own the reliability of GPU interconnect fabrics, NVLink/NVSwitch, InfiniBand and related networks across a large AI infrastructure.

This hands-on senior role emphasizes automation, software-defined remediation, and collaboration with vendor engineers to drive permanent fixes, runbooks, and resilient operations.

Qualifications

  • Extensive experience in high‑performance networking and production networks.
  • Deep knowledge of GPU interconnect fabrics (NVLink/NVSwitch) and topology.
  • Experience with InfiniBand, RoCE and related congestion control.
  • Experience with DPU/SmartNIC based host networking and firmware coordination.
  • Automation and infra‑as‑code practices in production environments.

Responsibilities

  • Ensure reliable operation and continuous improvement of GPU interconnect and network fabrics.
  • Develop guarded automation and remediation tooling for fabric faults.
  • Diagnose performance across the interconnect stack and tune networks.
  • Manage 24/7 incident response and on‑call escalation.
  • Lead vendor escalations and drive durable fixes and runbooks.

Skills

High-performance networking
Production network ownership
24/7 operations experience
Incident response
Automation and SRE practices
Scripting (Python/Go/Bash)
Vendor escalation
Technical communication

Tools

NVIDIA NVLink
NVSwitch
InfiniBand
RoCE
Spectrum-X
DPU/SmartNIC

Job description

Senior Platform Reliability Engineer (Fabric and Interconnect)

Firmus Technologies

Firmus Technologies is a global leader pioneering the development and operation of efficient AI infrastructure across Asia Pacific.

Founded in Australia in 2019, our mission is to create the most efficient AI infrastructure by combining cutting-edge technology with a steadfast commitment to sustainability.

At Firmus, we are unique in our approach. We design, build, and operate a new class of digital infrastructure – the AI Factory. Through our model-to-grid technology approach, we have pushed the boundaries of multi-generational liquid cooling systems, energy management, AI software orchestration, and construction. For our customers, this approach allows us to make every watt count and deliver low-cost AI tokens globally.

Firmus AI Cloud

Our large-scale GPU cloud platform, Firmus AI Cloud, is purpose-built to deliver energy-efficient AI compute at scale to customers.

It empowers developers, enterprises, educational institutions, and government users to train and deploy AI models with unmatched efficiency and cost savings. With an ever-growing suite of services and applications, we are committed to delivering a cloud experience that is market-leading, proprietary, and built to scale.

AI FactoryOS Operations

AI FactoryOS is Firmus' proprietary operating system for the AI Factory. It governs GPU telemetry, cooling, power and grid interaction as one integrated layer, so that every Firmus site can be optimised and monitored as a single system.

AI FactoryOS Operations runs that platform in production and owns the 24/7 reliability of AI FactoryOS, Firmus AI Cloud and the platforms built on them, together with the service levels the estate is measured against.

The remit is an engineering one. The function builds the guarded automation, remediation and operational tooling that turn manual response into a software-defined capability, and builds and operates the shared services the estate's own operation depends on. The function works closely with the engineering teams that build the platform, supplying the production evidence that shapes what they fix and what they build next.

Role Summary

Firmus runs large-scale, state-of-the-art AI infrastructure built on the latest generation of GPU rack-scale systems and operated as one estate to power the next generation of AI innovation. The Senior Platform Reliability Engineer, Fabric and Interconnect, owns the reliability of the fabrics this estate runs on: the GPU-to-GPU interconnect domains, the high-performance network fabrics carrying training and inference traffic, and the DPU-based host networking that binds compute to the rest of the platform.

This is a hands-on senior role with deep technical expertise. Automation is a first-class part of the role: the team builds and maintains the guarded automation and remediation tooling that turn manual fabric response into a self-healing capability, and the role engages fabric vendors at engineering level, reproducing faults to their standard and holding them to their answers.

Key Responsibilities

  • Responsible for the reliable operation, automation and continuous improvement of the estate's GPU interconnect and network fabrics (for example NVLink and NVSwitch domains, InfiniBand and Spectrum-X).
  • Build and maintain the guarded automation and remediation tooling for fabric faults, contributing to the software-driven remediation of AI clusters, including fault isolation and fabric reconvergence.
  • Diagnose and tune performance across the interconnect stack, from application collective communication down to link level, working with technologies including NVLink , InfiniBand, RoCE and congestion control tuning.
  • Operate DPU-based host networking across the fleet, including offload path configuration and driver and firmware compatibility.
  • Execute firmware upgrade waves, fabric expansions and capacity changes to the supported paths and scaling patterns defined by AI Infrastructure, owning the production window, the staged or canary path, verification against declared success criteria, and rollback execution.
  • Provide the deepest technical expertise for fabric and interconnect faults, correlating a collective communication failure to a specific physical link and diagnosing the faults that require internals-level knowledge to root cause, and driving the permanent fix to closure through AI Infrastructure .
  • Lead vendor escalations at engineering level, reproducing faults to the vendor's standard and pushing back credibly when a diagnosis does not explain the observed behaviour .
  • Lead technical recovery during major fabric incidents, drive the changes that remove repeat causes, share the follow-the-sun on-call roster, and mentor the engineers who carry frontline diagnosis, documenting operational procedures, runbooks and performance results.

Skills & Experience

  • Strong skills in high-performance networking and systems engineering, with 8+ years of experience including substantial ownership of production network or interconnect infrastructure in a 24/7 environment.
  • Deep operational experience with high-performance GPU interconnect fabrics (for example NVLink and NVSwitch ), including domain topology and failure diagnosis.
  • Extensive experience with high-performance networking fabrics (for example InfiniBand or RoCE-based Ethernet such as NVIDIA Spectrum-X), including routing internals and congestion control tuning.
  • Experience with DPU or SmartNIC ‑based host networking, including offload paths and driver and firmware coordination.
  • Experience planning and executing firmware upgrade waves and capacity expansions on production fabrics with defined rollback.
  • Strong skills in infrastructure automation and infrastructure‑as‑code practices, with change delivered through peer review and progressive rollout.
  • Practical experience with scripting or programming for operational automation and tooling, such as Python, Go or Bash.
  • Proven ability to act as a senior escalation point in production, including major incident response, on-call participation, vendor escalation at engineering level, and the production of runbooks that others can execute successfully.
  • Clear technical judgement and communication skills, with the ability to explain complex fabric failures to engineers and non-specialists.

Preferred Experience

  • Experience operating fabrics for large-scale distributed training or inference workloads.
  • Experience with NVIDIA rack-scale or multi-node GPU systems and their interconnect topology.
  • Experience in a multi-tenant service provider, cloud or colocation environment.
  • Knowledge of data centre and hardware fundamentals, including cabling and optics, firmware management and hardware fault workflows.

Location & Reporting

Location : Based in Australia or Singapore, with travel to Australian AI Factory sites as required.

On-call : The function runs 24/7. First line monitoring and first response sit with the operations centre. This role shares the after-hours escalation roster for its domain with the other senior engineers in the function.

Reporting to : Reports to the Head of AI FactoryOS Operations while the function is being established , working under broad direction with a high degree of autonomy and direct access to the decision makers. As the function reaches its planned structure, the role will report to the Infrastructure Operations Manager, with the Head of AI FactoryOS Operations remaining accountable for the function. The scope, level and remit of the role do not change under either arrangement.

Employment Basis

Permanent full-time

Diversity

At Firmus, we are committed to building a diverse and inclusive workplace. We encourage applications from candidates of all backgrounds who are passionate about creating a more sustainable future through innovative engineering solutions.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Firmus Technologies • City of Melbourne

On-site
AUD 180,000 - 260,000
Service Delivery Manager
Service Delivery Manager

Firmus Technologies • Sydney

On-site
AUD 140,000 - 180,000
Senior Platform Reliability Engineer
Senior Platform Reliability Engineer

Matchbox • City of Melbourne

Hybrid
AUD 180,000 - 260,000
Service Delivery ManagerNew
Service Delivery ManagerNew

Firmus Technologies • Sydney

Hybrid
AUD 140,000 - 190,000
Principal Network Architect, Backbone Network
Principal Network Architect, Backbone Network

Firmus Technologies • Council of the City of Sydney

On-site
AUD 180,000 - 280,000
Senior Kubernetes Platform Engineer
Senior Kubernetes Platform Engineer

Firmus Technologies • Sydney

On-site
AUD 180,000 - 240,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

Firmus Technologies • Sydney

On-site
AUD 150,000 - 210,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

Firmus • Sydney

On-site
AUD 130,000 - 180,000
Senior Security Engineer, Platform Engineering
Senior Security Engineer, Platform Engineering

re-zoo-me • Sydney

Hybrid
AUD 180,000 - 240,000
Site Reliability Engineer, AI Infrastructure
Site Reliability Engineer, AI Infrastructure

Firmus Technologies • City of Melbourne

On-site
AUD 120,000 - 180,000