Field Solution Engineer

Nityo Infotech

City of Melbourne

On-site

AUD 90,000 - 120,000

Full time

41 hours ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Nityo Infotech in Melbourne seeks a Data Centre Technician to support HPC and AI infrastructure across rack-and-stack, cabling, and out-of-band management. You will provision bare-metal nodes, hardware upgrades, and storage clients, while validating performance with standard benchmarks.

Responsibilities include fault triage for GPU, network, and storage, RMA coordination, and L2 support with offshore teams. Linux admin skills and Ansible experience are preferred.

Qualifications

  • Data Centre background with experience in rack-and-stack and field engineering in SI/OEM/colo environments.

Responsibilities

  • Rack-and-stack execution and verification of DGX/HGX/MGX nodes, switches, and PDUs.
  • Structured cabling and P2P validation; verify transceiver types and MPO sizes.
  • Bring up and validate out-of-band management network (BMC/IPMI) before in-band compute-fabric activation.
  • Bare-metal and fabric provisioning using BCM provisioning workflows and Ansible playbooks.
  • Configure host networking (IPoIB, Netplan) and mount high-performance storage clients (VAST, WEKA, Lustre).
  • Perform firmware and driver upgrades (SBIOS, BMC, GPU VBIOS, NVSwitch) per NVIDIA recipes.
  • Execute performance benchmarks (HPL, NCCL-tests, ib_write_bw/ib_send_bw, IOR/FIO) and compare to Gold Standard.
  • Triage hardware issues (GPU, links, storage mounts) and distinguish node vs network vs storage faults.
  • Coordinate RMA with OEM/vendor and track DOA cases; resolve L2 escalations with offshore HPC team.
  • Solid HPC/AI infra exposure and familiarity with NVIDIA BCM and DGX/HGX/MGX architectures.

Skills

Data centre background
Fault diagnosis
HPC & AI infrastructure
NVIDIA BCM familiarity
GPU/InfiniBand networking
Storage integration
RHEL/Ubuntu administration
Ansible
Git workflow

Tools

Ansible
Git
ibstat
ibdiagnet
iostat
iotop
nvidia-smi
dmesg
ethtool

Job description

Physical Infrastructure & Layer 1 Execution



  • Rack & Stack Execution:

  • Physically install and verify DGX/HGX/MGX nodes, switches, and PDUs against the current rack elevation and LLD.

  • Confirm floor-loading, bolting, and levelling before energisation; elevate any structural discrepancy to the Domain Architect AI Facilities.

  • Structured Cabling & P2P Validation:

  • Execute the point-to-point (P2P) cabling schedule; confirm transceiver type and MPO cable size against the code on the box, not the colour, before patching.

  • Clean and inspect optical connectors on every patch; validate seating and troubleshoot link-down, miswire, and link-flap faults by elimination (reseat, swap to a known-good port, clean, replace).

  • Bring up and validate the out-of-band (BMC/IPMI) management network ahead of in-band and compute-fabric activation, keeping it physically segregated per design.

  • Confirm node power state and basic health via BMC before handing off to HPC configuration. 2. HPC Compute, Network & Storage Configuration

  • Bare-Metal & Fabric Provisioning:

  • Execute NVIDIA Base Command Manager (BCM) provisioning workflows and Ansible playbooks supplied by the Domain Architects to bring compute nodes, switches, and storage clients into service.

  • Configure host-side networking (IPoIB, Netplan) and mount high-performance storage clients (VAST, WEKA, Lustre) to the current LLD.

  • Firmware & Driver Lifecycle:

  • Execute SBIOS, BMC, GPU VBIOS, and NVSwitch firmware upgrades per the NVIDIA firmware recipe across compute, network, and storage tiers.

  • Apply OS hardening, kernel patching, and driver updates during scheduled maintenance windows.

  • Performance Validation:

  • Execute and log HPL, NCCL-tests, ib_write_bw/ib_send_bw, and IOR/FIO benchmark suites; compare results against the Gold Standard and flag deviations to the relevant Domain Architect. 3. RMA/DOA Diagnosis & Operations Support

  • Fault Isolation:

  • Triage Dead-on-Arrival hardware, Xid errors, flapping links, and stale storage mounts, isolating whether the fault sits in compute, network, or storage before escalating.

  • Distinguish a \"Node Issue\" from a \"Network Issue\" or a \"Storage Issue\" using nvidia-smi/dmesg, ibstat/ibdiagnet, and iostat/iotop.

  • RMA Coordination:

  • Raise and track Return Merchandise Authorisation cases with OEM and vendor support; coordinate physical swap-out with minimal schedule impact.

  • L2 Ticket Resolution:

  • Handle on-site escalations passed from L1 support; close the loop with the offshore HPC Engineer squad on root cause and remediation. Technical CompetenciesEssential Skills

  • Data Centre Background: Prior Data Centre Technician, Field Engineer, or rack-and-stack experience within a System Integrator, OEM, or colocation environment.

  • Fault Diagnosis:

  • Comfortable triaging GPU faults (nvidia-smi, dmesg), link faults (ibstat, ibdiagnet, ethtool), and storage mount issues (iostat, iotop) to isolate a fault domain before escalating.

  • Physical Infrastructure (Layer 0/1):

  • Rack-and-stack procedures, structured cabling (OS2/OM4/DAC/AOC), MPO and OSFP transceiver handling, and cable pathway standards.

  • Site Acceptance Test (SAT) support and As-Built documentation capture. Desirable Experience

  • Solid RHEL/Ubuntu administration; ability to execute and troubleshoot Ansible playbooks.

  • Git workflow familiarity (pulling code, branching, committing configuration changes).

  • HPC & AI Infrastructure:

  • Working proficiency with NVIDIA Base Command Manager (BCM) for bare-metal provisioning.

  • Familiarity with DGX/HGX/MGX hardware architecture and standard benchmark suites (HPL, NCCL-tests, IOR/FIO).

  • Network Affinity: Hands-on exposure to InfiniBand/RoCEv2 cabling and switch-side transceiver handling.

  • Storage Integration: Parallel filesystem client mounting experience (VAST, WEKA, Lustre).

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Centre Field Solution Engineer or Data Centre Field Engineer/Technician
Data Centre Field Solution Engineer or Data Centre Field Engineer/Technician

World Wide Technology • City of Melbourne

Hybrid
AUD 120,000 - 160,000
Travel to data centres
Hybrid work model
Exposure to AI infrastructure
Field Solution Engineer or Data Centre Technician
Field Solution Engineer or Data Centre Technician

World Wide Technology • Sydney

On-site
AUD 140,000 - 210,000
Domain Architect - AI Network
Domain Architect - AI Network

World Wide Technology • Council of the City of Sydney

On-site
AUD 180,000 - 280,000
Senior High-Performance Storage Architect - NVIS
Senior High-Performance Storage Architect - NVIS

NVIDIA • Sydney

On-site
AUD 251,000 - 363,000
AI Domain Architect (AI Storage)
AI Domain Architect (AI Storage)

World Wide Technology • City of Melbourne

On-site
AUD 180,000 - 240,000
Senior Lead Field Deployment Engineer – AI Cluster Infrastructure
Senior Lead Field Deployment Engineer – AI Cluster Infrastructure

World Wide Technology • Sydney

On-site
AUD 180,000 - 240,000
AI Regional Supervisor
AI Regional Supervisor

World Wide Technology • Sydney

On-site
AUD 180,000 - 240,000
WHS site safety accreditation
NVIDIA AI Infrastructure certification
Structured cabling certification
Network Engineer - Compute / HPC
Network Engineer - Compute / HPC

Pathway Search • Sydney

On-site
AUD 120,000 - 190,000
Linux HPC Operations Engineer
Linux HPC Operations Engineer

Westbury Partners • Sydney

On-site
AUD 120,000 - 170,000
Data Center Engineer
Data Center Engineer

ScaleUp Technologies • Western Australia

On-site
AUD 70,000 - 100,000