Field Solution Engineer

N2S.Global

Sydney

On-site

AUD 120,000 - 180,000

Full time

10 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

N2S.Global is seeking an on-site Field Solutions Engineer in Australia to physically install and validate DGX/HGX/MGX nodes, cabling, and power systems. You will run bare-metal provisioning, configure networking and storage clients, and perform rigorous health checks before AI factory activation.

You will diagnose GPU, network, and storage faults on-site, coordinate RMA with OEMs, and provide timely escalation support to offshore teams. Prior data-centre technician experience is preferred.

Qualifications

  • Hands-on field engineering in on-site data centers in Australia/New Zealand.
  • Experience with bare-metal provisioning, networking and storage for HPC/AI hardware.
  • Familiarity with vendor coordination for RMA/DOA and on-site troubleshooting.

Responsibilities

  • physically install and verify DGX/HGX/MGX nodes and associated cabling.
  • execute provisioning workflows and configure host networking and storage clients.
  • perform fault isolation and coordinate RMA with vendors.
  • support L2 escalations with on-site diagnostics and remediation.

Skills

Rack and stack
NVIDIA BCM
Ansible
HPC/AI infrastructure
Linux system admin
Network cabling
NVIDIA GPUs

Education

RHCSA
NVIDIA-Certified Associate

Tools

IPMI tools
Git

Job description

The Field Solutions Engineer is the onshore hands-on execution engine that closes the gap between the offshore engineering squad and the physical reality of an Australian and New Zealand data centre floor. While the Domain Architects (Compute, Network, Storage) design the "Gold Standard" from the regional hub and the offshore HPC Engineer squad executes remote configuration and automation from India, you are the hands that cannot be replaced by a remote session: pulling a faulty transceiver, walking a rack elevation against the LLD, re-seating a cable, or supervising a burn-in test in person.

As a System Integrator, we do not simply manage a static cloud; we design and deliver bespoke, high-scale AI factories for the world's leading enterprises. In this role, you sit inside the AI Infrastructure team and work across NVIDIA SuperPOD, BasePOD, and Cisco AI Factory deployments as a generalist across the Compute-Network-Storage triad, rather than as a single-domain specialist, and you are the primary point of RMA/DOA diagnosis and remediation on the ground.

You operate with a 100% focus on Delivery, executing across Low-Level Designs (LLDs) assigned by whichever Domain Architect owns the active engagement, and providing Layer 1 QA support and OOB (out-of-band) standup ahead of AI Factory commissioning.

Key Responsibilities
1. Physical Infrastructure & Layer 1 Execution
  • Rack & Stack Execution:
  • Physically install and verify DGX/HGX/MGX nodes, switches, and PDUs against the current rack elevation and LLD.
  • Confirm floor-loading, bolting, and levelling before energisation; elevate any structural discrepancy to the Domain Architect – AI Facilities.
  • Structured Cabling & P2P Validation:
  • Execute the point-to-point (P2P) cabling schedule; confirm transceiver type and MPO cable size against the code on the box, not the colour, before patching.
  • Clean and inspect optical connectors on every patch; validate seating and troubleshoot link-down, miswire, and link-flap faults by elimination (reseat, swap to a known-good port, clean, replace).
  • Bring up and validate the out-of-band (BMC/IPMI) management network ahead of in-band and compute-fabric activation, keeping it physically segregated per design.
  • Confirm node power state and basic health via BMC before handing off to HPC configuration.
2. HPC Compute, Network & Storage Configuration
  • Bare-Metal & Fabric Provisioning:
  • Execute NVIDIA Base Command Manager (BCM) provisioning workflows and Ansible playbooks supplied by the Domain Architects to bring compute nodes, switches, and storage clients into service.
  • Configure host-side networking (IPoIB, Netplan) and mount high-performance storage clients (VAST, WEKA, Lustre) to the current LLD.
  • Firmware & Driver Lifecycle:
  • Execute SBIOS, BMC, GPU VBIOS, and NVSwitch firmware upgrades per the NVIDIA firmware recipe across compute, network, and storage tiers.
  • Apply OS hardening, kernel patching, and driver updates during scheduled maintenance windows.
  • Performance Validation:
  • Execute and log HPL, NCCL-tests, ib_write_bw/ib_send_bw, and IOR/FIO benchmark suites; compare results against the Gold Standard and flag deviations to the relevant Domain Architect.
3. RMA/DOA Diagnosis & Operations Support
  • Fault Isolation:
  • Triage Dead-on-Arrival hardware, Xid errors, flapping links, and stale storage mounts, isolating whether the fault sits in compute, network, or storage before escalating.
  • Distinguish a "Node Issue" from a "Network Issue" or a "Storage Issue" using nvidia-smi/dmesg, ibstat/ibdiagnet, and iostat/iotop.
  • RMA Coordination:
  • Raise and track Return Merchandise Authorisation cases with OEM and vendor support; coordinate physical swap-out with minimal schedule impact.
  • L2 Ticket Resolution:
  • Handle on-site escalations passed from L1 support; close the loop with the offshore HPC Engineer squad on root cause and remediation.
  • Data Centre Background: Prior Data Centre Technician, Field Engineer, or rack-and-stack experience within a System Integrator, OEM, or colocation environment.
  • Fault Diagnosis:
  • Comfortable triaging GPU faults (nvidia-smi, dmesg), link faults (ibstat, ibdiagnet, ethtool), and storage mount issues (iostat, iotop) to isolate a fault domain before escalating.
  • Physical Infrastructure (Layer 0/1):
  • Rack-and-stack procedures, structured cabling (OS2/OM4/DAC/AOC), MPO and OSFP transceiver handling, and cable pathway standards.
  • Site Acceptance Test (SAT) support and As-Built documentation capture.
  • Solid RHEL/Ubuntu administration; ability to execute and troubleshoot Ansible playbooks.
  • Git workflow familiarity (pulling code, branching, committing configuration changes).
  • HPC & AI Infrastructure:
  • Working proficiency with NVIDIA Base Command Manager (BCM) for bare-metal provisioning.
  • Familiarity with DGX/HGX/MGX hardware architecture and standard benchmark suites (HPL, NCCL-tests, IOR/FIO).
  • Network Affinity: Hands-on exposure to InfiniBand/RoCEv2 cabling and switch-side transceiver handling.
  • Storage Integration: Parallel filesystem client mounting experience (VAST, WEKA, Lustre).
Certifications
Highly Desirable:
  • NVIDIA-Certified Associate: AI Infrastructure and Operations (NCA-AIIO)
  • Red Hat Certified System Administrator (RHCSA)
  • RMA/DOA Turnaround: Time from fault identification to RMA lodgement and physical swap-out completion.
  • Layer 1 Accuracy: Zero unresolved miswires or "ghost links" handed to the validation/stress-test phase; cabling matches the current P2P schedule at handover.
  • Benchmark Pass Rate: Percentage of nodes and fabric segments passing HPL/NCCL/IOR-FIO validation on the first attempt.
  • Ticket Efficiency: Consistently meeting SLAs for L2 escalations requiring on-site hands.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Field Solution Engineer or Data Centre Technician
Field Solution Engineer or Data Centre Technician

World Wide Technology • Council of the City of Sydney

On-site
AUD 140,000 - 210,000
Domain Architect - AI Network
Domain Architect - AI Network

World Wide Technology • Council of the City of Sydney

On-site
AUD 180,000 - 280,000
Domain Architect AI Compute
Domain Architect AI Compute

World Wide Technology • City of Melbourne

On-site
AUD 260,000 - 380,000
AI Regional Supervisor
AI Regional Supervisor

World Wide Technology • Council of the City of Sydney

On-site
AUD 180,000 - 240,000
WHS site safety accreditation
NVIDIA AI Infrastructure certification
Structured cabling certification
AI Installation Supervisor
AI Installation Supervisor

Auxo Talent • Sydney

On-site
AUD 120,000 - 155,000
AI Data center Supervisor
AI Data center Supervisor

World Wide Technology • Sydney, City of Melbourne, Perth, City of Brisbane

On-site
AUD 150,000 - 210,000
Ai Data Center Supervisor
Ai Data Center Supervisor

World Wide Technology • Sydney, City of Melbourne, City of Brisbane

On-site
AUD 120,000 - 180,000
Solutions Architect - DevOps
Solutions Architect - DevOps

NVIDIA Pty. Ltd • Sydney, City of Melbourne

On-site
AUD 180,000 - 240,000
Senior Project Delivery Manager, NVIS
Senior Project Delivery Manager, NVIS

NVIDIA • Council of the City of Sydney

On-site
AUD 180,000 - 240,000
Senior Solution Architect, AI Compute Engineer - NVIS
Senior Solution Architect, AI Compute Engineer - NVIS

NVIDIA • Sydney

On-site
AUD 120,000 - 160,000