Domain Architect AI Compute

World Wide Technology

City of Melbourne

On-site

AUD 260,000 - 380,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

World Wide Technology is seeking a Domain Architect - AI Compute to own the lifecycle of high-performance GPU compute fleets across diverse client environments. The role blends complex AI infrastructure delivery (60%) with pre-sales SME responsibilities (40%), ensuring Day 2 readiness for NVIDIA/NVL2 deployments and guiding scope/cost estimates for future projects.

This position requires deep NVIDIA GPU platform knowledge, hands-on provisioning of DGX/BASEPOD/NVL72, Linux optimization, and

Qualifications

  • Deep architectural understanding of NVIDIA GPU platforms (Hopper, Grace-Hopper, Blackwell, Grace-Blackwell).
  • Mastery of NVL72 rack-scale integration and NVSwitch fabrics.
  • Expertise in CUDA, cuDNN, NCCL.
  • Linux distributions optimised for HPC/AI (Ubuntu, RHEL).
  • Kernel tuning, driver management and system hardening.
  • Proficiency in Python and Ansible for automation.
  • Experience in SI/MSP environments (desirable).
  • Hands-on experience with NVIDIA Base Command Manager (BCM) and NVIDIA Mission Control (desirable).
  • Knowledge of high-speed interconnects (InfiniBand NDR/HDR, RoCEv2).
  • Kubernetes/OpenShift installation and administration (desirable).

Responsibilities

  • Lead the physical provisioning of GPU compute clusters (NVIDIA NVL72, DGX SuperPOD, BasePOD, HGX, MGX, Cisco AI Factory).
  • Utilise BCM for diskless booting, firmware management, and OS hardening.
  • Establish monitoring with NVIDIA Mission Control.
  • Execute Zero Touch Provisioning (ZTP) workflows to production-ready nodes.
  • Define and enforce Fair Share policies and MIG-based quotas in multi-tenant environments.
  • Deploy and configure management planes for multi-cluster observability.
  • Implement telemetry with DCGM to monitor health and XID errors.
  • Assist sales by validating requirements and producing Labour Estimates for SOWs.

Skills

GPU architecture
NVSwitch
Linux HPC
Python
Ansible
System Integration

Tools

CUDA
cuDNN
NCCL
NVIDIA BCM
NVIDIA Mission Control

Job description

The Domain Architect - AI Compute acts as the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across diverse client environments, bridging the gap between architectural design and hands-on execution. Operating with a 60/40 split between delivering complex AI infrastructure (60%) and providing Pre-Sales Subject Matter Expertise (40%), you will lead the physical provisioning of NVIDIA SuperPOD, NVIDIA BasePOD, and Cisco AI Factory environments, ensuring clients receive "Day 2" ready AI factories, while assisting the sales team in defining the scope and cost of future deployments.

Key responsibilities

Lead the physical provisioning of clusters: NVIDIA NVL72, DGX SuperPOD, BasePOD, HGX, MGX, Cisco AI Factory

Utilise NVIDIA Base Command Manager (BCM) for diskless booting, firmware management, and OS hardening

Establish monitoring with NVIDIA Mission Control

Execute automated "Zero Touch Provisioning" (ZTP) workflows to transform bare-metal hardware into production-ready nodes

Define and enforce "Fair Share" policies, fractional GPU quotas using Multi-Instance GPU (MIG), and pre-emption logic for multi-tenant environments

Deploy and configure management planes like Rafay or Armada to enable multi-cluster management and observability

Implement high-fidelity telemetry using DCGM (Data Centre GPU Manager) to monitor GPU health, thermal throttling, and XID error rates

Conduct validation testing using NCCL-tests, HPL, and HPCG to verify cluster performance

Assist the sales team by validating customer technical requirements and producing accurate Labour Estimates (LOE) for Statements of Work (SOWs)

About you

Deep architectural understanding of NVIDIA GPU platforms (Hopper, Grace-Hopper, Blackwell, Grace-Blackwell)

Mastery of NVL72 rack-scale integration and NVSwitch fabrics

Expertise in the associated software stack (CUDA, cuDNN, NCCL)

Expert-level knowledge of Linux distributions (Ubuntu, RHEL) optimised for HPC/AI

Deep experience with kernel tuning, driver management, and system hardening

Proficiency in Python and Ansible for hardware configuration management and automation

Experience working within a System Integrator (SI) or Managed Service Provider (MSP) environment (desirable)

Hands-on experience with NVIDIA Base Command Manager (BCM) and/or NVIDIA Mission Control (desirable)

Solid understanding of high-speed interconnects (InfiniBand NDR/HDR, RoCEv2) and how they interface with host PCIe/NVLink topologies (desirable)

Experience with Kubernetes/Red Hat OpenShift installation and administration (desirable)

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Field Solution Engineer
Field Solution Engineer

N2S.Global • Sydney

On-site
AUD 120,000 - 180,000
Domain Architect - AI Network
Domain Architect - AI Network

World Wide Technology • Council of the City of Sydney

On-site
AUD 180,000 - 280,000
Field Solution Engineer or Data Centre Technician
Field Solution Engineer or Data Centre Technician

World Wide Technology • Council of the City of Sydney

On-site
AUD 140,000 - 210,000
Solutions Architect - DevOps
Solutions Architect - DevOps

NVIDIA Pty. Ltd • Sydney, City of Melbourne

On-site
AUD 180,000 - 240,000
Senior Solution Architect, AI Compute Engineer - NVIS
Senior Solution Architect, AI Compute Engineer - NVIS

NVIDIA • Sydney

On-site
AUD 120,000 - 160,000
AI Compute Domain Architect: Multi-Tenant GPU Infra Lead
AI Compute Domain Architect: Multi-Tenant GPU Infra Lead

World Wide Technology • City of Melbourne

On-site
AUD 260,000 - 380,000
Solutions Architect, Networking Ethernet
Solutions Architect, Networking Ethernet

NVIDIA • Sydney

On-site
AUD 120,000 - 160,000
AI Installation Supervisor
AI Installation Supervisor

Auxo Talent • Sydney

On-site
AUD 120,000 - 155,000
AI Compute Domain Architect: High-Performance GPU Infra
AI Compute Domain Architect: High-Performance GPU Infra

World Wide Technology • Australia

On-site
AUD 180,000 - 240,000
Senior AI & HPC Infrastructure Delivery Lead
Senior AI & HPC Infrastructure Delivery Lead

NVIDIA • Council of the City of Sydney

On-site
AUD 180,000 - 240,000