Sr. AI Validation Engineer

Mainz Brady Group

Portland (OR)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Mainz Brady Group is seeking a Senior AI Infrastructure Validation Engineer to validate the readiness, reliability, and performance of large-scale AI infrastructure environments before they reach production.

This role intersects systems engineering, networking, GPU infrastructure, container platforms, and distributed AI workloads, designing validation strategies that uncover failures early and accelerate root-cause analysis.

Qualifications

  • 7+ years of experience in infrastructure validation, systems validation, performance engineering, infrastructure QA, HPC, or AI infrastructure certification.
  • Strong hands-on troubleshooting experience across: Linux systems, GPU infrastructure, Network fabrics, Containers, Distributed workloads.
  • Demonstrated experience designing validation strategies, frameworks, and qualification approaches, rather than only executing predefined test cases.
  • Understanding of AI infrastructure dependencies including: NCCL, RDMA, Storage throughput, Multi-node communication, Distributed training behavior, Cluster orchestration.
  • Ability to distinguish infrastructure defects from application, framework, workload, and configuration issues.
  • Strong scripting and automation skills using Python, Bash, or similar technologies.
  • Experience automating test execution, evidence collection, analysis, and reporting.
  • Strong root-cause analysis and systems troubleshooting skills.
  • Excellent written communication skills with experience producing detailed defect reports, readiness assessments, and technical recommendations.

Responsibilities

  • Design and execute comprehensive validation plans for AI infrastructure spanning Compute nodes, GPU communication, Network fabric health, Storage access, Container orchestration, Distributed workload readiness.
  • Perform structured bring-up, soak, regression, qualification, and certification testing for new and modified AI cluster environments.
  • Reproduce, troubleshoot, and isolate failures involving: Distributed training workloads, GPU/node instability, NCCL and communication libraries, Kubernetes and container platforms, Storage paths and throughput, Ethernet and InfiniBand transport behavior.
  • Validate Ethernet and InfiniBand environments, including host readiness, RDMA connectivity, and end-to-end workload behavior.
  • Correlate failures across system logs, telemetry, firmware state, infrastructure health, and workload symptoms to accelerate root-cause analysis.
  • Determine whether failures originate from infrastructure, hardware, networking, software frameworks, workloads, or configuration.
  • Partner with Linux, networking, deployment, platform, and infrastructure engineering teams to close validation gaps before production handoff.
  • Define defect signatures, validation criteria, pass/fail thresholds, readiness assessments, and release recommendations.
  • Build and improve automation for: Cluster certification, Health scoring, Regression testing, Evidence collection, Post-change validation, Readiness reporting.
  • Improve repeatability and scalability of validation processes across large AI infrastructure environments.

Skills

Linux systems
GPU infrastructure
Network fabrics
Containers
Distributed workloads
Python
Bash
Automation
Root-cause analysis
Technical documentation

Tools

Kubernetes
Containers
NCCL
RDMA
InfiniBand
Ethernet
Python
Bash
Telemetry Platforms
Cluster Certification Frameworks

Job description

We are seeking a Senior AI Infrastructure Validation Engineer to validate the readiness, reliability, and performance of large-scale AI infrastructure environments before they reach production. This role sits at the intersection of systems engineering, networking, GPU infrastructure, container platforms, and distributed AI workloads. The focus is not simply executing predefined test cases, but designing validation strategies that uncover failure modes early, accelerate root-cause isolation, and provide clear evidence that clusters are ready for production workloads. The ideal candidate approaches infrastructure validation as an engineering discipline, with the ability to understand how compute, networking, storage, orchestration, firmware, and distributed workloads interact across complex AI environments.

Responsibilities
  • Design and execute comprehensive validation plans for AI infrastructure spanning:
    • Compute nodes
    • GPU communication
    • Network fabric health
    • Storage access
    • Container orchestration
    • Distributed workload readiness
  • Perform structured bring-up, soak, regression, qualification, and certification testing for new and modified AI cluster environments.
  • Reproduce, troubleshoot, and isolate failures involving:
    • Distributed training workloads
    • GPU/node instability
    • NCCL and communication libraries
    • Kubernetes and container platforms
    • Storage paths and throughput
    • Ethernet and InfiniBand transport behavior
  • Validate Ethernet and InfiniBand environments, including host readiness, RDMA connectivity, and end-to-end workload behavior.
  • Correlate failures across system logs, telemetry, firmware state, infrastructure health, and workload symptoms to accelerate root-cause analysis.
  • Determine whether failures originate from infrastructure, hardware, networking, software frameworks, workloads, or configuration.
  • Partner with Linux, networking, deployment, platform, and infrastructure engineering teams to close validation gaps before production handoff.
  • Define defect signatures, validation criteria, pass/fail thresholds, readiness assessments, and release recommendations.
  • Build and improve automation for:
    • Cluster certification
    • Health scoring
    • Regression testing
    • Evidence collection
    • Post-change validation
    • Readiness reporting
  • Improve repeatability and scalability of validation processes across large AI infrastructure environments.
Required Qualifications
  • 7+ years of experience in infrastructure validation, systems validation, performance engineering, infrastructure QA, HPC, or AI infrastructure certification.
  • Strong hands-on troubleshooting experience across:
    • Linux systems
    • GPU infrastructure
    • Network fabrics
    • Containers
    • Distributed workloads
  • Demonstrated experience designing validation strategies, frameworks, and qualification approaches, rather than only executing predefined test cases.
  • Understanding of AI infrastructure dependencies including:
    • NCCL
    • RDMA
    • Storage throughput
    • Multi-node communication
    • Distributed training behavior
    • Cluster orchestration
  • Ability to distinguish infrastructure defects from application, framework, workload, and configuration issues.
  • Strong scripting and automation skills using Python, Bash, or similar technologies.
  • Experience automating test execution, evidence collection, analysis, and reporting.
  • Strong root-cause analysis and systems troubleshooting skills.
  • Excellent written communication skills with experience producing detailed defect reports, readiness assessments, and technical recommendations.
Preferred Qualifications
  • Experience validating GPU clusters, large-scale AI training environments, HPC systems, or AI infrastructure prior to production deployment.
  • Experience with burn-in, soak testing, telemetry analysis, and cluster qualification.
  • Familiarity with hardware, firmware, driver, and software compatibility testing.
  • Experience building validation suites that support both:
    • Pre-production deployment readiness
    • Ongoing production health and regression validation
  • Experience supporting high-density GPU or distributed computing environments.
Tools & Technologies
Infrastructure & Platforms
  • Linux
  • GPU Infrastructure
  • Kubernetes
  • Containers
  • Distributed Training Systems
Networking
  • Ethernet
  • InfiniBand
  • RDMA
  • NCCL
Automation
  • Python
  • Bash
  • Automation Frameworks
  • Cluster Certification Frameworks
Validation & Observability
  • Telemetry Platforms
  • Performance Analysis Tools
  • Infrastructure Validation Tools
  • Burn-In / Soak Testing
  • Health Monitoring
  • Regression & Qualification Testing
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Socket.dev • Houston (TX)

On-site
USD 200,000 - 260,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Infrastructure Validation Engineer
Infrastructure Validation Engineer

Blue Signal Search • Phoenix (AZ)

On-site
USD 120,000 - 160,000
Infrastructure Validation Engineer
Infrastructure Validation Engineer

Blue Signal Search • Austin (TX)

On-site
USD 90,000 - 130,000
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
Senior AI Infrastructure Validation Architect
Senior AI Infrastructure Validation Architect

Mainz Brady Group • Portland (OR)

On-site
USD 150,000 - 190,000
Principal AI Cluster Performance & Validation Architect
Principal AI Cluster Performance & Validation Architect

Socket.dev • Houston (TX)

On-site
USD 200,000 - 260,000
Validation Engineer - Team Lead
Validation Engineer - Team Lead

Cloud Destinations • San Francisco (CA)

On-site