GPU Cluster Architect

Jobgether SRL

United States

Remote

USD 184,000 - 318,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision insurance
Remote work reimbursement
RSUs may be available
401(k) with company match
Parental leave
Career growth opportunities

Job summary

Jobgether SRL, on behalf of a partner, seeks a GPU Cluster Architect based in the United States. This is a remote, high-impact role focused on designing next-generation AI infrastructure at massive scale, making end-to-end architectural decisions across GPU compute, networking, storage, reliability, and control planes.

You will model workloads for LLM training and inference, collaborate with site reliability, networking, storage, and data center teams, and drive automation and telemetry

Qualifications

  • 5+ years designing and architecting large-scale GPU clusters.
  • Deep understanding of GPU architectures (NVIDIA/AMD).
  • Experience with InfiniBand and RoCE networking.
  • Ability to analyze workloads for latency and bandwidth.
  • Experience with automation and telemetry using Python or Go.

Responsibilities

  • Architect scalable GPU cluster topologies across compute, storage, and control planes.
  • Define architectures to support AI/ML workloads across data centers.
  • Model workload requirements for LLM training and inference.
  • Design high-throughput, low-latency networking at POD and data-center scale.
  • Collaborate with data center, networking, and storage teams to deploy architectures.
  • Contribute to automation and telemetry initiatives for visibility and reliability.

Skills

GPU cluster design
Networking
InfiniBand/RoCE
Automation scripting (Python/Go)

Tools

InfiniBand HDR/NDR

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a GPU Cluster Architect based in the United States.

This is a remote, high-impact architecture role focused on designing next-generation AI infrastructure at massive scale.

You will make end-to-end architectural decisions spanning GPU compute, high-performance networking, storage, reliability, and control planes.

Your work will shape how tens of thousands of GPUs are interconnected, powered, cooled, monitored, and optimized across multiple data center sites.

You will model demanding AI and machine-learning workloads, including large language model training and inference, to guide critical performance and infrastructure tradeoffs.

The role combines deep systems expertise with hands-on collaboration across networking, storage, site reliability, and data center engineering teams.

You will operate in a fast-moving, engineering-led environment where scalability, performance, reliability, and innovation are central to the work.

This is an opportunity to influence the architecture of large-scale AI infrastructure while solving complex distributed systems challenges.

Accountabilities
  • Architect scalable GPU cluster topologies encompassing compute nodes, high-performance interconnects, storage systems, and control planes.
  • Define and evaluate infrastructure architectures capable of supporting large-scale AI and machine-learning workloads across multiple data center sites.
  • Model workload requirements for applications such as large language model training and inference, using latency, bandwidth, GPU density, and other performance factors to guide architectural decisions.
  • Design and validate high-throughput, low-latency networking architectures at both POD and data-center scale, including InfiniBand and Ethernet-based environments.
  • Work with network architecture teams to evaluate and validate technologies such as InfiniBand HDR/NDR and RoCEv2.
  • Partner with storage engineering teams to optimize infrastructure for training datasets, checkpointing, and other demanding AI workloads.
  • Analyze monitoring and telemetry signals to identify design issues, reliability risks, and opportunities for architectural improvement.
  • Collaborate closely with site reliability, networking, storage, and data center engineering teams to operationalize, deploy, and scale infrastructure architectures.
  • Contribute to automation and telemetry initiatives that improve the visibility, performance, and reliability of large-scale GPU environments.
  • Make end-to-end architectural decisions that balance scalability, performance, reliability, operational complexity, and infrastructure efficiency.
Requirements
  • 5+ years of experience designing and architecting large-scale computing or GPU clusters.
  • Deep understanding of modern GPU architectures, including NVIDIA, AMD, or comparable platforms.
  • Strong experience with high-performance computing interconnects, particularly InfiniBand and RoCE.
  • Solid background in systems architecture, networking, hardware infrastructure, and hardware reliability.
  • Understanding of GPU cluster design principles, including compute topology, network architecture, storage integration, and control-plane considerations.
  • Experience evaluating infrastructure performance and making architecture decisions based on workload characteristics such as latency, bandwidth, and compute density.
  • Experience with scripting or software development for automation, telemetry, monitoring, or infrastructure tooling using languages such as Python or Go.
  • Ability to analyze technical signals and operational data to identify infrastructure issues and inform design improvements.
  • Strong cross-functional collaboration skills, with the ability to work effectively with networking, storage, site reliability, and data center engineering teams.
  • Strong analytical and problem-solving abilities, with a practical approach to complex infrastructure challenges.
  • Ability to work independently in a fast-moving environment while taking ownership of significant architectural decisions.
  • Excellent communication skills and the ability to explain complex technical architectures and tradeoffs to technical stakeholders.
Benefits
  • Competitive compensation ranging from $184,000 to $318,000 OTE, including base salary and performance bonus.
  • Equity in the form of RSUs may be available at certain salary grades.
  • 100% company-paid medical, dental, and vision insurance for employees and their families.
  • 401(k) plan with up to a 4% company match and immediate vesting.
  • Paid parental leave: 20 weeks for primary caregivers and 12 weeks for secondary caregivers.
  • Remote work reimbursement of up to $85 per month for mobile and internet expenses.
  • Company-paid short-term disability, long-term disability, and life insurance.
  • Remote work flexibility for U.S.-based employees.
  • Career growth and ongoing learning opportunities.
  • Opportunity to work on large-scale AI and machine-learning infrastructure projects.
  • Collaborative, international environment with experienced engineering and AI professionals.
  • High level of ownership and opportunities to contribute to technically ambitious infrastructure initiatives.
  • Work environment that encourages innovation, initiative, and continuous improvement.

#LI-CL1

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

GPU Cluster Architect (Fully Remote)
GPU Cluster Architect (Fully Remote)

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior Solutions Architect - AI Infrastructure
Senior Solutions Architect - AI Infrastructure

NVIDIA • California (MO)

On-site
USD 184,000 - 356,500
Equity
Benefits
Senior Solutions Architect, Generative AI
Senior Solutions Architect, Generative AI

NVIDIA • California (MO)

On-site
USD 184,000 - 356,500
Senior HPC AI Cluster Engineer
Senior HPC AI Cluster Engineer

NVIDIA • United States

On-site
USD 176,000 - 334,000
Equity
Benefits
Senior GPU System Architect
Senior GPU System Architect

NVIDIA • United States

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior GPU Cluster Architect for AI Infra at Scale
Senior GPU Cluster Architect for AI Infra at Scale

Partner Company • United States

Remote
USD 184,000 - 318,000
Medical insurance
Dental insurance
Vision insurance
+1
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Senior Solutions Architect - AI Infrastructure
Senior Solutions Architect - AI Infrastructure

NVIDIA • New York (NY)

On-site
USD 184,000 - 356,500