Software Engineer - AI Infrastructure

Hamilton Barnes Associates Limited

San Francisco (CA)

Hybrid

USD 300,000 - 500,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Early-stage equity
Founding engineer role
Equity package

Job summary

Hamilton Barnes Associates Limited is seeking a Principal Software Engineer to own the software architecture that connects thousands of GPUs, high‑performance networking, storage, schedulers and cloud orchestration platforms in a stealth‑mode hyperscale AI infrastructure project.

This hands‑on leadership role designs and builds infrastructure software for a platform enabling tens of thousands of GPUs to operate as a unified environment, with flexible remote or hybrid work and a substantial

Qualifications

  • 8+ years in distributed systems, infrastructure platforms, HPC, or large-scale cloud services.
  • Proficiency in Go, Rust, C++, Python or similar systems languages.
  • Experience building infrastructure software, not business apps.
  • Deep Linux knowledge, OS internals, networking, distributed computing.
  • Experience with HPC schedulers such as Slurm or large‑scale Kubernetes.
  • Strong understanding of APIs, SOA, and cloud-native platforms.
  • Ability to lead technically while remaining hands-on.

Responsibilities

  • Design and develop core platform software powering large-scale HPC and AI infra.
  • Build distributed systems for GPU scheduling, resource allocation, and workload orchestration.
  • Create integrations across Slurm, Kubernetes, cloud orchestration, and custom schedulers.
  • Develop services for provisioning, telemetry, observability, health, and remediation.
  • Optimize performance across GPU fabrics, storage, and distributed compute environments.
  • Collaborate with networking, platform, storage, and hardware teams to maximize efficiency.
  • Design APIs, automation, and tooling to improve operability and customer experience.
  • Contribute to long-term platform architecture for hyperscale AI and HPC.
  • Drive reliability, scalability, and operational excellence across thousands of nodes.
  • Mentor engineers and set software engineering standards.

Skills

Go
Rust
C++
Python
Distributed Systems
Linux
Kubernetes
Slurm
Networking
Leadership

Tools

Slurm
Kubernetes
Linux
Cloud Orchestration

Job description

Are you looking for an exciting new opportunity?

Join a stealth-mode hyperscale infrastructure startup building a 300MW+ AI compute platform designed to power the next generation of large-scale training, inference, and sovereign cloud environments. Backed by significant capital and industry-leading talent, the company is creating one of the most ambitious AI infrastructure projects currently under development.

The Principal Software Engineer will take ownership of the software architecture that connects thousands of GPUs, high-performance networking fabrics, storage systems, schedulers, and cloud orchestration platforms. This hands‑on technical leadership role will be responsible for designing and building the infrastructure software that powers one of the world's largest AI compute environments. The successful candidate will solve complex challenges across distributed systems, cluster orchestration, workload scheduling, GPU resource management, infrastructure automation, and platform scalability at exceptional scale.

This is a unique opportunity to design software that enables tens of thousands of GPUs to operate as a unified platform while helping shape the future of hyperscale AI infrastructure and supporting next-generation AI workloads.

Responsibilities:

  • Design and develop core platform software powering large-scale HPC and AI infrastructure environments.
  • Build distributed systems responsible for GPU scheduling, resource allocation, workload orchestration, and cluster management.
  • Develop software integrations across Slurm, Kubernetes, cloud orchestration platforms, and custom scheduling frameworks.
  • Create infrastructure services for provisioning, telemetry, observability, health monitoring, and automated remediation.
  • Optimize software performance across high-bandwidth GPU fabrics, storage platforms, and distributed compute environments.
  • Collaborate with networking, platform, storage, and hardware teams to maximize cluster efficiency and utilization.
  • Design APIs, automation frameworks, and developer tooling that improve infrastructure operability and customer experience.
  • Contribute to long-term platform architecture supporting hyperscale AI and HPC deployments.
  • Drive software reliability, scalability, and operational excellence across thousands of nodes and GPUs.
  • Mentor engineers and establish software engineering standards across the infrastructure organization.

Skills / Must Have:

  • 8+ years of experience developing software for distributed systems, infrastructure platforms, HPC environments, or large-scale cloud services.
  • Strong software engineering expertise in Go, Rust, C++, Python, or similar systems‑level languages.
  • Experience building infrastructure software rather than traditional business applications.
  • Deep understanding of Linux systems, operating systems internals, networking, and distributed computing concepts.
  • Experience with HPC schedulers such as Slurm or large‑scale Kubernetes environments.
  • Strong knowledge of infrastructure automation, APIs, service‑oriented architectures, and cloud‑native platforms.
  • Experience designing highly scalable, fault‑tolerant distributed systems.
  • Understanding of observability, telemetry, monitoring, and operational tooling at scale.
  • Ability to operate as a technical leader while remaining highly hands‑on.

Highly Desirable:

  • Experience supporting NVIDIA GPU environments including H100, H200, B200, GB200, DGX, or HGX platforms.
  • Knowledge of CUDA, NCCL, NVLink, NVSwitch, GPUDirect RDMA, or GPU resource management.
  • Experience with InfiniBand, RoCE, Spectrum-X, Cumulus Linux, or high‑performance networking fabrics.
  • Familiarity with AI infrastructure platforms, distributed training systems, and ML workload orchestration.
  • Experience working within hyperscalers, GPU cloud providers, HPC vendors, AI labs, or large‑scale infrastructure startups.
  • Contributions to open‑source infrastructure, Kubernetes, Slurm, or distributed systems projects.

Benefits:

  • Significant early-stage equity participation.
  • Founding engineer opportunity within a next‑generation hyperscale AI infrastructure platform.
  • Direct influence over architecture decisions across software, infrastructure, and AI platform design.
  • Opportunity to build systems operating at unprecedented GPU scale.
  • Flexible remote or hybrid working arrangements.
  • Work alongside industry leaders in AI infrastructure, cloud computing, and hyperscale operations.
  • Significant Equity Package
  • Executive‑Level Incentives Depending on Experience

Salary:

  • $300,000 - $500,000+
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer (SRE) - AI Inftastructure
Senior Site Reliability Engineer (SRE) - AI Inftastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 270,000 - 330,000
Equity
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Infrastructure Product Engineer - AI Infrastructure
Infrastructure Product Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 300,000 - 350,000
Equity
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
HPC Engineer - AI Infrastructure
HPC Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
AI Infrastructure Engineer
AI Infrastructure Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
Cash bonus
Founding engineer equity
Benefits
System Software Engineer - AI
System Software Engineer - AI

Entrada Ventures • Palo Alto (CA)

Hybrid
USD 140,000 - 200,000
Equity
401k
Benefits
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 213,000 - 288,000
Early-stage equity
Direct access to leadership
Staff Site Reliability Engineer - AI Infrastructure
Staff Site Reliability Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 297,500 - 402,500
Huge stock options
Company bonus
Unlimited PTO
+1
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000