Senior HPC Engineer

Hamilton Barnes ?

San Francisco (CA)

On-site

USD 180,000 - 250,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Relocation provided
Founding-level ownership

Job summary

Neocloud seeks a founding member to build and operate GPU-focused compute infrastructure in San Francisco. You will directly influence the platform that powers AI workloads, with ownership and visibility from the founders.

This onsite role relocates to San Francisco as needed. You will stand up production Kubernetes/Slurm clusters with custom GPU orchestration, design storage and networking for high-scale workloads, and own incident response end-to-end.

Qualifications

  • Operate production Kubernetes and/or Slurm clusters with custom GPU orchestration and lifecycle automation.
  • Deep hands-on experience with distributed object storage and NVMe storage clusters for large-scale AI workloads.
  • Design and operate scalable inference platforms for low-latency, high-throughput GPU acceleration.
  • Own end-to-end reliability and incident response; document runbooks and processes for new hires.

Responsibilities

  • Stand up and operate Kubernetes/Slurm clusters with custom GPU orchestration.
  • Design and implement storage architecture (object storage, NVMe clusters, high-bandwidth networking) for real customer workloads.
  • Translate workload requirements into infrastructure decisions with customers.
  • Own cluster reliability and production incident response end-to-end; document runbooks for future hires.

Skills

Kubernetes
Slurm
GPU orchestration
Cluster reliability
Distributed object storage
NVMe storage
High-bandwidth networking
AI workloads optimization
Incident response

Job description

Join a Stealth Neocloud building GPU compute infrastructure from the ground up in San Francisco. As one of the founding team, the systems you design and build this year are the systems the company runs on. You'll have direct access to and collaboration with the founders, high ownership, high visibility, and no bureaucracy standing between you and the rack.

Responsibilities:

Stand up and operate production Kubernetes and/or Slurm clusters with custom GPU orchestration (scheduling logic, topology-aware placement, GPU lifecycle automation) built or substantially customised by you, not tooling run out of the box.

Design and implement the storage architecture (object storage, NVMe clusters, high-bandwidth networking) that the platform depends on, tuning it for real customer workloads at scale.

Work directly with customers to translate workload requirements into infrastructure decisions: sizing clusters, configuring storage and networking, and adjusting orchestration to fit what they actually need to run.

Own cluster reliability and production incident response end-to-end, documenting runbooks and operational processes so the next hire doesn't start from zero.

Skills/Must have:

Operating Kubernetes and/or Slurm clusters at scale with custom GPU orchestration, scheduling logic, and lifecycle automation built or substantially customised in-house.

Deep hands-on experience with distributed object storage, NVMe storage clusters, high-bandwidth networking fabrics, and storage architectures optimised for large-scale AI workloads.

Designing and operating scalable inference platforms for deploying, serving, and optimising large AI models with low-latency, high-throughput GPU acceleration.

End-to-end ownership of distributed system reliability, incident response, and operational maturity; comfort being the first call when production breaks.

Based in or willing to relocate to San Francisco (onsite).

Founding-level ownership and visibility.

Direct access to and collaboration with the founders.

On-site role in San Francisco; relocation provided if needed.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding HPC Engineer - GPU Compute & Infra (SF)
Founding HPC Engineer - GPU Compute & Infra (SF)

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 180,000 - 250,000
Relocation provided
Founding-level ownership
Senior Network Solutions Architect
Senior Network Solutions Architect

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 190,000 - 260,000
HPC Engineer - AI Infrastructure
HPC Engineer - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 235,000 - 315,000
Founding engineer equity
Full benefits package
Founding GPU Cluster Architect — Hands-On, SF
Founding GPU Cluster Architect — Hands-On, SF

Hamilton Barnes ? • San Francisco (CA)

On-site
USD 190,000 - 260,000
AI Infrastructure Engineer
AI Infrastructure Engineer

Acceler8 Talent • San Francisco (CA), Northern (KY)

Hybrid
USD 180,000 - 300,000
Cash bonus
Founding engineer equity
Benefits
Senior GPU Infrastructure Engineer
Senior GPU Infrastructure Engineer

Hyperbolic • San Francisco (CA)

On-site
Senior HPC & GPU Cluster Architect
Senior HPC & GPU Cluster Architect

The Consensus • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Visa sponsorships
401(k) retirement matching
Medical, dental & vision insurance
+2
Senior Solution Architect - AI Infrastructure
Senior Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 233,000 - 316,000
Founding-level ownership and visible价值
Direct access to founders
Onsite role in San Francisco
HPC/ GPU Cluster Architect
HPC/ GPU Cluster Architect

Electric Capital • San Francisco (CA)

Hybrid
USD 220,000 - 300,000
Generous equity grant
401(k) matching
Comprehensive medical, dental, and vision insurance
+3
Senior HPC & GPU Cluster Architect — Scale & Automate
Senior HPC & GPU Cluster Architect — Scale & Automate

San Francisco Compute Company • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Generous equity grant
Competitive salary
Visa sponsorship
+6