Infrastructure, Large-scale Training San Jose

Hark, Inc.

San Jose (CA)

On-site

USD 180,000 - 450,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hark, Inc. is seeking a Member of Technical Staff, Infrastructure Compute in San Jose, California to lead the development of large-scale GPU computing clusters. The ideal candidate should have significant experience in systems engineering and machine learning.

Your responsibilities will include designing Infrastructure as Code, optimizing deployment pipelines, and collaborating with ML researchers. A competitive salary range of $180,000 - $450,000 is offered, reflecting your expertise in this vital role.

Qualifications

  • 5+ years of experience in infrastructure, systems, or platform engineering.
  • At least 2 years working in ML or HPC environments.
  • Demonstrated experience managing GPU clusters.

Responsibilities

  • Design, implement, and maintain Infrastructure as Code (IaC) best practices.
  • Enhance and harden CI/CD deployment pipelines.
  • Monitor system health and define SLOs.

Skills

Infrastructure engineering
Systems engineering
Machine learning infrastructure
GPU clusters
Networking fundamentals

Education

5+ years experience in relevant field

Tools

Kubernetes
Pulumi
Rust
Go
PyTorch
Ray

Job description

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

We are looking for a Member of Technical Staff, Infrastructure Compute to lead and manage large-scale GPU computing clusters powering our AI training and deployment workloads. You'll work at the intersection of systems engineering and machine learning infrastructure, owning the reliability, scalability, and efficiency of the compute platform that our research and engineering teams depend on. This is a high-impact, highly technical role suited for someone who thrives in complex distributed systems environments and cares deeply about infrastructure as a product.

Responsibilities
  • Design, implement, and maintain Infrastructure as Code (IaC) best practices to enable repeatable, auditable, and scalable cluster provisioning.
  • Enhance and harden CI/CD deployment pipelines to ensure robust, secure, and low-latency model service delivery across production environments.
  • Own and evolve stable training infrastructure operating at the scale of 10,000+ GPUs, including job scheduling, fault tolerance, and network fabric optimization.
  • Partner closely with ML researchers and engineers to understand compute bottlenecks and translate them into infrastructure improvements.
  • Monitor system health, define SLOs, and lead incident response for critical training and inference workloads.
  • Drive capacity planning, cost efficiency initiatives, and hardware lifecycle management across the GPU fleet.
  • Contribute to internal tooling and platform abstractions that improve developer experience for teams consuming compute resources.
Requirements
  • 5+ years of experience in infrastructure, systems, or platform engineering, with at least 2 years working in ML or HPC environments.
  • Demonstrated experience managing GPU clusters or large-scale distributed compute infrastructure.
  • Strong proficiency in at least one systems or infrastructure programming language.
  • Deep understanding of networking fundamentals (RDMA, InfiniBand, or RoCE a plus) relevant to high-throughput training workloads.
  • Experience with container orchestration, job scheduling, and multi-tenant resource management.
  • Proven track record owning production systems with high reliability requirements.
  • Strong debugging and observability skills across the full infrastructure stack.
Bonus Qualifications
  • Kubernetes (K8s) — particularly experience operating large, GPU-aware clusters.
  • Pulumi or similar modern IaC tooling.
  • Rust and/or Go for systems-level tooling and performance-critical services.
  • Familiarity with PyTorch and Ray for understanding workload patterns and integration requirements.
Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Infrastructure, Large-scale Training
Infrastructure, Large-scale Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Senior ML Infra Architect — Large-Scale GPU Training
Senior ML Infra Architect — Large-Scale GPU Training

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Staff Infra Engineer – Large-Scale GPU Training
Staff Infra Engineer – Large-Scale GPU Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff, Mid-training San Jose
Member of Technical Staff, Mid-training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Technical Lead, On-Device AI Inference San Jose
Technical Lead, On-Device AI Inference San Jose

Hark • San Jose (CA), Northern (KY)

Hybrid
USD 300,000 - 500,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Supercomputing
Software Engineer, Supercomputing

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Unlimited PTO
Parental leave
+1