Senior ML Infra Architect — Large-Scale GPU Training

Hark, Inc.

San Jose (CA)

On-site

USD 180,000 - 450,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hark, Inc. is seeking a Member of Technical Staff, Infrastructure Compute in San Jose, California to lead the development of large-scale GPU computing clusters. The ideal candidate should have significant experience in systems engineering and machine learning.

Your responsibilities will include designing Infrastructure as Code, optimizing deployment pipelines, and collaborating with ML researchers. A competitive salary range of $180,000 - $450,000 is offered, reflecting your expertise in this vital role.

Qualifications

  • 5+ years of experience in infrastructure, systems, or platform engineering.
  • At least 2 years working in ML or HPC environments.
  • Demonstrated experience managing GPU clusters.

Responsibilities

  • Design, implement, and maintain Infrastructure as Code (IaC) best practices.
  • Enhance and harden CI/CD deployment pipelines.
  • Monitor system health and define SLOs.

Skills

Infrastructure engineering
Systems engineering
Machine learning infrastructure
GPU clusters
Networking fundamentals

Education

5+ years experience in relevant field

Tools

Kubernetes
Pulumi
Rust
Go
PyTorch
Ray

Job description

Hark, Inc. is seeking a Member of Technical Staff, Infrastructure Compute in San Jose, California to lead the development of large-scale GPU computing clusters. The ideal candidate should have significant experience in systems engineering and machine learning.

Your responsibilities will include designing Infrastructure as Code, optimizing deployment pipelines, and collaborating with ML researchers. A competitive salary range of $180,000 - $450,000 is offered, reflecting your expertise in this vital role.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Infra Engineer – Large-Scale GPU Training
Staff Infra Engineer – Large-Scale GPU Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Infrastructure, Large-scale Training San Jose
Infrastructure, Large-scale Training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Infrastructure, Large-scale Training
Infrastructure, Large-scale Training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Senior ML Training Systems Engineer - Distributed GPU Infra
Senior ML Training Systems Engineer - Distributed GPU Infra

Baseten • San Francisco (CA)

On-site
USD 150,000 - 200,000
Competitive compensation, including equity
100% coverage of medical, dental, and vision insurance
Generous PTO policy
+2
Senior GPU ML Infra Engineer — Mid-Training & Inference
Senior GPU ML Infra Engineer — Mid-Training & Inference

Reflection AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML Infrastructure Engineer - GPU Training & MLOps
Senior ML Infrastructure Engineer - GPU Training & MLOps

Cssmerge • San Francisco (CA)

On-site
USD 224,000 - 280,000
Medical, Dental, Vision insurance
401(k)
Equity awards
+2
GPU Infra Solutions Architect for Large-Scale AI Clusters
GPU Infra Solutions Architect for Large-Scale AI Clusters

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform
Senior ML Infra Engineer: GPU-Optimized Kubernetes Platform

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
Senior AI Infra Engineer: GPU Compute on Kubernetes
Senior AI Infra Engineer: GPU Compute on Kubernetes

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000