Member of Technical Staff - AI Infrastructure

Veeda AI

California

On-site

USD 150,000 - 210,000

Full time

37 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Veeda AI in the United States (California) is assembling a small, high‑impact team to architect and operate scalable HPC infrastructure for cutting‑edge multimodal AI workloads. You’ll own GPU clusters, networking, storage, and security, ensuring reliable data delivery and observability across systems.

Ideal candidates have deep Linux and HPC experience, plus automation skills and familiarity with PyTorch/NCCL, Slurm, and cloud PoCs.

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in HPC or infrastructure engineering.
  • Strong troubleshooting skills in low-level Linux networking, kernel and PCIe/NUMA tuning, hardware diagnostics, and distributed storage (NFS, NVMe-oF, Lustre, Ceph).
  • Proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting), and treats cluster configuration as reviewed, version-controlled code.
  • Experience supporting distributed deep learning workloads (PyTorch/NCCL, Ray, DeepSpeed) closely enough to distinguish infrastructure faults from model bugs with controlled benchmarks.
  • Experience administering Linux-based HPC clusters (Slurm/Kubernetes) or cloud PoCs and security policy management.

Responsibilities

  • GPU Cluster Operations: Design, deploy, and operate Slurm-on-Kubernetes GPU clusters, including compute health, interconnects, storage, and scheduling.
  • Access Control and Security: Design and operate identity provisioning and access control for source code, cluster, data storages, and dev tools.
  • Network and Data: Design and maintain distributed data transmission and caching services for just-in-time data delivery to clusters.
  • Observability & Hardware Health: Build telemetry pipelines (Prometheus, Grafana, DCGM, Redfish) to detect hardware issues and coordinate with cloud vendors.
  • Resource Forecasting: Work with Model and Data teams to determine resource needs and lead PoC rounds with vendors.
  • Infra-OPEX: Prepare operational expenses forecasts and provide expert opinion on infra spending decisions.
  • End-to-End Ownership: Lead projects through the full software lifecycle including CI/CD and production observability.

Skills

HPC administration
Linux networking
Automation tooling
Distributed DL workloads
Security policies

Education

Bachelor's in CS/CE or equivalent

Tools

Slurm
Kubernetes
GitLab CI
Terraform
Ansible

Job description

About Us
Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.
About Us
Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.
Responsibilities
  • GPU Cluster Operations: Design, deploy, and operate Slurm-on-Kubernetes GPU clusters. Own the compute (CPU/GPU device health), interconnect (nvlink, Infiniband, RoCE), storage (linux file systems, Lustre), and scheduler (topology aware scheduling, QoS, partitions, fair-share) resources.
  • Access Control and Security: Design, deploy, and operate identity provisioning and access control for our source code, cluster, knowledge bases, data storages, and dev tools.
  • Network and Data: Design and maintain distributed data transmission and caching services that provide just-in-time delivery of data and checkpoints to the clusters in a timely and cost-effective manner. Collaborate with internal stakeholders and external vendors to find the optimal path between data and compute.
  • Observability & Hardware Health: Build the telemetry pipeline (Prometheus, Grafana, DCGM, BMC/Redfish) that catches Xid and ECC errors, thermal throttling, and link flaps, and work with cloud vendors to diagnose and resolve issues in timely manner.
  • Resource Forecasting: Work with Model and Data teams to determine resource requirements and delivery cadence for the next month/quarter/year. Lead PoC rounds with vendors for compute and data resources.
  • Infra-OPEX: Prepare operational expenses forecast and reports. Provide expert opinion on key infra spending decisions.
  • End-to-End Ownership: Lead projects through the complete software lifecycle, including technical specs, implementation, CI/CD, on-call support, and production observability.
Requirements
  • Bachelor's degree in Computer Science, Computer Engineering, or equivalent hands-on experience in high-performance computing (HPC) or infrastructure engineering.
  • Strong troubleshooting skills below the framework layer: low-level Linux networking, kernel and PCIe/NUMA tuning, hardware diagnostics, and network or distributed storage (NFS, NVMe-oF, Lustre, Ceph).
  • Proficient in automation and infrastructure-as-code tools (e.g., Ansible, Terraform, Helm, Python/Bash scripting), and treats cluster configuration as reviewed, version-controlled code.
  • Experience supporting distributed deep learning workloads (PyTorch/NCCL, Ray, DeepSpeed) closely enough to tell an infrastructure fault from a model bug, and able to prove which with a controlled benchmark.
  • One of the following:
    • Deep hands-on expertise administering Linux-based HPC clusters running Slurm or Kubernetes (job scheduling, cgroups, fair-share policies, topology configuration, GPU device plugins).
    • Experience working with cloud providers on workload profiling, Proof of Concepts (PoC), resource cost strategies.
    • Experience managing security policies. Hands‑on‑experience with building security layers from overall strategy to IaC. Educating fast moving Dev teams of security practices.
Nice to Have
  • Experience managing high-density GPU infrastructure (NVIDIA H100/H200, B200, and GB200 NVL72 systems, DGX/HGX architectures, liquid‑cooled racks).
  • Built container and image supply chains for HPC workloads (Enroot, Pyxis, Docker, custom Kubernetes operators).
  • Experience running hybrid capacity, combining owned hardware with cloud or neocloud burst under a single scheduler.
  • Experience running secure multi‑tenant research environments with SSO, per‑team quota, and interactive access that stays fast.
  • Fluency with Prometheus, Grafana, and OpenTelemetry, and instrument workloads before its first outage.
  • Contributed to Slurm plugins, Kubernetes operators, or open‑source cluster and observability tooling.
  • Experience with GitLab, especially GitLab CI, for managing infrastructure‑as‑code and automation pipelines.
  • Experience benchmarking fabric, storage, or scheduler performance and publishing results internally or externally to settle a design or procurement decision.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda AI • Seattle (WA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - AI Infrastructure
Member of Technical Staff - AI Infrastructure

Veeda Innovation • California (MO)

Hybrid
USD 180,000 - 240,000
Member of Technical Staff - ML Operations
Member of Technical Staff - ML Operations

Veeda AI • California

On-site
USD 120,000 - 190,000
Senior/Staff Software Engineer, Kubernetes Infrastructure
Senior/Staff Software Engineer, Kubernetes Infrastructure

Kindredventures • United States

On-site
USD 140,000 - 190,000
Member of Technical Staff — Compute Cluster
Member of Technical Staff — Compute Cluster

Kindredventures • San Francisco (CA)

On-site
USD 140,000 - 230,000
Principal Infrastructure Engineer, AI Cluster Performance & Validation
Principal Infrastructure Engineer, AI Cluster Performance & Validation

Nscale • New York (NY), San Francisco (CA), Seattle (WA)

On-site
USD 180,000 - 240,000
HPC Infrastructure Engineer
HPC Infrastructure Engineer

Arcadia • San Francisco (CA)

On-site
USD 180,000 - 260,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Hyperbolic • San Francisco (CA)

On-site
USD 180,000 - 260,000
Manager of Technical Support Engineering (Bare Metal)
Manager of Technical Support Engineering (Bare Metal)

Coreweave • Sunnyvale (CA)

On-site
USD 180,000 - 240,000
Sr HPC Hardware Engineer
Sr HPC Hardware Engineer

Career Techniques • Dallas (TX)

On-site
USD 120,000 - 180,000