Remote ML Infrastructure Engineer — GPU Clusters & AI Platform

United States Digital Space LLC

United States

Remote

USD 100,000 - 150,000

Full time

5 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

BV Teck seeks an AI Infrastructure Engineer to design, build, and operate the platform powering large‑scale AI training and inference workloads. You will focus on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers, with emphasis on reliability and cost efficiency.

The role requires hands‑on experience with production AI infrastructure at scale, understanding hardware, kernel, scheduler, and ML frameworks, and strong software

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider’s ML infrastructure offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.

Responsibilities

  • Design and operate GPU and accelerator infrastructure for training and inference across on‑prem, cloud, and hybrid setups.
  • Build scheduling, queuing, and resource-sharing systems to maximize accelerator utilization.
  • Integrate PyTorch, JAX, DeepSpeed, Megatron-LM, and Ray Train into a unified platform.
  • Operate high-performance storage systems and data pipelines for fast data access.
  • Design networking architectures with RDMA, InfiniBand, NCCL, and high-bandwidth communication.
  • Develop observability for AI workloads including utilization and failure analytics.
  • Implement checkpointing, restart, and fault-tolerance for long-running training jobs.
  • Drive cost optimization across compute, storage, and networking via scheduling and spot capacity.
  • Develop tooling and workflows to empower researchers to launch experiments safely.
  • Plan capacity with research teams for upcoming training runs.
  • Implement security controls and multi-tenant access management.
  • Automate cluster provisioning, lifecycle management, and config enforcement.
  • Maintain runbooks and capacity dashboards for the AI platform.

Skills

Python
Systems programming (Go/C++)
Distributed training
CI/CD
Linux internals
Communication

Education

Bachelor’s or Master’s degree in Computer Science or related field
Six+ years of infrastructure/engineering experience

Tools

Kubernetes
Slurm
Ray
NCCL

Job description

BV Teck seeks an AI Infrastructure Engineer to design, build, and operate the platform powering large‑scale AI training and inference workloads. You will focus on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers, with emphasis on reliability and cost efficiency.

The role requires hands‑on experience with production AI infrastructure at scale, understanding hardware, kernel, scheduler, and ML frameworks, and strong software

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Systems Engineer — Scale GPU ML Infra
Remote AI Systems Engineer — Scale GPU ML Infra

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 90,000 - 100,000
Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Senior AI Infrastructure Engineer - Remote & Scaled GPU
Senior AI Infrastructure Engineer - Remote & Scaled GPU

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 160,000
ML Infrastructure Engineer: Build Scalable GPU Clusters
ML Infrastructure Engineer: Build Scalable GPU Clusters

Cursor • California (MO)

On-site
USD 140,000 - 185,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Remote ML Infrastructure Engineer: Scale AI Inference
Remote ML Infrastructure Engineer: Scale AI Inference

JobCubby • Redwood City (CA)

Hybrid
USD 105,000 - 143,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Senior ML Infrastructure Engineer – Remote
Senior ML Infrastructure Engineer – Remote

Bright Vision Technologies • Redwood City (CA), San Mateo (CA)

On-site
USD 105,000 - 143,000
Senior Remote ML Infrastructure Engineer
Senior Remote ML Infrastructure Engineer

BairesDev • Peru (IL)

On-site
USD 120,000 - 170,000
Remote work
Payment in USD
Home setup provided
+3
Senior AI Platform Engineer - Remote Cloud-Native ML Infra
Senior AI Platform Engineer - Remote Cloud-Native ML Infra

Bright Vision Technologies • Columbus (OH), Dublin (OH)

On-site
USD 130,000 - 180,000