Remote ML Infrastructure Engineer: GPU Clusters & Scale

Bright Vision Technologies

Eden Prairie (MN)

Remote

USD 100,000 - 150,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an ML Infrastructure Engineer for 100% remote work in the U.S., focusing on GPU clusters, distributed training, and a unified AI platform. You will design and operate the platform layer powering large-scale AI workloads, with emphasis on reliability, efficiency, and cost control.

The role requires 6+ years of experience, strong software engineering practices, and collaboration with research and ML teams to plan capacity and ensure robust infrastructure for

Qualifications

  • Degree in computer science or related field.
  • 6+ years in infrastructure, platform, or HPC engineering.
  • Hands-on GPU clusters or large-scale ML infrastructure experience.
  • Proficiency in Python and at least one systems language (Go or C++).
  • Distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray for ML workloads.
  • Strong Linux, networking, and high-performance storage knowledge.
  • Experience with cloud ML infrastructure offerings from major providers.
  • CI/CD, testing, and code review practices.
  • Excellent communication and cross-functional collaboration.

Responsibilities

  • Design and operate GPU and accelerator infrastructure for training and inference.
  • Build scheduling, queueing, and resource-sharing systems for accelerator utilization.
  • Integrate frameworks like PyTorch, JAX, DeepSpeed, Megatron-LM, and Ray Train.
  • Operate high-performance storage and data pipelines for near-line-rate data.</li>
  • Design networking architectures with RDMA, InfiniBand, NCCL.
  • Build observability for AI workloads including utilization and training stability.
  • Implement checkpointing, restart, and fault-tolerance patterns for long-running training jobs.
  • Drive cost optimization across compute, storage, and networking via scheduling and spot capacity.
  • Develop developer tooling and paved-road workflows for researchers.
  • Plan capacity with research and applied ML teams.
  • Implement security, isolation, and multi-tenant access controls.
  • Automate cluster provisioning, lifecycle management, and configuration enforcement.
  • Maintain runbooks and capacity dashboards for the AI platform.
  • Stay current with AI infra research and open-source tooling.

Skills

Python
Go
C++
Kubernetes
Slurm
Ray
Linux
Networking
Storage
Cloud ML
CI/CD
Collaboration
Communication

Education

Bachelor’s or Master’s in CS

Job description

Bright Vision Technologies is seeking an ML Infrastructure Engineer for 100% remote work in the U.S., focusing on GPU clusters, distributed training, and a unified AI platform. You will design and operate the platform layer powering large-scale AI workloads, with emphasis on reliability, efficiency, and cost control.

The role requires 6+ years of experience, strong software engineering practices, and collaboration with research and ML teams to plan capacity and ensure robust infrastructure for

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Infra Engineer — Scale GPU Clusters
Remote AI Infra Engineer — Scale GPU Clusters

Bright Vision Technologies • Monroeville

Remote
USD 100,000 - 160,000
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior Remote ML Infrastructure Engineer
Senior Remote ML Infrastructure Engineer

BairesDev • Peru (IL)

On-site
USD 120,000 - 170,000
Remote work
Payment in USD
Home setup provided
+3
Senior ML Inference Platform Engineer — Remote
Senior ML Inference Platform Engineer — Remote

Bright Vision Technologies • Beaverton (OR)

Remote
USD 105,000 - 143,000
ML Infrastructure Engineer
ML Infrastructure Engineer

Bright Vision Technologies • Eden Prairie (MN)

Remote
USD 100,000 - 150,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Remote AI Research Clusters Engineer - ML Infra & GPU
Remote AI Research Clusters Engineer - ML Infra & GPU

NEPSE Trading • Northern (KY)

Hybrid
USD 124,000 - 196,000