Remote AI Systems Engineer — Scalable GPU Infra

Bright Vision Technologies

United States

Remote

USD 90,000 - 100,000

Full time

5 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bright Vision Technologies seeks an AI Systems Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role emphasizes GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control.

The ideal candidate has built or operated production AI infrastructure at scale, understands hardware,

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider’s ML infrastructure offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.

Responsibilities

  • Design, build, and operate the platform layer powering large-scale AI training and inference workloads.
  • Focus on GPU clusters, distributed training frameworks, scheduling, and storage performance.
  • Improve developer experience for ML engineers and researchers with reliability and cost control in mind.

Skills

GPU clusters
Python
Go/C++
Distributed training
Linux internals
Kubernetes/Slurm
CI/CD
Communication

Education

Bachelor’s/Master’s in CS or related

Tools

Ray
Slurm
Kubernetes

Job description

Bright Vision Technologies seeks an AI Systems Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role emphasizes GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control.

The ideal candidate has built or operated production AI infrastructure at scale, understands hardware,

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Senior AI Platform Engineer - Remote
Senior AI Platform Engineer - Remote

Bright Vision Technologies • Edison (NJ)

Remote
USD 100,000 - 160,000
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Remote GPU Systems Engineer for AI & HPC Optimization
Remote GPU Systems Engineer for AI & HPC Optimization

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Remote ML Platform Engineer - Scalable Inference & Systems
Remote ML Platform Engineer - Scalable Inference & Systems

Bright Vision Technologies • Edison (NJ)

Remote
USD 100,000 - 160,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Remote AI Pipeline Engineer - Petabyte-Scale Data
Remote AI Pipeline Engineer - Petabyte-Scale Data

Bright Vision Technologies • Maple Grove (MN)

Remote
USD 100,000 - 150,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Remote MLOps Engineer - Build Scalable AI Serving
Remote MLOps Engineer - Build Scalable AI Serving

Bright Vision Technologies • United States

Remote
USD 100,000 - 150,000