Remote AI Systems Engineer - Scale GPU ML Infra

Bright-Vision-Technologies

United States

Remote

USD 90,000 - 100,000

Full time

12 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an AI Systems Engineer to design, build, and operate the platform layer powering large-scale AI training and inference workloads. The role emphasizes GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers.

The ideal candidate has built or operated production AI infrastructure at scale, understands interactions between hardware, kernel, scheduler, and ML framework, and brings

Qualifications

  • Bachelor's or Master's degree in Computer Science or a related field.
  • 6+ years of infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infra.
  • Strong Python and one systems language (Go or C++).
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray or similar ML schedulers.
  • Linux internals, networking, and high-performance storage knowledge.
  • Experience with at least one major cloud provider's ML infra offerings.
  • Strong software engineering practices: testing, CI/CD, code review.
  • Excellent communication and cross-functional collaboration.

Responsibilities

  • Design, build, and operate the platform layer powering large-scale AI training and inference workloads.
  • Focus on GPU clusters, distributed training frameworks, scheduling, and storage performance.
  • Enhance developer experience for ML engineers and researchers with reliability and cost control.
  • Understand hardware, kernel, scheduler, and ML framework interactions; apply strong software engineering discipline.

Skills

Python
Go
C++
distributed training
CI/CD
testing
code review
communication
cross-functional collaboration
rapid problem solving
cloud infrastructure

Education

Bachelor's or Master's degree in Computer Science or related field

Tools

Kubernetes
Slurm
Ray
InfiniBand / RDMA

Job description

Bright Vision Technologies is seeking an AI Systems Engineer to design, build, and operate the platform layer powering large-scale AI training and inference workloads. The role emphasizes GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers.

The ideal candidate has built or operated production AI infrastructure at scale, understands interactions between hardware, kernel, scheduler, and ML framework, and brings

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Systems Engineer — Scale GPU ML Infra
Remote AI Systems Engineer — Scale GPU ML Infra

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 90,000 - 100,000
Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Remote ML Infrastructure Engineer: Scale AI Inference
Remote ML Infrastructure Engineer: Scale AI Inference

JobCubby • Redwood City (CA)

Hybrid
USD 105,000 - 143,000
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform
Remote ML Infrastructure Engineer — GPU Clusters & AI Platform

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Remote ML Systems Engineer Scalable Inference Platforms
Remote ML Systems Engineer Scalable Inference Platforms

Bright Vision Technologies • Round Rock (TX)

On-site
USD 145,000 - 165,000
Senior AI Platform Engineer - Remote Cloud-Native ML Infra
Senior AI Platform Engineer - Remote Cloud-Native ML Infra

Bright Vision Technologies • Columbus (OH), Dublin (OH)

On-site
USD 130,000 - 180,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Remote Model Serving Engineer - Scale ML Inference
Remote Model Serving Engineer - Scale ML Inference

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 74,000 - 98,000
Remote AI Performance Engineer - GPU & ML Systems
Remote AI Performance Engineer - GPU & ML Systems

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 75,000 - 100,000
Remote work
Remote AI Data Operations Engineer - Scale ML Pipelines
Remote AI Data Operations Engineer - Scale ML Pipelines

Bright Vision Technologies • Euless (TX), Bedford (TX)

On-site
USD 150,000 - 165,000