Remote AI Systems Engineer — Scale GPU ML Infra

Bright Vision Technologies

Palo Alto (CA)

On-site

USD 90,000 - 100,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an AI Systems Engineer to design, build, and operate the platform layer that powers large‑scale AI training and inference workloads. You will focus on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, emphasizing reliability and cost efficiency.

The ideal candidate has built or operated production AI infrastructure at scale, and possesses strong software engineering

Qualifications

  • Bachelor’s or Master’s degree in CS or a related field.
  • 6+ years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider’s ML infrastructure offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.

Responsibilities

  • Design, build, and operate the platform layer powering AI training workloads.
  • Manage GPU clusters, distributed training frameworks, and scheduling.
  • Improve storage performance and developer experience for ML engineers.
  • Collaborate cross-functionally with hardware, kernel, and ML teams.
  • Ensure reliability, efficiency, and cost control of AI infrastructure.

Skills

Python
Go
C++
Kubernetes
Slurm
Ray
Distributed training
Linux internals
Networking
Cloud ML infra
CI/CD
Code review
Communication

Education

Bachelor’s or Master’s in Computer Science

Tools

Kubernetes
Slurm
Ray

Job description

Bright Vision Technologies is seeking an AI Systems Engineer to design, build, and operate the platform layer that powers large‑scale AI training and inference workloads. You will focus on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, emphasizing reliability and cost efficiency.

The ideal candidate has built or operated production AI infrastructure at scale, and possesses strong software engineering

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer - Remote & Scaled GPU
Senior AI Infrastructure Engineer - Remote & Scaled GPU

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 160,000
Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Remote ML Infrastructure Engineer: Scale AI Inference
Remote ML Infrastructure Engineer: Scale AI Inference

JobCubby • Redwood City (CA)

Hybrid
USD 105,000 - 143,000
Remote AI Platform Engineer - Scalable ML Inference
Remote AI Platform Engineer - Scalable ML Inference

Bright Vision Technologies • Charlotte (NC)

On-site
USD 100,000 - 150,000
Remote ML Systems Engineer Scalable Inference Platforms
Remote ML Systems Engineer Scalable Inference Platforms

Bright Vision Technologies • Round Rock (TX)

On-site
USD 145,000 - 165,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
Senior ML Infrastructure Engineer – Remote
Senior ML Infrastructure Engineer – Remote

Bright Vision Technologies • Redwood City (CA), San Mateo (CA)

On-site
USD 105,000 - 143,000
Remote AI Performance Engineer - GPU & ML Systems
Remote AI Performance Engineer - GPU & ML Systems

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 75,000 - 100,000
Remote work
Remote Model Serving Engineer - Scale ML Inference
Remote Model Serving Engineer - Scale ML Inference

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 74,000 - 98,000
Remote AI Data Operations Engineer - Scale ML Pipelines
Remote AI Data Operations Engineer - Scale ML Pipelines

Bright Vision Technologies • Euless (TX), Bedford (TX)

On-site
USD 150,000 - 165,000