Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies

Framingham (MA)

On-site

USD 100,000 - 160,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with emphasis on reliability and cost control.

The ideal candidate has built or operated production AI infrastructure at scale, understands

Qualifications

  • Bachelor’s or Master’s degree in CS or related field.
  • 10+ years in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infra.
  • Proficiency in Python and at least one systems language (Go or C++).
  • Understanding of distributed training, accelerators, and collective communication.
  • Experience with Kubernetes, Slurm, Ray for ML workloads.
  • Strong Linux internals, networking, and high-performance storage.
  • Experience with cloud ML infra offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration.

Responsibilities

  • Design and operate GPU and accelerator infra for training and inference.
  • Build scheduling, queueing, and resource sharing to maximize utilization.
  • Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform.
  • Operate high-performance storage and data pipelines for near-line-rate training data.
  • Design networking architectures supporting RDMA, InfiniBand, NCCL.
  • Develop observability for AI workloads including throughput and failure analytics.
  • Implement checkpointing, restart, and fault-tolerance for long-running training jobs.
  • Drive cost optimization across compute, storage, and networking.
  • Develop tooling and paved-road workflows for researchers to run experiments safely.
  • Plan capacity with research and applied ML teams for upcoming runs.
  • Implement security controls and multi-tenant isolation.
  • Automate cluster provisioning, lifecycle management, and configuration enforcement.
  • Maintain runbooks, capacity dashboards, and operational documentation.
  • Stay current with AI infra research, accelerator hardware, and open-source tooling.

Skills

Python
Go or C++
Linux internals
Distributed training
GPU clusters
CI/CD
Communication

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

Kubernetes
Slurm
Ray
PyTorch
DeepSpeed
NCCL
Megatron-LM
FSDP
InfiniBand

Job description

Bright Vision Technologies is seeking an AI Infrastructure Engineer to design, build, and operate the platform layer that powers large-scale AI training and inference workloads. The role focuses on GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with emphasis on reliability and cost control.

The ideal candidate has built or operated production AI infrastructure at scale, understands

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Senior Remote ML Infrastructure Engineer: GPU & Scale
Senior Remote ML Infrastructure Engineer: GPU & Scale

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Remote AI Systems Engineer – Scalable ML Infra
Remote AI Systems Engineer – Scalable ML Infra

Bright Vision Technologies • Mountain View (CA)

On-site
USD 90,000 - 100,000
Senior AI Data Infrastructure Engineer – Remote
Senior AI Data Infrastructure Engineer – Remote

Bright Vision Technologies • Euless (TX), Bedford (TX)

On-site
USD 125,000 - 170,000
Senior ML Infrastructure Engineer — Remote
Senior ML Infrastructure Engineer — Remote

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 105,000 - 143,000
Senior ML Infra Engineer — Remote AI Serving
Senior ML Infra Engineer — Remote AI Serving

United States Digital Space LLC • United States

Remote
USD 90,000 - 150,000
Senior AI Data Pipeline Engineer (Remote)
Senior AI Data Pipeline Engineer (Remote)

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Senior AI Data Engineer: Scalable Pipelines (Remote)
Senior AI Data Engineer: Scalable Pipelines (Remote)

United States Digital Space LLC • United States

Remote
USD 150,000 - 165,000
Remote AI Data Infrastructure Engineer
Remote AI Data Infrastructure Engineer

United States Digital Space LLC • United States

Remote
USD 125,000 - 170,000
Remote AI Operations Engineer for Large-Scale ML Pipelines
Remote AI Operations Engineer for Large-Scale ML Pipelines

Bright Vision Technologies • Flower Mound (TX)

On-site
USD 150,000 - 165,000