Senior AI Infrastructure Engineer - Remote & Scaled GPU

Bright Vision Technologies

Nashua (NH)

On-site

USD 100,000 - 160,000

Full time

7 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an AI Infrastructure Engineer to design, build, and operate the platform layer powering large-scale AI training and inference workloads. The role emphasizes GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control.

The ideal candidate has built or operated production AI infrastructure at scale, understands

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 10+ years of experience in infrastructure, platform, or HPC engineering.
  • Hands-on experience operating GPU clusters or large-scale ML training infrastructure.
  • Strong proficiency in Python and at least one systems language such as Go or C++.
  • Deep understanding of distributed training, accelerator architectures, and collective communication.
  • Experience with Kubernetes, Slurm, Ray, or similar scheduling systems for ML workloads.
  • Strong understanding of Linux internals, networking, and high-performance storage.
  • Experience with at least one major cloud provider’s ML infrastructure offerings.
  • Strong software engineering practices including testing, CI/CD, and code review.

Responsibilities

  • Design and operate GPU and accelerator infrastructure for training and inference across on-prem, cloud, and hybrid configurations.
  • Build scheduling, queueing, and resource-sharing systems to maximize accelerator utilization.
  • Integrate PyTorch, JAX, DeepSpeed, FSDP, Megatron-LM, and Ray Train into a unified platform.
  • Operate high-performance storage systems and data pipelines for near-line-rate data feeds.
  • Design networking architectures supporting RDMA, InfiniBand, NCCL, and high-bandwidth IPC.
  • Develop observability for AI workloads including utilization, throughput, and fault analytics.
  • Implement checkpointing, restart, and fault-tolerance for long-running training jobs.
  • Drive cost optimization through scheduling, spot capacity, and right-sizing.
  • Develop developer tooling and paved-road workflows for researchers.

Skills

Python
Go
C++
Distributed systems
Linux
CI/CD
Kubernetes
Ray
Networking

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

Kubernetes
Slurm
Ray
PyTorch
DeepSpeed

Job description

Bright Vision Technologies is seeking an AI Infrastructure Engineer to design, build, and operate the platform layer powering large-scale AI training and inference workloads. The role emphasizes GPU clusters, distributed training frameworks, scheduling, storage performance, and developer experience for ML engineers and researchers, with strong emphasis on reliability, efficiency, and cost control.

The ideal candidate has built or operated production AI infrastructure at scale, understands

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Remote AI Systems Engineer — Scale GPU ML Infra
Remote AI Systems Engineer — Scale GPU ML Infra

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 90,000 - 100,000
Remote AI Platform Engineer - Scalable ML Inference
Remote AI Platform Engineer - Scalable ML Inference

Bright Vision Technologies • Charlotte (NC)

On-site
USD 100,000 - 150,000
Senior AI Systems Performance Engineer (Remote)
Senior AI Systems Performance Engineer (Remote)

Bright Vision Technologies • Huntersville (NC)

On-site
USD 130,000 - 180,000
Remote work
Senior ML Infrastructure Engineer – Remote
Senior ML Infrastructure Engineer – Remote

Bright Vision Technologies • Redwood City (CA), San Mateo (CA)

On-site
USD 105,000 - 143,000
Senior AI Performance Engineer — Remote
Senior AI Performance Engineer — Remote

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 150,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
Remote AI Performance Engineer - GPU & ML Systems
Remote AI Performance Engineer - GPU & ML Systems

Bright Vision Technologies • Farmington Hills (MI)

On-site
USD 75,000 - 100,000
Remote work
Remote ML Infrastructure Engineer: Scale AI Inference
Remote ML Infrastructure Engineer: Scale AI Inference

JobCubby • Redwood City (CA)

Hybrid
USD 105,000 - 143,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1