Remote GPU & Cloud Infrastructure Engineer for AI

Innomium

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote-first policy
Annual learning budget
Company-provided equipment
Internet/home-office stipend
Paid time off
Equity/bonuses where applicable
Retirement contributions

Job summary

Innomium is seeking a GPU and Cloud Infrastructure Engineer to build the compute foundation for model teams, covering training, evaluation, fine-tuning, and inference.

You will design GPU workloads, containerized environments, storage and networking paths, and automation to reproduce experiments and releases across programs, while collaborating with researchers to optimize utilization, memory, and cost.

Qualifications

  • Experience with cloud infrastructure for GPU workloads.
  • Strong Linux, containers, networking, storage.
  • IaC fundamentals across cloud platforms.
  • Ability to diagnose failures across infra layers and document work.

Responsibilities

  • Build and operate GPU training, evaluation, fine-tuning, and inference environments.
  • Automate provisioning, container images, dependencies, secrets, networking, and storage.
  • Profile utilization, memory, throughput, queue behavior, data transfer, and cost.
  • Design model-serving and batch-execution paths with observability and recovery.
  • Collaborate on CUDA, PyTorch, kernel, and framework compatibility.
  • Create runbooks, capacity models, security controls, and reproducible docs.
  • Improve developer workflows for launching and debugging GPU jobs.
  • Surface cost and performance trade-offs to stakeholders.

Skills

GPU workloads
Linux
Containers
Networking
Storage
Infrastructure as code
Cloud platforms
Container orchestration
NVIDIA drivers/CUDA
PyTorch
Diagnostics

Tools

Kubernetes GPU scheduling
Slurm
Ray

Job description

Innomium is seeking a GPU and Cloud Infrastructure Engineer to build the compute foundation for model teams, covering training, evaluation, fine-tuning, and inference.

You will design GPU workloads, containerized environments, storage and networking paths, and automation to reproduce experiments and releases across programs, while collaborating with researchers to optimize utilization, memory, and cost.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

GPU & Cloud Infrastructure Engineer
GPU & Cloud Infrastructure Engineer

Innomium • Northern (KY)

Hybrid
USD 150,000 - 210,000
Remote-first policy
Annual learning budget
Company-provided equipment
+4
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
Remote Infra Operations Engineer - GPU Cloud
Remote Infra Operations Engineer - GPU Cloud

Nscale • Barstow (TX)

Remote
USD 80,000 - 110,000
Competitive pay with equity
Flexible workplace
Dynamic progression plan
Remote GPU Cloud Platform Engineer: Scale AI Compute
Remote GPU Cloud Platform Engineer: Scale AI Compute

Yotta Labs • United States

Remote
USD 120,000 - 160,000
Flexible remote work environment
Innovative team collaboration
Cutting-edge technology challenges
AI Infra Architect — GPU HPC & Cloud/On‑Prem
AI Infra Architect — GPU HPC & Cloud/On‑Prem

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)
Senior Full-Stack Engineer — AI Infra & GPU Cloud (Equity)

Nvidia Corporation • Santa Clara (CA)

On-site
USD 184,000 - 357,000
Staff Engineer, GPU Inference & Training Platform
Staff Engineer, GPU Inference & Training Platform

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud
Senior ML Engineer - GPU Inference & Large-Scale AI Cloud

Nebius • United States

Remote
USD 180,000 - 250,000
Remote GPU Cloud Infrastructure Engineer
Remote GPU Cloud Infrastructure Engineer

Runpod • United States

On-site
USD 120,000 - 180,000
Home Office & Equipment Stipend
Equity in company
Comprehensive health plans
+3