GPU & Cloud Infrastructure Engineer

Innomium

Northern (KY)

Hybrid

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Remote-first policy
Annual learning budget
Company-provided equipment
Internet/home-office stipend
Paid time off
Equity/bonuses where applicable
Retirement contributions

Job summary

Innomium is seeking a GPU and Cloud Infrastructure Engineer to build the compute foundation for model teams, covering training, evaluation, fine-tuning, and inference.

You will design GPU workloads, containerized environments, storage and networking paths, and automation to reproduce experiments and releases across programs, while collaborating with researchers to optimize utilization, memory, and cost.

Qualifications

  • Experience with cloud infrastructure for GPU workloads.
  • Strong Linux, containers, networking, storage.
  • IaC fundamentals across cloud platforms.
  • Ability to diagnose failures across infra layers and document work.

Responsibilities

  • Build and operate GPU training, evaluation, fine-tuning, and inference environments.
  • Automate provisioning, container images, dependencies, secrets, networking, and storage.
  • Profile utilization, memory, throughput, queue behavior, data transfer, and cost.
  • Design model-serving and batch-execution paths with observability and recovery.
  • Collaborate on CUDA, PyTorch, kernel, and framework compatibility.
  • Create runbooks, capacity models, security controls, and reproducible docs.
  • Improve developer workflows for launching and debugging GPU jobs.
  • Surface cost and performance trade-offs to stakeholders.

Skills

GPU workloads
Linux
Containers
Networking
Storage
Infrastructure as code
Cloud platforms
Container orchestration
NVIDIA drivers/CUDA
PyTorch
Diagnostics

Tools

Kubernetes GPU scheduling
Slurm
Ray

Job description

Engineer secure, efficient GPU and cloud environments for training, evaluation, fine-tuning, and inference across Innomium programs.

Innomium is an applied AI research and engineering company that turns ambitious technical ideas into dependable, production-ready systems.

We bring together AI research, product engineering, data, cloud infrastructure, evaluation, and operational delivery within one accountable program. Our teams work with startups, product companies, and enterprises to build custom AI models, software products, deployment pipelines, integrations, and reproducible evaluation systems.

Our work spans language models, AI agents, computer vision, retrieval systems, cloud and edge deployments, open research releases, and engineering contributions. Through Innomium Arena, we also create structured opportunities for builders to contribute to challenging technical projects. Through Innomium Compute, we provide on-demand GPU capacity for training and inference.

We focus on measurable outcomes, inspectable evidence, and software that teams can operate and improve—not prototypes that stop at the demonstration stage.

The Role

As a GPU and Cloud Infrastructure Engineer, you build the compute foundation that model and product teams need to be fast, observable under load, and disciplined about cost and security.

You will design and operate GPU workloads, containerized environments, storage and network paths, schedulers, model-serving infrastructure, and the automation required to reproduce experiments and releases—across Innomium Agency programs and the wider Compute product.

You will partner directly with researchers and AI engineers to understand the execution path rather than treating workloads as anonymous jobs. You will also make trade-offs visible: utilization, queue time, memory, throughput, data movement, image provenance, dependency compatibility, and cost per useful result.

What Strong Performance Looks Like

Researchers can launch reproducible work without manually rebuilding environments. Serving paths are benchmarked and observable, and infrastructure failures are diagnosable. Capacity decisions are supported by evidence instead of intuition.

You write infrastructure code, debug drivers and containers, improve developer workflows, and document operating procedures for others.

Over time, you raise the reliability and efficiency of Innomium’s GPU platform and reduce the operational friction around training and inference.

How We Work

Innomium operates through small, accountable teams with direct access to the technical problem.

We value:

  • Clear ownership and reliable execution.
  • Written decisions and reviewable technical reasoning.
  • Measurable acceptance criteria.
  • Honest communication about risks and limitations.
  • Practical solutions over unnecessary complexity.
  • Documentation and handover from the beginning of a project.
  • Engineering decisions connected to user and operating outcomes.

Remote collaboration requires dependable communication, thoughtful handoffs, and agreed working-hour overlap with the relevant delivery team.

Compensation and Benefits

Compensation range: $150,000–$210,000 USD (base), depending on experience, location, and engagement type. Total compensation may include performance-based bonuses or equity participation where applicable.

Employment arrangement: Full-time

Location and working hours: Remote. United States preferred; international candidates are considered subject to work authorization, contracting or employment availability, and required overlap with team working hours.

Health and wellness: Medical, dental, and vision coverage (or equivalent stipend for international contractors), plus access to mental health and wellness support programs.

Paid time off: Flexible paid time off policy, including vacation, sick leave, and company holidays. Parental leave provided in accordance with local regulations and role type.

Professional development: Annual learning and development budget for courses, certifications, books, and conferences. Support for attending relevant industry events and technical communities.

Equipment and remote-work support: Company-provided laptop and necessary development equipment. Monthly stipend for internet and home-office setup where applicable. Access to required software and cloud tools.

Additional benefits: Retirement or pension contributions where applicable, remote-first flexibility, and potential performance-based bonuses or equity participation depending on role and engagement type.

What You Will Own

The work this role is expected to own.

  • Build and operate GPU training, evaluation, fine-tuning, and inference environments.
  • Automate provisioning, container images, dependency pinning, secrets, networking, and storage.
  • Profile utilization, memory, throughput, queue behavior, data transfer, and workload cost.
  • Design model-serving and batch-execution paths with observability and failure recovery.
  • Collaborate on CUDA, PyTorch, kernel, and framework compatibility across hardware.
  • Create runbooks, capacity models, security controls, and reproducible environment documentation.
  • Improve developer and researcher workflows for launching and debugging GPU jobs.
  • Surface cost and performance trade-offs clearly to technical and product stakeholders.
Required Qualifications

Capabilities and experience that support success in this role.

  • Professional cloud or infrastructure engineering experience with GPU workloads.
  • Strong Linux, containers, networking, storage, and infrastructure-as-code fundamentals.
  • Experience with one or more major cloud platforms and container orchestration.
  • Practical understanding of NVIDIA drivers, CUDA environments, PyTorch workloads, and GPU profiling.
  • Ability to diagnose failures across application, container, node, network, and storage layers.
  • Strong automation and technical-documentation habits.
  • Ability to work effectively in a remote environment with autonomy and accountability.
Preferred Qualifications

Valuable adjacent experience, but not a substitute for the core requirements.

  • Experience with Kubernetes GPU scheduling, Slurm, Ray, or distributed-training stacks.
  • Experience operating model servers or high-throughput inference systems.
  • Knowledge of Triton kernels, NCCL, topology, or multi-node training.
  • Experience with GPU rental platforms, quota systems, or multi-tenant compute products.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote GPU & Cloud Infrastructure Engineer for AI
Remote GPU & Cloud Infrastructure Engineer for AI

Innomium • Northern (KY)

Hybrid
USD 150,000 - 210,000
Remote-first policy
Annual learning budget
Company-provided equipment
+4
Senior Product Engineer — Full Stack
Senior Product Engineer — Full Stack

Innomium • Northern (KY)

Hybrid
USD 150,000 - 200,000
Health and wellness
Paid time off
Professional development
+3
Member of Technical Staff (Software Engineer, Inference & Training Platform)
Member of Technical Staff (Software Engineer, Inference & Training Platform)

United States Digital Space LLC • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
Cluster Engineer
Cluster Engineer

STN Inc • San Francisco (CA)

On-site
USD 180,000 - 240,000
Machine Learning Engineer — Computer Vision
Machine Learning Engineer — Computer Vision

Innomium • Northern (KY)

Hybrid
USD 150,000 - 210,000
Remote-first flexibility
Health and wellness stipends
Annual learning budget
+2
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure
Member of Technical Staff – Software Engineer, GPU Cluster Infrastructure

Perplexity • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • San Francisco (CA)

On-site
USD 150,000 - 300,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity • New York (NY)

On-site
USD 250,000 - 485,000
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

United States Digital Space LLC • San Francisco (CA)

On-site
USD 180,000 - 260,000