Remote AI Infra Engineer — GPU Clusters & Model Hub

DeWinter Group

Campbell (CA)

Remote

USD 68,880 - 241,080

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology consulting firm is looking for an AI Infrastructure / ML Infrastructure Engineer for a 12-month contract. This remote position involves managing and optimizing high-performance GPU clusters for AI applications. Candidates should have over 5 years of Cloud Infrastructure or DevOps experience, along with expertise in AWS, Kubernetes, and GPU orchestration. Ideal candidates can work autonomously and are expected to deliver results quickly in a collaborative environment.

Qualifications

  • 5+ years of experience in Cloud Infrastructure or DevOps.
  • Deep expertise in AWS/Azure/GCP, Kubernetes (EKS/GKE), and GPU orchestration.
  • Demonstrated ability to work autonomously and manage time effectively.
  • Experience with Terraform, Docker, and monitoring tools.

Responsibilities

  • Provisioning and managing high-performance GPU clusters using Terraform or CloudFormation.
  • Building and maintaining the internal 'Model Hub' for AI models.
  • Optimizing networking and storage for multi-node distributed training.
  • Implementing autoscaling logic for managing inference costs.
  • Designing high-availability infrastructure for AI applications.

Skills

Cloud Infrastructure
DevOps
AWS
Azure
GCP
Kubernetes
GPU orchestration
Terraform
Docker
Prometheus
Grafana

Job description

Title:AI Infrastructure / ML Infrastructure Engineer
Job Type:Contract
Contract Length:12 Months
Pay Range:$50/hr – $175/hr
Start Date:ASAP
Location:Remote

About the Opportunity:

Our client, a leader in AI testing, is looking for a skilled AI Infrastructure / ML Infrastructure Engineerto join their team for a 12-month engagement. This project involves provisioning, managing, and optimizing high-performance GPU clusters and infrastructure to support mission-critical AI applications. This is a high-impact role that requires a self-motivated professional who can hit the ground running and deliver results quickly.

Key Responsibilities & Deliverables:

This role is focused on the successful completion of specific tasks and deliverables. Your responsibilities will include:

  • Provisioning and managing high-performance GPU clusters using Terraform or CloudFormation.
  • Building and maintaining the internal "Model Hub" for versioning and deploying AI models across the company.
  • Optimizing the networking and storage layers to support multi-node distributed training.
  • Implementing autoscaling logic to manage inference costs while meeting peak user demand.
  • Designing high-availability infrastructure for mission-critical AI applications.
Required Skills & Experience:

We are looking for someone with a proven track record of successful contract engagements. The ideal candidate will have:
  • 5+ years of experience in Cloud Infrastructure or DevOps.
  • Deep expertise in AWS/Azure/GCP, Kubernetes (EKS/GKE), and GPU orchestration. This isn't a learning role—you need to be a subject matter expert.
  • Demonstrated ability to work autonomously and manage your own time effectively to meet project goals.
  • Experience with Terraform, Docker, and monitoring tools like Prometheus/Grafana.
  • Strong communication skills to provide clear and concise status updates to the project team.
#LI-NK1
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure / ML Infrastructure Engineer
AI Infrastructure / ML Infrastructure Engineer

DeWinter Group • Campbell (CA)

On-site
Remote ML Engineer — Production AI/ML, 12-Month Contract
Remote ML Engineer — Production AI/ML, 12-Month Contract

DeWinter Group • Campbell (CA)

Remote
USD 68,880 - 241,080
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Software Engineer, AI Infrastructure
Software Engineer, AI Infrastructure

Harell Data • Palo Alto (CA)

On-site
USD 180,000 - 260,000
AI Infrastructure Lead
AI Infrastructure Lead

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Remote AI Infrastructure Engineer: GPU Clusters & MLOps
Remote AI Infrastructure Engineer: GPU Clusters & MLOps

Bonfirevc • United States

On-site
USD 120,000 - 150,000
Competitive salary
Stock options
Health/dental/vision insurance
+2
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options