AI Infrastructure / ML Infrastructure Engineer

DeWinter Group

Campbell (CA)

On-site

USD 68,880 - 241,080

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A technology consulting firm is looking for an AI Infrastructure / ML Infrastructure Engineer for a 12-month contract. This remote position involves managing and optimizing high-performance GPU clusters for AI applications. Candidates should have over 5 years of Cloud Infrastructure or DevOps experience, along with expertise in AWS, Kubernetes, and GPU orchestration. Ideal candidates can work autonomously and are expected to deliver results quickly in a collaborative environment.

Qualifications

  • 5+ years of experience in Cloud Infrastructure or DevOps.
  • Deep expertise in AWS/Azure/GCP, Kubernetes (EKS/GKE), and GPU orchestration.
  • Demonstrated ability to work autonomously and manage time effectively.
  • Experience with Terraform, Docker, and monitoring tools.

Responsibilities

  • Provisioning and managing high-performance GPU clusters using Terraform or CloudFormation.
  • Building and maintaining the internal 'Model Hub' for AI models.
  • Optimizing networking and storage for multi-node distributed training.
  • Implementing autoscaling logic for managing inference costs.
  • Designing high-availability infrastructure for AI applications.

Skills

Cloud Infrastructure
DevOps
AWS
Azure
GCP
Kubernetes
GPU orchestration
Terraform
Docker
Prometheus
Grafana

Job description

Title: AI Infrastructure / ML Infrastructure Engineer

Job Type: Contract

Contract Length: 12 Months

Pay Range: $50/hr – $175/hr

Start Date: ASAP

Location: Remote

About the Opportunity

Our client, a leader in AI testing, is looking for a skilled AI Infrastructure / ML Infrastructure Engineer to join their team for a 12‑month engagement. This project involves provisioning, managing, and optimizing high-performance GPU clusters and infrastructure to support mission‑critical AI applications. This is a high‑impact role that requires a self‑motivated professional who can hit the ground running and deliver results quickly.

Key Responsibilities & Deliverables
  • Provisioning and managing high-performance GPU clusters using Terraform or CloudFormation.
  • Building and maintaining the internal "Model Hub" for versioning and deploying AI models across the company.
  • Optimizing the networking and storage layers to support multi‑node distributed training.
  • Implementing autoscaling logic to manage inference costs while meeting peak user demand.
  • Designing high‑availability infrastructure for mission‑critical AI applications.
Required Skills & Experience
  • 5+ years of experience in Cloud Infrastructure or DevOps.
  • Deep expertise in AWS/Azure/GCP, Kubernetes (EKS/GKE), and GPU orchestration. This isn't a learning role—you need to be a subject matter expert.
  • Demonstrated ability to work autonomously and manage your own time effectively to meet project goals.
  • Experience with Terraform, Docker, and monitoring tools such as Prometheus and Grafana.
  • Strong communication skills to provide clear and concise status updates to the project team.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Infra Engineer — GPU Clusters & Model Hub
Remote AI Infra Engineer — GPU Clusters & Model Hub

DeWinter Group • Campbell (CA)

Remote
Machine Learning Engineer / AI Engineer
Machine Learning Engineer / AI Engineer

DeWinter Group • Campbell (CA)

On-site
USD 68,880 - 241,080
AI/ML Infra Engineer - Hosting
AI/ML Infra Engineer - Hosting

Hamilton Barnes Associates Limited • San Francisco (CA)

On-site
USD 225,000 - 275,000
Stock options
ML Infrastructure Engineer
ML Infrastructure Engineer

Strativ Group • Menlo Park (CA)

On-site
USD 250,000 - 320,000
Solutions Architect – AI / ML
Solutions Architect – AI / ML

DeWinter Group • Campbell (CA)

On-site
Solution Architect - AI Infrastructure
Solution Architect - AI Infrastructure

Hamilton Barnes Associates Limited • Town of Texas (WI)

On-site
USD 283,500 - 346,500
Equity (RSUs)
AI Infrastructure Lead
AI Infrastructure Lead

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Remote ML Engineer — Production AI/ML, 12-Month Contract
Remote ML Engineer — Production AI/ML, 12-Month Contract

DeWinter Group • Campbell (CA)

Remote
USD 68,880 - 241,080
MLOps / AI Ops Engineer
MLOps / AI Ops Engineer

DeWinter Group • Campbell (CA)

Remote
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2