AI Infrastructure Engineer — Cloud, GPU & MLOps

dicedemo

Boston (CT)

On-site

USD 130,000 - 170,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

dicedemo is seeking an AI Infrastructure Engineer to design, build, and scale the infrastructure that powers our AI and ML workloads. You will work across AI/ML, cloud infrastructure, distributed systems, and DevOps to ensure workloads run reliably and securely at scale.

You will partner with ML engineers, data scientists, and platform teams to deploy GPU-based training and inference environments, implement IaC, CI/CD, and monitoring, and optimize for performance, cost, and security.

Qualifications

  • 3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure.
  • Experience with at least one major cloud platform: AWS, Azure, or GCP.
  • Experience with Kubernetes and Docker.
  • Experience with Infrastructure-as-Code tools such as Terraform.
  • Strong scripting/programming in Python, Bash, Go, or similar.
  • Experience building CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or similar.
  • Knowledge of networking, Linux systems, distributed computing, and cloud architecture.
  • Experience implementing monitoring and observability solutions.
  • Understanding of machine learning development and deployment workflows.

Responsibilities

  • Design, build, and maintain scalable infrastructure for AI, machine learning, and Generative AI workloads
  • Build and manage cloud infrastructure across AWS, Azure, and/or Google Cloud Platform
  • Deploy and operate GPU-based compute environments for model training and inference
  • Design infrastructure supporting LLMs, model training, fine-tuning, inference, and AI applications
  • Build and manage containerized workloads using Docker and Kubernetes
  • Develop infrastructure-as-code using Terraform, CloudFormation, or Pulumi
  • Build CI/CD and MLOps pipelines supporting model development and deployment
  • Optimize GPU/CPU utilization, infrastructure performance, scalability, and cloud costs
  • Implement monitoring, logging, observability, and alerting for AI infrastructure and services
  • Support distributed training and high-performance computing environments
  • Build secure, highly available systems capable of supporting production AI workloads
  • Partner with ML Engineers and Data Scientists to move models from experimentation into production
  • Troubleshoot infrastructure, networking, performance, and deployment issues
  • Evaluate emerging AI infrastructure technologies and recommend improvements to the platform

Skills

Cloud Infrastructure
DevOps / SRE / MLOps
Python / Bash / Go

Tools

Kubernetes
Docker
Terraform
GitHub Actions

Job description

dicedemo is seeking an AI Infrastructure Engineer to design, build, and scale the infrastructure that powers our AI and ML workloads. You will work across AI/ML, cloud infrastructure, distributed systems, and DevOps to ensure workloads run reliably and securely at scale.

You will partner with ML engineers, data scientists, and platform teams to deploy GPU-based training and inference environments, implement IaC, CI/CD, and monitoring, and optimize for performance, cost, and security.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infra DevOps & Backend Engineer
AI Infra DevOps & Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

dicedemo • Boston (CT)

On-site
USD 130,000 - 170,000
AI Systems Engineer: Cloud, GPUs & Automation
AI Systems Engineer: Cloud, GPUs & Automation

MCI • United States

On-site
USD 90,000 - 130,000
Infra DevOps and Backend Engineer
Infra DevOps and Backend Engineer

GMI Cloud • Mountain View (CA)

On-site
USD 160,000 - 210,000
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior ML Engineer — Production AI & Scalable Systems
Senior ML Engineer — Production AI & Scalable Systems

dicedemo • Boston (AL)

On-site
USD 130,000 - 190,000
AI Infra Architect — GPU HPC & Cloud/On‑Prem
AI Infra Architect — GPU HPC & Cloud/On‑Prem

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm
AI Infra Engineer — GPU Cloud, Kubernetes/Slurm

Blue Signal Search • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual bonus
Equity participation
Comprehensive benefits
+1
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops
AI/ML Infrastructure Lead — GPU Cluster & Cloud Ops

Oracle • United States

On-site
USD 121,000 - 307,000
Medical, dental, and vision insurance
401(k) Savings and Investment Plan
Paid time off
Remote AI Infrastructure Engineer — Scale ML & Inference
Remote AI Infrastructure Engineer — Scale ML & Inference

Vantaca, LLC • Redwood City (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Medical, Dental, Vision