AI Infrastructure Engineer

dicedemo

Boston (CT)

On-site

USD 130,000 - 170,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

dicedemo is seeking an AI Infrastructure Engineer to design, build, and scale the infrastructure that powers our AI and ML workloads. You will work across AI/ML, cloud infrastructure, distributed systems, and DevOps to ensure workloads run reliably and securely at scale.

You will partner with ML engineers, data scientists, and platform teams to deploy GPU-based training and inference environments, implement IaC, CI/CD, and monitoring, and optimize for performance, cost, and security.

Qualifications

  • 3+ years of experience in Cloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure.
  • Experience with at least one major cloud platform: AWS, Azure, or GCP.
  • Experience with Kubernetes and Docker.
  • Experience with Infrastructure-as-Code tools such as Terraform.
  • Strong scripting/programming in Python, Bash, Go, or similar.
  • Experience building CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or similar.
  • Knowledge of networking, Linux systems, distributed computing, and cloud architecture.
  • Experience implementing monitoring and observability solutions.
  • Understanding of machine learning development and deployment workflows.

Responsibilities

  • Design, build, and maintain scalable infrastructure for AI, machine learning, and Generative AI workloads
  • Build and manage cloud infrastructure across AWS, Azure, and/or Google Cloud Platform
  • Deploy and operate GPU-based compute environments for model training and inference
  • Design infrastructure supporting LLMs, model training, fine-tuning, inference, and AI applications
  • Build and manage containerized workloads using Docker and Kubernetes
  • Develop infrastructure-as-code using Terraform, CloudFormation, or Pulumi
  • Build CI/CD and MLOps pipelines supporting model development and deployment
  • Optimize GPU/CPU utilization, infrastructure performance, scalability, and cloud costs
  • Implement monitoring, logging, observability, and alerting for AI infrastructure and services
  • Support distributed training and high-performance computing environments
  • Build secure, highly available systems capable of supporting production AI workloads
  • Partner with ML Engineers and Data Scientists to move models from experimentation into production
  • Troubleshoot infrastructure, networking, performance, and deployment issues
  • Evaluate emerging AI infrastructure technologies and recommend improvements to the platform

Skills

Cloud Infrastructure
DevOps / SRE / MLOps
Python / Bash / Go

Tools

Kubernetes
Docker
Terraform
GitHub Actions

Job description

AI Infrastructure Engineer
Position Overview

We are seeking anAI Infrastructure Engineer to design, build, and scale the infrastructure that powers our artificial intelligence and machine learning workloads. This role sits at the intersection ofAI/ML, cloud infrastructure, distributed systems, and DevOps/MLOps.

The ideal candidate has experience building highly available, scalable infrastructure for training, deploying, and operating machine learning and generative AI applications. You will partner closely with Machine Learning Engineers, Data Scientists, Software Engineers, and Platform Engineering teams to ensure AI workloads can run reliably, securely, and efficiently at scale.

Key Responsibilities
  • Design, build, and maintain scalable infrastructure forAI, machine learning, and Generative AI workloads
  • Build and manage cloud infrastructure acrossAWS, Azure, and/or Google Cloud Platform
  • Deploy and operate GPU-based compute environments for model training and inference
  • Design infrastructure supportingLLMs, model training, fine-tuning, inference, and AI applications
  • Build and manage containerized workloads usingDocker and Kubernetes
  • Develop infrastructure-as-code using tools such asTerraform, CloudFormation, or Pulumi
  • Build CI/CD and MLOps pipelines supporting model development and deployment
  • Optimize GPU/CPU utilization, infrastructure performance, scalability, and cloud costs
  • Implement monitoring, logging, observability, and alerting for AI infrastructure and services
  • Support distributed training and high-performance computing environments
  • Build secure, highly available systems capable of supporting production AI workloads
  • Partner with ML Engineers and Data Scientists to move models from experimentation into production
  • Troubleshoot infrastructure, networking, performance, and deployment issues
  • Evaluate emerging AI infrastructure technologies and recommend improvements to the platform
Required Qualifications
  • 3+ years of experience inCloud Infrastructure, DevOps, Platform Engineering, SRE, MLOps, or AI/ML Infrastructure
  • Strong experience with at least one major cloud platform:AWS, Azure, or GCP
  • Experience withKubernetes and Docker
  • Experience with Infrastructure-as-Code tools such asTerraform
  • Strong scripting/programming skills inPython, Bash, Go, or similar languages
  • Experience building CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or similar
  • Knowledge of networking, Linux systems, distributed computing, and cloud architecture
  • Experience implementing monitoring and observability solutions
  • Understanding of machine learning development and deployment workflows
Preferred Qualifications
  • Experience managingGPU infrastructure, including NVIDIA GPUs and CUDA environments
  • Experience with AI/ML frameworks such asPyTorch, TensorFlow, JAX, or Hugging Face
  • Experience supportingLLM training, fine-tuning, RAG, or inference workloads
  • Experience with MLOps platforms such asMLflow, Kubeflow, SageMaker, Vertex AI, or Azure Machine Learning
  • Experience with distributed training technologies such asRay, DeepSpeed, PyTorch Distributed, or Horovod
  • Familiarity with AI inference technologies such asvLLM, NVIDIA Triton, or TensorRT
  • Experience managing Kubernetes-based GPU clusters
  • Understanding of model serving, vector databases, and modern Generative AI architecture
  • Experience optimizing infrastructure for performance and cloud/GPU cost efficiency
What Success Looks Like

In this role, you will help create the infrastructure foundation that allows AI teams toexperiment faster, train models efficiently, deploy AI applications reliably, and scale them into production. You will reduce friction between AI development and production while improving reliability, performance, security, and infrastructure cost.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Platform Engineer
AI Platform Engineer

Park Place Technologies in • Highland Heights (OH)

On-site
USD 90,000 - 130,000
Staff Software Engineer (AI Infrastructure)
Staff Software Engineer (AI Infrastructure)

DeepRec.ai • Palo Alto (CA)

On-site
USD 180,000 - 320,000
AI Infrastructure & Platform Engineer
AI Infrastructure & Platform Engineer

International Materials, LLC • Delray Beach (FL), Northern (KY)

Hybrid
USD 120,000 - 180,000
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)
AI/ML Engineer – (Next-Generation AI Platforms & Workloads)

VeeAR Projects Inc. • Sunnyvale (CA)

On-site
USD 140,000 - 210,000
AI Engineer
AI Engineer

Compunnel, Inc. • San Francisco (CA)

On-site
USD 120,000 - 160,000
Sr AI Platform Engineer
Sr AI Platform Engineer

BravoTECH • Richardson (TX)

On-site
USD 170,000 - 250,000
AI Infrastructure / ML Infrastructure Engineer
AI Infrastructure / ML Infrastructure Engineer

DeWinter Group • Campbell (CA)

On-site
Senior AI DevOps Engineer (AI Ops / Platform Engineering)
Senior AI DevOps Engineer (AI Ops / Platform Engineering)

DeepCamp • Tucker (GA)

On-site
USD 96,000 - 165,000
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Senior Developer
Senior Developer

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000