AI Infra Engineer: GPU Clusters, Automation & HPC

Accenture

Detroit (MI)

On-site

USD 110,000 - 210,000

Full time

6 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Medical benefits
Dental benefits
401(k) plan

Job summary

Accenture is seeking an experienced Senior AI Infrastructure Engineer to design and operate accelerated‑computing infrastructure for AI training and inference across on‑premises, cloud, and hybrid environments. You will deploy GPU clusters, manage scheduling with Kubernetes and Slurm, and build automation tools to ensure performance, resiliency, and scalability.

The role requires 5+ years on accelerated infrastructure, with hands‑on experience in GPUs, DPUs, and CPUs, plus proficiency in

Qualifications

  • Minimum of 5+ years of experience designing, deploying, and managing accelerated-computing infrastructure across on-premises, cloud, and hybrid environments.
  • Minimum of 5+ years of hands‑on experience with accelerated‑computing platforms, including GPUs, DPUs, and CPUs, high‑bandwidth network fabrics, and AI based storage architectures.
  • Minimum of 5+ years of experience with cluster management, workload scheduling, orchestration, observability, and infrastructure automation, including building operational tools and automation workflows with platforms such as Kubernetes, Slurm, Run:ai
  • Minimum 6 months hands‑on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting.
  • Bachelor's degree or equivalent (minimum 12 years) work experience.

Responsibilities

  • Design and implement accelerated-computing infrastructure solutions aligned to system architecture, deployment roadmaps, performance, scalability, resiliency, and governance requirements.
  • Deploy, configure, and operate GPU-based clusters across bare-metal and containerized environments, using workload schedulers and Kubernetes orchestration to support AI training, inference, and high-performance compute workloads.
  • Integrate infrastructure platforms with enterprise systems, data platforms, security frameworks, service-management processes, and governance controls.
  • Design, build, and maintain reusable tools, scripts, self-service capabilities, and automation workflows for infrastructure operations, including provisioning, configuration management, validation, capacity planning, monitoring, incident management, reporting, and recurring remediation.
  • Establish repeatable operational processes for cluster provisioning, configuration management, patching, capacity planning, monitoring, incident response, and lifecycle management.
  • Perform and automate GPU, compute, storage, and network benchmarking and validation; diagnose performance issues across multi-node AI training, inference, and distributed compute workloads.
  • Develop and maintain architecture diagrams, configuration baselines, operational runbooks, and support documentation.
  • Provide technical guidance, troubleshooting, and optimization for GPU clusters supporting AI training, inference, high-performance computing, and multi-node simulation workloads, with emphasis on availability, resiliency, scalability, energy efficiency, and cost management.

Skills

GPU clusters
Cluster management
Workload orchestration
Kubernetes
Slurm
Run:ai
Claude Code
Python
Bash scripting

Education

Bachelor's degree or equivalent (12 years)

Tools

Terraform
Ansible
Claude Code

Job description

Accenture is seeking an experienced Senior AI Infrastructure Engineer to design and operate accelerated‑computing infrastructure for AI training and inference across on‑premises, cloud, and hybrid environments. You will deploy GPU clusters, manage scheduling with Kubernetes and Slurm, and build automation tools to ensure performance, resiliency, and scalability.

The role requires 5+ years on accelerated infrastructure, with hands‑on experience in GPUs, DPUs, and CPUs, plus proficiency in

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: GPU Clusters & Automation
AI Infrastructure Engineer: GPU Clusters & Automation

Accenture • Irving (TX)

On-site
USD 90,000 - 260,000
Medical coverage
Dental coverage
Vision coverage
+3
AI Infra Ops Engineer: GPU Clusters & Automation
AI Infra Ops Engineer: GPU Clusters & Automation

Accenture • Carmel (IN)

Hybrid
USD 100,000 - 210,000
Medical, dental, vision insurance
Life insurance
Long-term disability
+4
AI Infra Operations Engineer - GPU Clusters and Automation
AI Infra Operations Engineer - GPU Clusters and Automation

Accenture • City of Albany (NY)

On-site
USD 87,000 - 266,000
Medical insurance
Dental insurance
Vision insurance
+6
AI Infrastructure & GPU Compute Engineer
AI Infrastructure & GPU Compute Engineer

Accenture • Seattle (WA)

On-site
USD 101,000 - 245,000
AI Infrastructure Ops Engineer: GPU Clusters, Hybrid Cloud
AI Infrastructure Ops Engineer: GPU Clusters, Hybrid Cloud

Accenture • Miami (FL)

On-site
USD 87,000 - 266,000
AI Infrastructure Engineer: GPU Clusters & Automation
AI Infrastructure Engineer: GPU Clusters & Automation

Socket.dev • Missouri

On-site
USD 87,000 - 266,000
Medical, dental, vision
401(k)
Paid holidays & time off
Senior AI Infrastructure & GPU Compute Engineer
Senior AI Infrastructure & GPU Compute Engineer

Accenture • Austin (TX)

On-site
USD 140,000 - 260,000
Medical benefits
401(k) program
Paid holidays & time off
AI Infrastructure Engineer: GPU Clusters & Hybrid Cloud
AI Infrastructure Engineer: GPU Clusters & Hybrid Cloud

Accenture • Redmond (WA)

On-site
USD 101,000 - 245,000
Senior GPU Compute Solutions Architect
Senior GPU Compute Solutions Architect

Computacenter AG & Co. oHG • Northern (KY)

Hybrid
USD 190,000 - 230,000
AI Infra Architect — GPU HPC & Cloud/On‑Prem
AI Infra Architect — GPU HPC & Cloud/On‑Prem

NVIDIA • California (MO)

On-site
USD 152,000 - 287,500
Equity
Benefits