AI Infrastructure Operations Engineer – GPU Compute

Accenture

Bentonville (AR)

On-site

USD 120,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Accenture is seeking an experienced professional to join the Global AI Infrastructure team to design, deploy, and manage accelerated-computing infrastructure across on-prem, cloud, and hybrid environments. You will build and operate GPU clusters, automate provisioning, and optimize performance for AI training and inference at scale.

The role emphasizes governance, security, capacity planning, and reproducible platform operations, with travel 25-60% depending on client needs.

Qualifications

  • Minimum of 5+ years of experience designing, deploying, and managing accelerated-computing infrastructure across on-premises, cloud, and hybrid environments.

Responsibilities

  • Design and implement accelerated-computing infrastructure solutions aligned to system architecture, deployment roadmaps, performance, scalability, resiliency, and governance requirements.
  • Deploy, configure, and operate GPU-based clusters across bare-metal and containerized environments, using workload schedulers and Kubernetes orchestration to support AI training, inference, and high-performance compute workloads.
  • Integrate infrastructure platforms with enterprise systems, data platforms, security frameworks, service-management processes, and governance controls.
  • Design, build, and maintain reusable tools, scripts, self-service capabilities, and automation workflows for infrastructure operations, including provisioning, configuration management, validation, capacity planning, monitoring, incident management, reporting, and recurring remediation.
  • Establish repeatable operational processes for cluster provisioning, configuration management, patching, capacity planning, monitoring, incident response, and lifecycle management.
  • Perform and automate GPU, compute, storage, and network benchmarking and validation; diagnose performance issues across multi-node AI training, inference, and distributed compute workloads.
  • Develop and maintain architecture diagrams, configuration baselines, operational runbooks, and support documentation.
  • Provide technical guidance, troubleshooting, and optimization for GPU clusters supporting AI training, inference, high-performance computing, and multi-node simulation workloads.

Skills

GPU clusters
Cluster management
Kubernetes
AI infrastructure
Python

Education

Bachelor's degree

Tools

Terraform
Ansible
Python
Bash

Job description

Accenture is seeking an experienced professional to join the Global AI Infrastructure team to design, deploy, and manage accelerated-computing infrastructure across on-prem, cloud, and hybrid environments. You will build and operate GPU clusters, automate provisioning, and optimize performance for AI training and inference at scale.

The role emphasizes governance, security, capacity planning, and reproducible platform operations, with travel 25-60% depending on client needs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infrastructure Engineer: GPU-HPC Clusters & MLOps
AI Infrastructure Engineer: GPU-HPC Clusters & MLOps

Freelio • Northern (KY)

Hybrid
USD 90,000 - 230,000
AI Cloud Engineer: GPU Kubernetes & AI Infra
AI Cloud Engineer: GPU Kubernetes & AI Infra

Accenture • Hill Air Force Base (UT)

On-site
USD 106,000 - 206,000
AI & HPC Infra Engineer: GPU Compute & Cloud Automation
AI & HPC Infra Engineer: GPU Compute & Cloud Automation

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000
Senior AI Infra Architect GPU & Cloud
Senior AI Infra Architect GPU & Cloud

Nvidia Corporation in • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits
AI Infrastructure Architect — Scalable GPU Compute
AI Infrastructure Architect — Scalable GPU Compute

EngineersOfAI • Sunnyvale (CA)

On-site
USD 150,000 - 200,000
Senior AI Compute Infra Engineer (Hybrid)
Senior AI Compute Infra Engineer (Hybrid)

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Recruitment accommodations
Senior GPU Compute Solutions Architect
Senior GPU Compute Solutions Architect

Computacenter AG & Co. oHG • Northern (KY)

Hybrid
USD 190,000 - 230,000
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior AI Infrastructure Architect — Enterprise GPU Clusters
Senior AI Infrastructure Architect — Enterprise GPU Clusters

NVIDIA • California (MO)

On-site
USD 184,000 - 287,500
Equity
Benefits
Senior AI Infrastructure Engineer - GPU & Kubernetes
Senior AI Infrastructure Engineer - GPU & Kubernetes

HCL Technologies Limited • California (MO)

On-site
USD 120,000 - 180,000
401(k) retirement plan
Paid time off (PTO)
Paid holidays
+1