AI Infrastructure Engineer - GPU Clusters & Automation

Accenture

Walnut Creek (CA)

On-site

USD 94,000 - 266,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Medical, dental, vision
401(k)
Bonus opportunities
Paid holidays
Paid time off

Job summary

Accenture’s Global AI Infrastructure team seeks experts to design and operate accelerated-computing infrastructure across on‑prem, cloud, and hybrid environments. You’ll deploy GPU clusters, integrate with security and governance, and build reusable automation tools to run AI training and HPC workloads.

The role emphasizes scalability, reliability, and cost management, with extensive collaboration across the technology ecosystem. Travel may be required 25%–60% depending on client needs.

Qualifications

  • Minimum of 5+ years designing, deploying, and managing accelerated-computing infra across on‑prem, cloud, and hybrid environments for large clients.
  • Minimum of 5+ years hands‑on experience with GPUs, DPUs, CPUs, high‑bandwidth networks, and AI storage architectures.
  • Minimum of 5+ years experience with cluster management, workload scheduling, orchestration, observability, and automation (Kubernetes, Slurm, Run:ai).
  • Minimum 6 months hands‑on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting.
  • Bachelor's degree or equivalent (minimum 12 years) work experience.

Responsibilities

  • Design and implement accelerated‑computing infra solutions aligned to architecture, roadmaps, performance, scalability, resiliency, and governance.
  • Deploy, configure, and operate GPU clusters across bare‑metal and containerized environments using schedulers and Kubernetes for AI workloads.
  • Integrate infra platforms with enterprise systems, data platforms, security and governance controls.
  • Build reusable tools, scripts, self‑service capabilities, and automation workflows for provisioning, configuration, validation, capacity planning, monitoring, and incident management.
  • Establish repeatable processes for cluster provisioning, patching, and lifecycle management.
  • Perform and automate GPU/compute/storage/network benchmarking and diagnose performance issues.
  • Develop architecture diagrams, runbooks, and support documentation.
  • Provide troubleshooting and optimization for GPU clusters with emphasis on availability, resiliency, energy efficiency, and cost management.

Skills

Infra design
Kubernetes
Slurm
Run:ai
Python
Bash scripting
Terraform
Ansible
GPU clustering

Education

Bachelor's degree or equivalent (minimum 12 years)

Tools

Base Command Manager (BCM)
NGC
NCCL
CUDA-X
NVAIE
Dynamo
OpenAPI
JSON/YAML

Job description

Accenture’s Global AI Infrastructure team seeks experts to design and operate accelerated-computing infrastructure across on‑prem, cloud, and hybrid environments. You’ll deploy GPU clusters, integrate with security and governance, and build reusable automation tools to run AI training and HPC workloads.

The role emphasizes scalability, reliability, and cost management, with extensive collaboration across the technology ecosystem. Travel may be required 25%–60% depending on client needs.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Operations Engineer: GPU Clusters & Automation
AI Infra Operations Engineer: GPU Clusters & Automation

Accenture • New York (NY)

On-site
USD 87,000 - 266,000
Medical insurance
Dental insurance
Vision insurance
+6
AI Infrastructure Engineer: GPU Clusters & Automation
AI Infrastructure Engineer: GPU Clusters & Automation

Accenture • San Francisco (CA)

On-site
USD 94,000 - 266,000
Medical benefits
Dental & Vision insurance
Life insurance
+3
AI Infra Operations Engineer: GPU Clusters & Automation
AI Infra Operations Engineer: GPU Clusters & Automation

Accenture • Oklahoma City (OK)

On-site
USD 120,000 - 210,000
AI Infra Ops Engineer: GPU Clusters & HPC
AI Infra Ops Engineer: GPU Clusters & HPC

Accenture • Mountain View (CA)

On-site
USD 94,000 - 266,000
Senior AI Infra Engineer - GPU Clusters & HPC
Senior AI Infra Engineer - GPU Clusters & HPC

Accenture • Overland Park (KS)

On-site
USD 120,000 - 230,000
Senior AI Infra Engineer — GPU Clusters & Automation
Senior AI Infra Engineer — GPU Clusters & Automation

Accenture • Minneapolis (MN)

On-site
USD 94,000 - 230,000
AI Infra Ops Engineer: GPU Clusters & Automation
AI Infra Ops Engineer: GPU Clusters & Automation

Accenture • Saint Petersburg (FL)

On-site
USD 120,000 - 260,000
AI Infrastructure Engineer: HPC, GPUs & Agentic AI
AI Infrastructure Engineer: HPC, GPUs & Agentic AI

Accenture • Walnut Creek (CA)

On-site
USD 94,000 - 266,000
Senior AI Infrastructure & HPC Architect
Senior AI Infrastructure & HPC Architect

Accenture • Saint Petersburg (FL)

Hybrid
USD 140,000 - 230,000
Senior AI Infra Architect GPU & Cloud
Senior AI Infra Architect GPU & Cloud

Nvidia Corporation in • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits