Senior AI Infra Engineer - GPU Clusters & HPC

Accenture

Overland Park (KS)

On-site

USD 120,000 - 230,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Accenture is seeking an experienced principal engineer to design and operate accelerated‑computing infrastructure for AI training, inference, and HPC workloads across on‑premises, cloud, and hybrid environments. You will deploy GPU clusters, orchestrate with Kubernetes and Run:AI, and automate provisioning, monitoring, and governance processes.

Join a team building scalable, secure infrastructure with emphasis on performance, energy efficiency, and cost management, while collaborating across the

Qualifications

  • Minimum of 5+ years designing, deploying, and managing accelerated‑computing infra across on‑prem, cloud, and hybrid environments.
  • Hands‑on with GPUs, DPUs, CPUs, high‑bandwidth fabrics, and AI storage like parallel file systems.
  • 5+ years in cluster management, scheduling, orchestration, observability, and automation (Kubernetes, Slurm, Run:ai).
  • 6+ months hands‑on with Claude Code, AI automation tools, Terraform, Ansible, Python and Bash.
  • Bachelor's degree or equivalent work experience (min 12 years).

Responsibilities

  • Design and implement accelerated‑computing infra aligned to system architecture, roadmaps, performance, and governance.
  • Deploy and operate GPU clusters across bare‑metal and containerized environments with schedulers.
  • Integrate infra platforms with enterprise systems, data platforms, security frameworks, and governance.
  • Develop reusable tools, scripts, and automation for provisioning, config, validation, and capacity planning.
  • Establish repeatable processes for cluster provisioning, patching, monitoring, and lifecycle management.
  • Benchmark and validate GPU, compute, storage, and network performance for AI workloads.
  • Create architecture diagrams, runbooks, and support documentation.
  • Provide guidance and optimization for GPU clusters with emphasis on availability, resiliency, energy efficiency, and cost.

Skills

Accelerated computing infra design
GPU cluster management
Kubernetes orchestration
Python scripting
Bash scripting
Observability & automation

Education

Bachelor's degree or equivalent work experience

Tools

Kubernetes
Slurm
Run:AI
Terraform
Ansible

Job description

Accenture is seeking an experienced principal engineer to design and operate accelerated‑computing infrastructure for AI training, inference, and HPC workloads across on‑premises, cloud, and hybrid environments. You will deploy GPU clusters, orchestrate with Kubernetes and Run:AI, and automate provisioning, monitoring, and governance processes.

Join a team building scalable, secure infrastructure with emphasis on performance, energy efficiency, and cost management, while collaborating across the

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Infra Ops Engineer: GPU Clusters & HPC
AI Infra Ops Engineer: GPU Clusters & HPC

Accenture • Mountain View (CA)

On-site
USD 94,000 - 266,000
AI Infra Operations Engineer: GPU Clusters & Automation
AI Infra Operations Engineer: GPU Clusters & Automation

Accenture • Oklahoma City (OK)

On-site
USD 120,000 - 210,000
AI Infrastructure Engineer - GPU Clusters & Automation
AI Infrastructure Engineer - GPU Clusters & Automation

Accenture • Walnut Creek (CA)

On-site
USD 94,000 - 266,000
Medical, dental, vision
401(k)
Bonus opportunities
+2
AI Infra Operations Engineer: GPU Clusters & Automation
AI Infra Operations Engineer: GPU Clusters & Automation

Accenture • New York (NY)

On-site
USD 87,000 - 266,000
Medical insurance
Dental insurance
Vision insurance
+6
AI Infrastructure Engineer: GPU Clusters & Automation
AI Infrastructure Engineer: GPU Clusters & Automation

Accenture • San Francisco (CA)

On-site
USD 94,000 - 266,000
Medical benefits
Dental & Vision insurance
Life insurance
+3
Senior AI Infra Engineer — GPU Clusters & Automation
Senior AI Infra Engineer — GPU Clusters & Automation

Accenture • Minneapolis (MN)

On-site
USD 94,000 - 230,000
Senior AI Infrastructure & HPC Architect
Senior AI Infrastructure & HPC Architect

Accenture • Saint Petersburg (FL)

Hybrid
USD 140,000 - 230,000
AI Infra Ops Engineer: GPU Clusters & Automation
AI Infra Ops Engineer: GPU Clusters & Automation

Accenture • Saint Petersburg (FL)

On-site
USD 120,000 - 260,000
AI Infrastructure Engineer: HPC, GPUs & Agentic AI
AI Infrastructure Engineer: HPC, GPUs & Agentic AI

Accenture • Walnut Creek (CA)

On-site
USD 94,000 - 266,000
Senior AI Compute Infra Engineer (Hybrid)
Senior AI Compute Infra Engineer (Hybrid)

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Recruitment accommodations