Senior AI Infra Engineer - GPU Clusters & HPC Ops

Accenture

Boston (MA)

On-site

USD 94,000 - 245,000

Full time

8 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Accenture is seeking an experienced Global AI Infrastructure engineer to design and operate large-scale GPU clusters across on-prem, cloud, and hybrid environments. You will implement automated tooling, improve performance, and ensure governance and cost efficient operations.

Travel 25-60% may be required; 5+ years with GPUs, Kubernetes, and automation is expected. This role is based in Massachusetts with competitive benefits and career growth opportunities.

Qualifications

  • Minimum 5+ years designing, deploying accelerated infrastructure across on-prem, cloud, and hybrid.
  • Hands-on with GPUs/DPUs, high-bandwidth networks, AI storage architectures.
  • 5+ years with cluster management, orchestration, automation (Kubernetes, Slurm, Run:AI).

Responsibilities

  • Design and implement accelerated-computing infrastructure aligned to architecture and governance.
  • Deploy GPU-based clusters across bare-metal and containerized environments using Kubernetes.
  • Integrate infrastructure with security, data platforms, and operations processes.
  • Build reusable tools and automation workflows for provisioning, monitoring, and remediation.
  • Establish repeatable processes for provisioning, patching, and lifecycle management.
  • Benchmark and diagnose performance across multi-node AI training workloads.
  • Develop architecture diagrams, runbooks, and support documentation.
  • Provide guidance and optimization for GPU clusters supporting AI training/inference.

Skills

Kubernetes
Python
Terraform
Ansible
Slurm
Bash
GPU clusters
Run:ai

Education

Bachelor's degree or equivalent (minimum 12 years) work experience

Tools

Kubernetes
Slurm
Terraform
Ansible
Python
Bash
NVIDIA tools (NGC, BCM)

Job description

Accenture is seeking an experienced Global AI Infrastructure engineer to design and operate large-scale GPU clusters across on-prem, cloud, and hybrid environments. You will implement automated tooling, improve performance, and ensure governance and cost efficient operations.

Travel 25-60% may be required; 5+ years with GPUs, Kubernetes, and automation is expected. This role is based in Massachusetts with competitive benefits and career growth opportunities.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Infra Architect GPU & Cloud
Senior AI Infra Architect GPU & Cloud

Nvidia Corporation in • Washington

On-site
USD 184,000 - 357,000
Equity
Benefits
Senior AI Infrastructure Lead - GPU Clusters & Model Serving
Senior AI Infrastructure Lead - GPU Clusters & Model Serving

Outsourceit • San Francisco (CA)

On-site
USD 120,000 - 170,000
Senior AI Compute Infra Engineer (Hybrid)
Senior AI Compute Infra Engineer (Hybrid)

Arm Limited • Seattle (WA)

Hybrid
USD 209,000 - 283,000
Relocation package
Recruitment accommodations
Staff Compute Infra Engineer - GPU & AI Systems
Staff Compute Infra Engineer - GPU & AI Systems

xAI • Palo Alto (CA)

On-site
USD 180,000 - 440,000
AI Infrastructure Engineer - GPU & Kubernetes
AI Infrastructure Engineer - GPU & Kubernetes

HCLTech • California (MO)

On-site
USD 150,000 - 210,000
Medical Insurance
Dental Insurance
Vision Insurance
+2
Senior AI Infrastructure Engineer - GPU Compute
Senior AI Infrastructure Engineer - GPU Compute

Unchain Data • United States

On-site
USD 120,000 - 160,000
Senior GPU Compute Solutions Architect
Senior GPU Compute Solutions Architect

Computacenter AG & Co. oHG • Northern (KY)

Hybrid
USD 190,000 - 230,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote
Senior AI Infrastructure Engineer — Scale GPU Clusters Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 280,000 - 420,000
Equity options
Health, vision, dental benefits
Unlimited PTO
+2
AI & HPC Infra Engineer: GPU Compute & Cloud Automation
AI & HPC Infra Engineer: GPU Compute & Cloud Automation

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000