Technology Platform Engineer

Accenture PLC

Bengaluru

On-site

INR 3,500,000 - 5,500,000

Full time

2 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Accenture PLC is seeking a Technology Platform Engineer to design and deploy HPC and AI infrastructure across on-premises, cloud, and hybrid environments. You will build scalable clusters and manage XPU-based systems.

You will optimize performance, ensure security, and collaborate with analysts and engineers to run large-scale models and simulations. Strong Linux and CUDA GPU expertise required; cloud and DevOps skills preferred.

Qualifications

  • Experience in enterprise-wide HPC strategy and architecting next-gen supercomputing environments across the full stack.
  • Primary Skills: Linux Administrator/Architect, Advanced CUDA/GPU on H100/A100, HPC cluster design (SLURM/PBS Pro)
  • Secondary Skills: Parallel programming: MPI, OpenMP, CUDA, SYCL, NVIDIA DGX SuperPOD & BCM, InfiniBand NDR 400Gb/s & RoCEv2 design, MLOps & HPC integration (Kubeflow), Containerization: Singularity, Kubernetes, DevOps: Ansible, Terraform for HPC, UFM management, RunAI, Azure ML integration: distributed training, MLflow, Terraform / Bicep IaC for Azure HPC, FinOps: Reserved Instances, Spot VM strategies, Hybrid cloud HPC: on-prem to Azure/AWS burst

Responsibilities

  • Design and implement HPC and AI infrastructure solutions, aligning system architecture and deployment roadmaps to industry-specific performance and scalability needs
  • Deploy, configure, and manage XPU-based clusters (CPU/GPU/accelerators) using schedulers, VM/K8s orchestration platforms, Slurm, and containerized platforms in scalable designs to provide MaaS, GPUaaS, AIaaS, and other offerings
  • Optimize cluster performance, scalability, energy, and cost efficiency across on-premises, cloud, and hybrid environments
  • Integrate AI and HPC platforms with existing IT systems, data pipelines, and security frameworks
  • Monitor, troubleshoot, and tune infrastructure to ensure high availability, low-latency networking, and workload resiliency
  • Develop and maintain documentation including architecture diagrams, configuration baselines, and operational runbooks
  • Provide technical guidance and support to users, enabling efficient execution of HPC/AI workloads, large-scale models, and simulations.

Skills

Linux Administrator/Architect
CUDA/GPU
HPC cluster design

Education

15 years full time education

Tools

SLURM
PBS Pro
Kubernetes
Docker
Terraform
Ansible
Run:AI

Job description

Project Role : Technology Platform Engineer

Project Role Description : Creates production and non-production cloud environments using the proper software tools such as a platform for a project or product. Deploys the automation pipeline and automates environment creation and configuration.

Must have skills : Linux Architecture

Good to have skills : Machine Learning (ML), Cloud Technology Architecture, Docker Kubernetes Administration, Edge Computing, Microsoft Agentic AI Architect

Minimum 12 year(s) of experience is required

Educational Qualification : 15 years full time education

Key Responsibilities
  • 1) Design and implement HPC and AI infrastructure solutions, aligning system architecture and deployment roadmaps to industry-specific performance and scalability needs
  • 2) Deploy, configure, and manage XPU-based clusters (CPU/GPU/accelerators) using schedulers, VM/K8s orchestration platforms, Slurm, and containerized platforms in scalable designs to provide Metal as a Service (MaaS), GPUaaS, AIaaS, and other offerings
  • 3) Optimize cluster performance, scalability, energy, and cost efficiency across on-premises, cloud, and hybrid environments
  • 4) Integrate AI and HPC platforms with existing IT systems, data pipelines, and security frameworks
  • 5) Monitor, troubleshoot, and tune infrastructure to ensure high availability, low-latency networking, and workload resiliency
  • 6) Develop and maintain documentation including architecture diagrams, configuration baselines, and operational runbooks
  • 7) Provide technical guidance and support to users, enabling efficient execution of HPC/AI workloads, large-scale models, and simulations.
Required Skills and Qualifications
  • 1) Experience in enterprise-wide HPC strategy and architecting next-gen supercomputing environments across the full stack.
  • 2) Primary Skills: Linux Administrator/Architect, Advanced CUDA/GPU on H100/A100, HPC cluster design (SLURM/PBS Pro)
  • 3) Secondary Skills: Parallel programming: MPI, OpenMP, CUDA, SYCL, NVIDIA DGX SuperPOD & BCM, InfiniBand NDR 400Gb/s & RoCEv2 design, MLOps & HPC integration (Kubeflow), Containerization: Singularity, Kubernetes, DevOps: Ansible, Terraform for HPC, UFM management, RunAI, Azure ML integration: distributed training, MLflow, Terraform / Bicep IaC for Azure HPC, FinOps: Reserved Instances, Spot VM strategies, Hybrid cloud HPC: on-prem to Azure/AWS burst
  • 4) Proven hands-on experience designing, deploying, and managing HPC and AI infrastructure across on-premises, cloud, and hybrid environments in 2 or more segments: hyperscaler, neocloud, large Enterprise, Telco/Mobile, supporting key industries such as Financial Services, Life Sciences, Manufacturing, and Retail
  • 5) Deep knowledge of accelerated computing architectures (GPUs, XPUs, DPUs), high-performance fabrics (InfiniBand, Ethernet), SONiC, networking, and modern storage/data platforms (e.g. NVMe-oF, Lustre, GPFS, BeeGFS, VAST, DDN, Weka) to build robust solutions
  • 6) Proficiency with cluster management and orchestration (e.g. Slurm, Run:ai, Kubernetes, Docker), real-time performance monitoring, and observability frameworks
  • 7) Hands-on experience with cloud and virtualization platforms (e.g. AWS, Azure, GCP, VMware, Nutanix) and expertise in automation and optimization using scripting (Python, AI tools) with foundational Infrastructure-as-Code tools such as Terraform and Ansible.
Preferred Skills and Qualifications
  • 1) Experience managing the deployment of 1,000+ GPU clusters for HPC and AI workloads with various infrastructure services enabled
  • 2) Experience with GPU computing libraries and accelerators (e.g., NVIDIA CUDA, Dynamo, AMD ROCm).
  • 3) Experience with AI and HPC Networking (e.g., RoCE, InfiniBand, muti-planar/multi-rail designs, platform buffer architectures)
  • 4) Knowledge of Machine Learning and AI frameworks (e.g., TensorFlow, PyTorch, JAX), Jupyter notebooks / Google Colab environments
  • 5) Familiarity with DevOps practices and tools (e.g., Ansible, Terraform) for infrastructure automation
  • 6) Experience with AgenticAI and associated technologies to leverage and build agents for workflow automation and observability
  • 7) Industry certifications in NVIDIA infrastructure, public cloud providers, Data Science, etc. are a plus
  • 8) Strong problem-solving, troubleshooting, communication, and collaboration skills to deliver reliable, scalable, and high-performance infrastructure solutions in fast-paced, dynamic environments that reward technical talent

15 years full time education

Equal Employment Opportunity Statement

All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by federal, state, or local law.

Please read Accenture’s Recruiting and Hiring Statement for more information on how we process your data during the Recruiting and Hiring process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technology Platform Engineer
Technology Platform Engineer

Accenture PLC • Gurugram District

On-site
INR 4,000,000 - 7,500,000
AI Infrastructure Architect
AI Infrastructure Architect

Accenture PLC • Bengaluru

On-site
INR 3,500,000 - 6,000,000
Cloud Platform Architect
Cloud Platform Architect

Accenture • India

On-site
INR 4,000,000 - 7,000,000
AI Infrastructure Architect
AI Infrastructure Architect

Accenture • Mumbai

On-site
INR 4,000,000 - 8,000,000
Cloud HPC Infrastructure Engineer
Cloud HPC Infrastructure Engineer

Amgen Inc • Hyderabad

On-site
INR 1,800,000 - 2,800,000
Senior AI Compute Engineer
Senior AI Compute Engineer

Neysa • Mumbai

On-site
INR 900,000 - 1,500,000
Specialist Software Engineer
Specialist Software Engineer

Amgen SA • Hyderabad

On-site
INR 3,500,000 - 7,000,000
AI / ML Engineer
AI / ML Engineer

8114 ASOL - Pune 3 SEZ Company • Pune District

On-site
INR 3,000,000 - 6,000,000
Infrastructure Engineer
Infrastructure Engineer

Accenture • Ahmedabad District

On-site
INR 1,200,000 - 2,100,000
AI Infrastructure Architect
AI Infrastructure Architect

Accenture PLC • Gurugram District

On-site
INR 4,000,000 - 6,500,000