Technology Platform Engineer

Accenture PLC

Gurugram District

On-site

INR 4,000,000 - 7,500,000

Full time

2 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Accenture is hiring a Technology Platform Engineer to design and deploy HPC and AI infrastructure, including XPU-based clusters, using Slurm, Kubernetes, and containerized platforms. You will optimize performance across on-prem, cloud, and hybrid environments and integrate AI/HPC workflows into existing IT systems.

You will mentor users, develop runbooks, and contribute to architecture blueprints for scalable, secure, and cost-efficient solutions that support large-scale models and simulations.

Qualifications

  • Experience in enterprise-wide HPC strategy and architecting next-gen supercomputing environments across the full stack.
  • Primary skills: Linux Administrator/Architect, Advanced CUDA/GPU on H100/A100, HPC cluster design (SLURM/PBS Pro).
  • Secondary skills: Parallel programming – MPI, OpenMP, CUDA, SYCL; InfiniBand/NDR 400Gb/s, Kubeflow integration.
  • Hands-on design, deployment, and management of HPC/AI infra across on-prem, cloud, and hybrid environments in multiple segments.
  • Deep knowledge of GPUs/XPUs, high-performance fabrics, SONiC, networking, and storage platforms (NVMe-oF, Lustre, GPFS).
  • Proficiency with cluster orchestration (Slurm, Run:ai, Kubernetes, Docker) and real-time performance monitoring.
  • Experience with cloud/virtualization platforms (AWS, Azure, GCP, VMware) and IaC tools like Terraform/Ansible.

Responsibilities

  • Design and implement HPC and AI infrastructure solutions aligned to performance and scalability needs.
  • Deploy, configure, and manage XPU-based clusters using schedulers and containerized platforms; provide MaaS/AIaaS.
  • Optimize cluster performance, energy, and cost across on-prem, cloud, and hybrid environments.
  • Integrate AI and HPC platforms with IT systems, data pipelines, and security frameworks.
  • Monitor, troubleshoot, and tune infrastructure for high availability and low-latency workloads.
  • Develop and maintain architecture diagrams, baselines, and operational runbooks.
  • Provide technical guidance and support for HPC/AI workloads, large-scale models, and simulations.

Skills

Linux Architecture
Machine Learning (ML)
Cloud Technology Architecture
Docker Kubernetes Administration
Edge Computing
Microsoft Agentic AI Architect

Education

15 years full time education

Tools

Slurm
PBS Pro
Kubeflow
Singularity
Kubernetes
Docker
RunAI
Terraform
Ansible
Azure ML

Job description

Project Role:

Technology Platform Engineer

Project Role Description:

Creates production and non-production cloud environments using the proper software tools such as a platform for a project or product. Deploys the automation pipeline and automates environment creation and configuration.

Must have skills:

Linux Architecture

Good to have skills:

Machine Learning (ML), Cloud Technology Architecture, Docker Kubernetes Administration, Edge Computing, Microsoft Agentic AI Architect

Minimum 5 year(s) of experience is required

Educational Qualification:

15 years full time education

Key Responsibilities:
  • 1) Design and implement HPC and AI infrastructure solutions, aligning system architecture and deployment roadmaps to industry-specific performance and scalability needs
  • 2) Deploy, configure, and manage XPU-based clusters (CPU/GPU/accelerators) using schedulers, VM/K8s orchestration platforms, Slurm, and containerized platforms in scalable designs to provide Metal as a Service (MaaS), GPUaaS, AIaaS, and other offerings
  • 3) Optimize cluster performance, scalability, energy, and cost efficiency across on-premises, cloud, and hybrid environments
  • 4) Integrate AI and HPC platforms with existing IT systems, data pipelines, and security frameworks
  • 5) Monitor, troubleshoot, and tune infrastructure to ensure high availability, low-latency networking, and workload resiliency
  • 6) Develop and maintain documentation including architecture diagrams, configuration baselines, and operational runbooks
  • 7) Provide Provide technical guidance and support to users, enabling efficient execution of HPC/AI workloads, large-scale models, and simulations
Required Skills and Qualifications:
  • 1) Experience in enterprise-wide HPC strategy and architecting next-gen supercomputing environments across the full stack.
  • 2) Primary Skills: Linux Administrator/Architect, Advanced CUDA/GPU on H100/A100, HPC cluster design (SLURM/PBS Pro)
  • 3) Secondary Skills: Parallel programming: MPI, OpenMP, CUDA, SYCL, NVIDIA DGX SuperPOD & BCM, InfiniBand NDR 400Gb/s & RoCEv2 design, MLOps & HPC integration (Kubeflow), Containerization: Singularity, Kubernetes, DevOps: Ansible, Terraform for HPC, UFM management, RunAI, Azure ML integration: distributed training, MLflow, Terraform / Bicep IaC for Azure HPC, FinOps: Reserved Instances, Spot VM strategies, Hybrid cloud HPC: on-prem to Azure/AWS burst
  • 4) Proven hands-on experience designing, deploying, and managing HPC and AI infrastructure across on-premises, cloud, and hybrid environments in 2 or more segments: hyperscaler, neocloud, large Enterprise, Telco/Mobile, supporting key industries such as Financial Services, Life Sciences, Manufacturing, and Retail
  • 5) Deep knowledge of accelerated computing architectures (GPUs, XPUs, DPUs), high-performance fabrics (InfiniBand, Ethernet), SONiC, networking, and modern storage/data platforms (e.g. NVMe-oF, Lustre, GPFS, BeeGFS, VAST, DDN, Weka) to build robust solutions
  • 6) Proficiency with cluster management and orchestration (e.g. Slurm, Run:ai, Kubernetes, Docker), real-time performance monitoring, and observability frameworks
  • 7) Hands-on experience with cloud and virtualization platforms (e.g. AWS, Azure, GCP, VMware, Nutanix) and expertise in automation and optimization using scripting (Python, AI tools) with foundational Infrastructure-as-Code tools such as Terraform and Ansible.
Preferred Skills and Qualifications:
  • 1) Experience managing the deployment of 1,000+ GPU clusters for HPC and AI workloads with various infrastructure services enabled
  • 2) Experience with GPU computing libraries and accelerators (e.g., NVIDIA CUDA, Dynamo, AMD ROCm).
  • 3) Experience with AI and HPC Networking (e.g., RoCE, InfiniBand, muti-planar/multi-rail designs, platform buffer architectures)
  • 4) Knowledge of Machine Learning and AI frameworks (e.g., TensorFlow, PyTorch, JAX), Jupyter notebooks / Google Colab environments
  • 5) Familiarity with DevOps practices and tools (e.g., Ansible, Terraform) for infrastructure automation
  • 6) Experience with AgenticAI and associated technologies to leverage and build agents for workflow automation and observability
  • 7) Industry certifications in NVIDIA infrastructure, public cloud providers, Data Science, etc. are a plus
  • 8) Strong problem-solving, troubleshooting, communication, and collaboration skills to deliver reliable, scalable, and high-performance infrastructure solutions in fast-paced, dynamic environments that reward technical talent

15 years full time education

Important Notice

We have been alerted to the existence of fraudulent messages asking job seekers to set up payment to cover various costs associated with establishing employment at Accenture. No one is ever required to pay for employment at Accenture. If you are contacted by someone asking for payment, please do not respond, and contact us at india.fc.check@accenture.com immediately.

Equal Employment Opportunity Statement

All employment decisions shall be made without regard to age, race, creed, color, religion, sex, national origin, ancestry, disability status, veteran status, sexual orientation, gender identity or expression, genetic information, marital status, citizenship status or any other basis as protected by federal, state, or local law.

Please read Accenture's Recruiting and Hiring Statement for more information on how we process your data during the Recruiting and Hiring process.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technology Platform Engineer
Technology Platform Engineer

Accenture PLC • Bengaluru

On-site
INR 3,500,000 - 5,500,000
Operations Engineer
Operations Engineer

Accenture PLC • Gurugram District

On-site
INR 4,000,000 - 6,000,000
DevOps Architect
DevOps Architect

Accenture • India

On-site
INR 2,800,000 - 4,800,000
Data Platform Architect
Data Platform Architect

Accenture • Gurugram District

On-site
INR 4,000,000 - 7,000,000
AI/ML Computational Science Specialist
AI/ML Computational Science Specialist

Accenture • Mumbai

On-site
INR 2,500,000 - 4,500,000
AppModernization_AWS_Solution_Architect
AppModernization_AWS_Solution_Architect

Accenture PLC • New Town

On-site
INR 2,400,000 - 3,600,000
S&C GN - TS&T –Enterprise AI Value Strategy - Consultant
S&C GN - TS&T –Enterprise AI Value Strategy - Consultant

Accenture PLC • Gurugram District

On-site
INR 2,400,000 - 4,200,000
Technology Platform Engineer
Technology Platform Engineer

Accenture PLC • Hyderabad

On-site
INR 3,000,000 - 6,000,000
S&C Global Network - AI - Hi Tech - Data Science Analyst 2
S&C Global Network - AI - Hi Tech - Data Science Analyst 2

Accenture PLC • Bengaluru

Hybrid
INR 1,400,000 - 2,100,000
AppModernization_Agentic_AI_Solution_Architect
AppModernization_Agentic_AI_Solution_Architect

Accenture PLC • New Town

On-site
INR 2,500,000 - 4,000,000