AI & HPC Infrastructure Engineer

Prodapt ASIC services (Formerly Innovative Logic)

San Jose (CA)

On-site

USD 150,000 - 210,000

Full time

7 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Prodapt is seeking an AI & HPC Infrastructure Engineer to design, build, and operate GPU-enabled compute platforms across on-prem and cloud environments. You will automate deployments, monitor performance, and contribute to scalable, cost-effective AI/ML infrastructure for engineering and R&D teams.

You will work at the intersection of AI infrastructure, cloud platforms, and operations, enabling hybrid compute strategies and ensuring high availability for AI-driven development workflows.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or related technical discipline.
  • 3+ years of hands-on infrastructure engineering, platform engineering, cloud operations, or HPC experience.
  • Strong Linux-based infrastructure administration and cloud platform experience (AWS/Azure/GCP).

Responsibilities

  • Build, configure, and operate GPU and HPC clusters across compute, storage, and networking.
  • Deploy, monitor, and optimize AI/ML workloads and HPC workloads for reliability and cost-efficiency.
  • Implement IaC and automation for provisioning and operations.
  • Manage multi-cloud and on-prem environments (AWS, Azure, GCP).
  • Collaborate with engineering teams to enable AI-driven development workflows.
  • Ensure high availability, performance tuning, and capacity planning.
  • Develop monitoring, logging, and alerting solutions for platform reliability.
  • Support AI services deployment, including LLM APIs and AI/agent platforms.

Skills

Linux admin
GPU/HPC infra
Kubernetes/Slurm
Automation/IaC
Public cloud ops

Education

Bachelor's degree in CS/Engineering

Tools

Kubernetes
Slurm
Terraform
Linux

Job description

Prodapt is the largest specialized player in the Connectedness industry. As an AI-first strategic technology partner, Prodapt provides consulting, business reengineering, and managed services for the largest telecom and tech enterprises building networks and digital experiences of tomorrow. A ServiceNow-invested company, Prodapt has been recognized by Gartner as a Large, Telecom-Native, Regional IT Service Provider. Prodapt’s ASIC Services is a leading provider of SoC/ASIC RTL Design, UVM based verification, Emulation, FPGA based validation, DFT, RTL2GDSII, Physical Design using ICC2 and Innovus, Mask Layout, Firmware, Silicon Bringup, and Analog mask layout. Our embedded services include device drivers, RTOS porting, and board bring-up. A “Great Place To Work® Certified™” company, Prodapt employs over 5,000 technology and domain experts in 30+ countries. Prodapt is part of the 130-year-old business conglomerate The Jhaver Group, which employs over 32,000 people across 80+ locations globally.

We are looking for an AI & HPC Infrastructure Engineer to design, build, and operate the compute platforms that power our AI and High-Performance Computing (HPC) workloads.

In this role, you will deploy, automate, and manage GPU-enabled infrastructure across on-premises and cloud environments, enabling scalable, reliable, and cost-effective compute resources for engineering and R&D teams. You will work at the intersection of AI infrastructure, cloud platforms, automation, and operations to support next-generation AI and engineering applications.

Key Responsibilities
  • Build, configure, and operate GPU and HPC clusters across compute, storage, and networking environments.
  • Support capacity planning, performance tuning, and resource optimization for AI training, inference, and compute-intensive workloads.
  • Monitor infrastructure health and ensure high availability and performance.
  • Deploy and manage compute environments across on-premises and public cloud platforms (AWS, Azure, or GCP).
  • Contribute to infrastructure modernization, scalability, and resiliency initiatives.
  • Support cloud adoption and hybrid computing strategies.
  • Implement Infrastructure-as-Code (IaC) and automation frameworks for provisioning and operations.
  • Develop monitoring, logging, and alerting solutions to improve platform reliability.
  • Drive continuous improvements in operational efficiency and resource utilization.
  • Deploy, integrate, and support AI services, including LLM APIs, coding assistants, and AI/agent platforms.
  • Collaborate with engineering teams to enable AI-driven development workflows.
  • Support AI/ML infrastructure requirements and best practices.
  • Troubleshoot and resolve infrastructure, networking, and platform issues.
  • Create and maintain technical documentation, standards, and operational runbooks.
  • Partner with engineering, IT, and platform teams to deliver secure and scalable solutions.
Required Qualifications
  • Bachelor's degree in Computer Science, Engineering, or a related technical discipline.
  • 3+ years of hands-on experience in Infrastructure Engineering, Platform Engineering, Cloud Operations, or HPC environments.
  • Strong experience with Linux-based infrastructure administration.
  • Experience working with public cloud platforms such as AWS, Azure, or GCP.
  • Hands-on experience with GPU/HPC environments and workload orchestration platforms such as Kubernetes or Slurm.
  • Experience with automation, infrastructure provisioning, monitoring, and performance optimization.
  • Solid understanding of compute, storage, networking, virtualization, and container technologies.
Preferred Qualifications
  • Experience supporting AI/ML workloads and GPU-based infrastructure.
  • Knowledge of Kubernetes, Docker, Infrastructure-as-Code tools, and observability platforms.
  • Familiarity with AI platforms, LLM deployment, and modern engineering productivity tools.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI & HPC Infra Engineer: GPU Compute & Cloud Automation
AI & HPC Infra Engineer: GPU Compute & Cloud Automation

Prodapt ASIC services (Formerly Innovative Logic) • San Jose (CA)

On-site
USD 150,000 - 210,000
Senior Solutions Engineer, AI Infrastructure
Senior Solutions Engineer, AI Infrastructure

VAST Data • New York (NY)

On-site
USD 150,000 - 200,000
AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saige Partners • San Jose (CA)

On-site
USD 140,000 - 190,000
HPC AI Systems Administrator
HPC AI Systems Administrator

MRE Consulting • Houston (TX)

On-site
USD 120,000 - 180,000
Competitive salary
Comprehensive benefits
Professional development support
Cluster Design
Cluster Design

Blue Signal Search • San Francisco (CA)

On-site
USD 150,000 - 230,000
Head of AI Data Center Infrastructure Platforms and Software
Head of AI Data Center Infrastructure Platforms and Software

Summit Group Solutions, LLC • United States

On-site
USD 150,000 - 350,000
AI Systems Administrator
AI Systems Administrator

Murray Resources - Best Staffing Agency • Houston (TX)

On-site
USD 100,000 - 140,000
401K
AI/HPC Systems Engineer
AI/HPC Systems Engineer

Saigepartners • San Jose (CA)

Hybrid
USD 120,000 - 180,000
Member of Technical Staff - GPU Infrastructure
Member of Technical Staff - GPU Infrastructure

Prime Intellect • United States

On-site
USD 120,000 - 150,000
AI Kernel / Cluster Engineer
AI Kernel / Cluster Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 150,000 - 210,000