AI/ML Engineer for HPC & Distributed Systems (Semi-Remote)

Hewlett Packard Enterprise

Spring (TX)

On-site

USD 121,000 - 277,000

Full time

13 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Hewlett Packard Enterprise is seeking an experienced HPC/AI Engineer to advance high-performance computing and AI workloads. You will install, configure, and optimize complex IT infrastructure, develop automation scripts, and study performance of large language models on HPE GPU servers.

The role requires 5+ years in ML/AI and HPC, experience with containers, Slurm, and AI benchmarks. Strong Python/C/C++ skills and a self-starter attitude are essential for a semi-remote team.

Qualifications

  • Master's degree or PhD in CS, Engineering, IT, or Systems.
  • 5+ years of experience.
  • Experience with NCCL, HPL, and AI benchmarks.
  • Experience with containers and distributed DL/transformers.
  • Strong Python, C, C++ programming.
  • CI/CD and scripting experience desired.
  • Self-starter able to work in a semi-remote setting.

Responsibilities

  • Installs and configures complex IT infrastructure components (servers, storage, network).
  • Develop software scripts and configurations for automating deployment.
  • Study and improve the performance of Large Language Models on HPE GPU servers.
  • Analyze server workloads and DL/ML code on HPC platforms via InfiniBand.
  • Write white papers and guidance on AI workload and model selection.
  • Capture and review performance data, logs, and traces.
  • Develop scripts to analyze AI workload performance data.
  • Communicate technical work clearly to non-technical colleagues.
  • Collaborate with software/hardware partners to optimize systems.
  • Document and report issues found during testing and evaluation.
  • Provide status updates to management in a timely manner.
  • Mentor less-experienced staff members.
  • Run AI and HPC benchmarks.

Skills

ML/AI
HPC
Python
C/C++
NCCL
HPL
Containers
Slurm
NTFS/Lustre
CI/CD
Transformers
Self-motivation

Education

Masters/PhD in CS/Engineering

Tools

Docker
Slurm
NCCL
HPL
NTFS
Lustre

Job description

Hewlett Packard Enterprise is seeking an experienced HPC/AI Engineer to advance high-performance computing and AI workloads. You will install, configure, and optimize complex IT infrastructure, develop automation scripts, and study performance of large language models on HPE GPU servers.

The role requires 5+ years in ML/AI and HPC, experience with containers, Slurm, and AI benchmarks. Strong Python/C/C++ skills and a self-starter attitude are essential for a semi-remote team.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Hybrid AI & HPC ML Engineer
Hybrid AI & HPC ML Engineer

Hewlett Packard Enterprise Company • Springs (NY)

Hybrid
USD 121,000 - 277,000
Health & Wellbeing
Professional development
Unconditional inclusion
Hybrid AI & HPC Engineer, Finance
Hybrid AI & HPC Engineer, Finance

Hewlett Packard Enterprise Company in • Spring (TX)

Hybrid
USD 121,000 - 277,000
Hybrid work model
Hybrid AI/ML Engineer — Scale AI & HPC Systems
Hybrid AI/ML Engineer — Scale AI & HPC Systems

Hewlett Packard Enterprise Development LP • Fort Collins (CO)

Hybrid
USD 144,000 - 273,000
Health & wellbeing
Career development
Inclusive culture
AI and Machine Learning Engineer
AI and Machine Learning Engineer

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 120,000 - 160,000
Hybrid AI & ML Engineer — Optimize LLMs & Systems
Hybrid AI & ML Engineer — Optimize LLMs & Systems

Hewlett Packard Enterprise Development LP • Spring (TX)

Hybrid
USD 120,000 - 160,000
AI and Machine Learning Engineer
AI and Machine Learning Engineer

Hewlett Packard Enterprise Company • Springs (NY)

Hybrid
USD 121,000 - 277,000
Health & Wellbeing
Professional development
Unconditional inclusion
AI and Machine Learning Engineer
AI and Machine Learning Engineer

Hewlett Packard Enterprise Company in • Spring (TX)

Hybrid
USD 121,000 - 277,000
Hybrid work model
Onsite HPC & AI Performance Engineer
Onsite HPC & AI Performance Engineer

Hewlett Packard Enterprise • Spring (TX)

On-site
USD 105,000 - 243,000
Health & Wellbeing benefits
Personal & Professional Development programs
Inclusive workplace culture
Hybrid AI HPC Infrastructure Engineer (GPU/ML)
Hybrid AI HPC Infrastructure Engineer (GPU/ML)

Analysis Group, Inc. • Boston (MA)

On-site
USD 150,000 - 170,000
Discretionary annual bonus
Benefits package
HPC AI Systems Architect (On-Prem GPU Cluster)
HPC AI Systems Architect (On-Prem GPU Cluster)

MRE Consulting • Houston (TX)

On-site
USD 95,000 - 140,000