On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America

Charlotte (NC)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

NTT DATA North America is seeking an On-Premise LLM Inference & GPU Systems Engineer in Charlotte, North Carolina. The role focuses on building and maintaining large-scale on-prem LLM infrastructure using NVIDIA H200 GPU clusters.

Responsibilities include optimizing runtime efficiencies, managing inference serving, and overseeing the complete Hugging Face model lifecycle within the OpenShift AI ecosystem. Ideal candidates should have extensive experience with NVIDIA clusters and modern inference frameworks.

Qualifications

  • 5+ years expertise as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
  • 5+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques.
  • 3+ years experience in OpenShift AI and GPU orchestration tools.

Responsibilities

  • Drive extreme runtime efficiency and optimization for the token generation pipeline.
  • Deploy and manage inference engines including vLLM and TensorRT-LLM.
  • Optimize GPU throughput tuning, batching strategies, and latency optimization.

Skills

NVIDIA H200 clusters
Inference frameworks (vLLM, TensorRT-LLM)
OpenShift AI
GPU orchestration techniques

Tools

Kubernetes
RunAI

Job description

NTT DATA North America is seeking an On-Premise LLM Inference & GPU Systems Engineer in Charlotte, North Carolina. The role focuses on building and maintaining large-scale on-prem LLM infrastructure using NVIDIA H200 GPU clusters.

Responsibilities include optimizing runtime efficiencies, managing inference serving, and overseeing the complete Hugging Face model lifecycle within the OpenShift AI ecosystem. Ideal candidates should have extensive experience with NVIDIA clusters and modern inference frameworks.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Onsite LLM Inference Architect for NVIDIA GPU Infra
Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
Onsite LLM Inference and GPU Systems Engineer
Onsite LLM Inference and GPU Systems Engineer

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 190,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
LLM Inference & GPU Systems Consultant
LLM Inference & GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 190,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 140,000 - 180,000
Health And Wellbeing Benefits
Personal And Professional Development
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits