On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America

Charlotte (NC)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA North America is seeking an On-Premise LLM Inference & GPU Systems Engineer in Charlotte, North Carolina. The role focuses on building and maintaining large-scale on-prem LLM infrastructure using NVIDIA H200 GPU clusters.

Responsibilities include optimizing runtime efficiencies, managing inference serving, and overseeing the complete Hugging Face model lifecycle within the OpenShift AI ecosystem. Ideal candidates should have extensive experience with NVIDIA clusters and modern inference frameworks.

Qualifications

  • 5+ years expertise as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
  • 5+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques.
  • 3+ years experience in OpenShift AI and GPU orchestration tools.

Responsibilities

  • Drive extreme runtime efficiency and optimization for the token generation pipeline.
  • Deploy and manage inference engines including vLLM and TensorRT-LLM.
  • Optimize GPU throughput tuning, batching strategies, and latency optimization.

Skills

NVIDIA H200 clusters
Inference frameworks (vLLM, TensorRT-LLM)
OpenShift AI
GPU orchestration techniques

Tools

Kubernetes
RunAI

Job description

NTT DATA North America is seeking an On-Premise LLM Inference & GPU Systems Engineer in Charlotte, North Carolina. The role focuses on building and maintaining large-scale on-prem LLM infrastructure using NVIDIA H200 GPU clusters.

Responsibilities include optimizing runtime efficiencies, managing inference serving, and overseeing the complete Hugging Face model lifecycle within the OpenShift AI ecosystem. Ideal candidates should have extensive experience with NVIDIA clusters and modern inference frameworks.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Onsite LLM Inference Architect for NVIDIA GPU Infra
Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Prem LLM Platform Engineer (OpenShift AI & GPUs)
On-Prem LLM Platform Engineer (OpenShift AI & GPUs)

Infosys • Charlotte (NC)

On-site
USD 80,000 - 120,000
Medical/Dental/Vision Insurance
401(k) plan with contributions
Paid holidays plus Paid Time Off
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior LLM Inference Architect
Senior LLM Inference Architect

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
On-Prem LLM Platform Engineer — OpenShift AI & GPU
On-Prem LLM Platform Engineer — OpenShift AI & GPU

Infosys Limited • Charlotte (NC)

On-site
USD 80,000 - 120,000
Long-term disability
Health reimbursement accounts
Insurance offerings
+1
GenAI Platform: LLM Inference Engineer (Cloud)
GenAI Platform: LLM Inference Engineer (Cloud)

Infosys • Charlotte (NC)

On-site
USD 90,000 - 120,000
Medical/Dental/Vision/Life Insurance
401(k) plan
Paid Time Off