Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc

Charlotte (NC)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Delan Associates, Inc. in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and maintain on-prem LLM infrastructure on NVIDIA H200 clusters with an OpenShift AI deployment.

The role focuses on production inference, not training, requiring onsite presence 3 days/week and hands-on optimization of GPU workloads, vLLM, TensorRT-LLM, and Hugging Face lifecycle. Candidates should have 8+ years in LLM systems or AI infra, experience with OpenShift AI and RunAI, and strong

Qualifications

  • 8+ years of experience as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
  • Hands-on experience with NVIDIA H200 clusters and runtime optimization (KV Cache, prefill/decode).
  • Proficiency with OpenShift AI and RunAI for GPU orchestration.
  • Experience with modern inference frameworks: vLLM and TensorRT-LLM.
  • Experience managing the Hugging Face deployment lifecycle.

Responsibilities

  • NVIDIA GPU Runtime Optimization: Drive runtime efficiency and optimize token generation with prefill/decode and KV cache management.
  • Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.
  • Hardware Utilization: Optimize GPU throughput, batching, latency; manage RunAI and Kubernetes GPU orchestration.
  • Model Lifecycle Management: Oversee Hugging Face model onboarding, deployment, retirement.
  • Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

Skills

LLM Systems Engineer
AI Infrastructure Runtime
Runtime optimization
GPU orchestration
OpenShift AI
RunAI
vLLM
TensorRT-LLM
Hugging Face lifecycle

Tools

NVIDIA H200 GPUs
OpenShift AI
RunAI
Kubernetes GPU orchestration
vLLM
TensorRT-LLM

Job description

Delan Associates, Inc. in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and maintain on-prem LLM infrastructure on NVIDIA H200 clusters with an OpenShift AI deployment.

The role focuses on production inference, not training, requiring onsite presence 3 days/week and hands-on optimization of GPU workloads, vLLM, TensorRT-LLM, and Hugging Face lifecycle. Candidates should have 8+ years in LLM systems or AI infra, experience with OpenShift AI and RunAI, and strong

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Prem LLM Inference & GPU Systems Architect
On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior LLM Inference Architect
Senior LLM Inference Architect

NVIDIA • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior DL Systems Engineer, LLM & GPU Performance
Senior DL Systems Engineer, LLM & GPU Performance

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 272,000 - 432,000
Equity
Benefits
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Staff Software Engineer: LLM Inference & GPU Infra (Equity)
Staff Software Engineer: LLM Inference & GPU Infra (Equity)

DevHub • San Francisco (CA)

On-site
USD 180,000 - 240,000
Annual performance bonus
Equity