Onsite LLM Inference and GPU Systems Engineer

Delan Associates, Inc

Charlotte (NC)

On-site

USD 150,000 - 190,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Delan Associates, Inc. in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and sustain large-scale on-prem AI infrastructure. The role focuses on production inference with NVIDIA H200 clusters and OpenShift AI, with no training pipelines.

You will optimize runtime, deploy vLLM and TensorRT-LLM, manage KV cache and prefill/decode, and oversee the Hugging Face deployment lifecycle while ensuring efficient GPU utilization and Kubernetes orchestration.

Qualifications

  • 8+ years of hands-on experience in LLM infrastructure and AI runtime engineering.
  • 8+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).
  • Proficiency in OpenShift AI and GPU orchestration tools like RunAI.
  • Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.
  • Proven track record managing the Hugging Face deployment lifecycle.

Responsibilities

  • NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.
  • Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.
  • Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.
  • Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.
  • Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.

Education

8+ years LLM Systems Engineer / AI Infrastructure Runtime experience

Tools

NVIDIA H200 clusters
OpenShift AI
RunAI
vLLM
TensorRT-LLM
Hugging Face deployment

Job description

Delan Associates, Inc. in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and sustain large-scale on-prem AI infrastructure. The role focuses on production inference with NVIDIA H200 clusters and OpenShift AI, with no training pipelines.

You will optimize runtime, deploy vLLM and TensorRT-LLM, manage KV cache and prefill/decode, and oversee the Hugging Face deployment lifecycle while ensuring efficient GPU utilization and Kubernetes orchestration.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Onsite LLM Inference Architect for NVIDIA GPU Infra
Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Prem LLM Inference & GPU Systems Architect
On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
LLM Inference & GPU Systems Consultant
LLM Inference & GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 190,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 140,000 - 180,000
Health And Wellbeing Benefits
Personal And Professional Development
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
Senior DL Inference Engineer - GPU/LLM Performance & Equity
Senior DL Inference Engineer - GPU/LLM Performance & Equity

NVIDIA • Washington

On-site
USD 184,000 - 288,000
Equity
Benefits