Get more replies from employers
Send a job-specific resume in minutes.
Delan Associates, Inc. in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and maintain on-prem LLM infrastructure on NVIDIA H200 clusters with an OpenShift AI deployment.
The role focuses on production inference, not training, requiring onsite presence 3 days/week and hands-on optimization of GPU workloads, vLLM, TensorRT-LLM, and Hugging Face lifecycle. Candidates should have 8+ years in LLM systems or AI infra, experience with OpenShift AI and RunAI, and strong
Job Title: LLM Inference & GPU Systems Consultant
Location: Charlotte, NC (Onsite)
Duration: 6+ Months
Must be onsite at client in Charlotte, NC at least 3 days/week
We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.