Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Delan Associates, Inc. in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and sustain large-scale on-prem AI infrastructure. The role focuses on production inference with NVIDIA H200 clusters and OpenShift AI, with no training pipelines.
You will optimize runtime, deploy vLLM and TensorRT-LLM, manage KV cache and prefill/decode, and oversee the Hugging Face deployment lifecycle while ensuring efficient GPU utilization and Kubernetes orchestration.
Delan Associates, Inc. in Charlotte, NC is seeking an LLM Inference & GPU Systems Consultant to build and sustain large-scale on-prem AI infrastructure. The role focuses on production inference with NVIDIA H200 clusters and OpenShift AI, with no training pipelines.
You will optimize runtime, deploy vLLM and TensorRT-LLM, manage KV cache and prefill/decode, and oversee the Hugging Face deployment lifecycle while ensuring efficient GPU utilization and Kubernetes orchestration.