Get more replies from employers
Send a job-specific resume in minutes.
Compunnel, Inc. in Charlotte, North Carolina is seeking an On-Premise LLM Inference & GPU Systems Engineer. This role involves building, optimizing, and supporting a large-scale enterprise Generative AI infrastructure utilizing NVIDIA H200 GPU clusters and OpenShift AI. Candidates should have significant experience in GPU runtime optimization, Kubernetes orchestration, and managing open-source LLMs.
Key qualifications include 5+ years of experience in relevant roles and hands-on expertise with NVIDIA GPU environments. This position offers a contract opportunity with significant responsibilities across enterprise AI workloads.
North Carolina, Charlotte
06/05/2026
Contract
Active
We are seeking an On-Premise LLM Inference & GPU Systems Engineer to build, optimize, and support a large-scale enterprise Generative AI infrastructure environment. This role is focused exclusively on Large Language Model (LLM) inference operations within a private on-premises ecosystem utilizing NVIDIA H200 GPU clusters and OpenShift AI. The ideal candidate will possess deep expertise in GPU runtime optimization, inference serving platforms, Kubernetes-based orchestration, and production-scale deployment of open-source LLMs. This position will be responsible for maximizing inference performance, operational efficiency, and platform reliability across enterprise AI workloads.