On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc.

Charlotte (NC)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Compunnel, Inc. in Charlotte, North Carolina is seeking an On-Premise LLM Inference & GPU Systems Engineer. This role involves building, optimizing, and supporting a large-scale enterprise Generative AI infrastructure utilizing NVIDIA H200 GPU clusters and OpenShift AI. Candidates should have significant experience in GPU runtime optimization, Kubernetes orchestration, and managing open-source LLMs.

Key qualifications include 5+ years of experience in relevant roles and hands-on expertise with NVIDIA GPU environments. This position offers a contract opportunity with significant responsibilities across enterprise AI workloads.

Qualifications

  • 5+ years as an LLM Systems Engineer or related role.
  • Hands-on experience with NVIDIA GPU environments.
  • Expertise in deploying and managing inference frameworks.

Responsibilities

  • Design and maintain large-scale LLM inference infrastructure.
  • Optimize performance of token generation pipelines.
  • Manage workload scheduling using Kubernetes.

Skills

NVIDIA GPU environments
Runtime optimization techniques
Token generation pipelines
Kubernetes
OpenShift AI

Tools

vLLM
TensorRT-LLM
RunAI

Job description

North Carolina, Charlotte

06/05/2026

Contract

Active

Job Description:
Job Summary

We are seeking an On-Premise LLM Inference & GPU Systems Engineer to build, optimize, and support a large-scale enterprise Generative AI infrastructure environment. This role is focused exclusively on Large Language Model (LLM) inference operations within a private on-premises ecosystem utilizing NVIDIA H200 GPU clusters and OpenShift AI. The ideal candidate will possess deep expertise in GPU runtime optimization, inference serving platforms, Kubernetes-based orchestration, and production-scale deployment of open-source LLMs. This position will be responsible for maximizing inference performance, operational efficiency, and platform reliability across enterprise AI workloads.

Key Responsibilities
  • Design, deploy, and maintain large-scale on-premises LLM inference infrastructure supporting enterprise Generative AI workloads.
  • Optimize runtime performance of token generation pipelines, including prefill/decode optimization and KV cache management.
  • Deploy, configure, and manage inference serving platforms such as vLLM and TensorRT-LLM.
  • Optimize GPU utilization, throughput, batching strategies, latency, and resource efficiency across production inference environments.
  • Manage workload scheduling and orchestration using Kubernetes-based GPU orchestration platforms and RunAI.
  • Oversee the complete lifecycle of open-source language models, including onboarding, deployment, version management, monitoring, and retirement.
  • Manage and support enterprise Hugging Face model deployment workflows and operational processes.
  • Operate, maintain, and optimize the OpenShift AI ecosystem supporting Generative AI applications and services.
  • Monitor platform performance, identify bottlenecks, and implement optimization strategies to improve inference efficiency and scalability.
  • Collaborate with AI, platform engineering, infrastructure, and operations teams to ensure reliable service delivery.
  • Implement operational best practices related to platform availability, monitoring, security, and governance.
  • Develop automation, deployment processes, and operational procedures to support platform scalability and maintainability.
  • Troubleshoot and resolve infrastructure, inference, performance, and deployment issues across the AI ecosystem.
  • Create and maintain technical documentation, operational runbooks, and platform standards.
Required Qualifications
  • 5+ years of experience as an LLM Systems Engineer, AI Infrastructure Engineer, AI Platform Engineer, or related role.
  • 5+ years of hands‑on experience supporting NVIDIA GPU environments and runtime optimization techniques.
  • Experience optimizing token generation pipelines, including KV cache management and prefill/decode optimization strategies.
  • Strong experience deploying and managing inference frameworks such as vLLM and TensorRT-LLM.
  • 3+ years of experience with OpenShift AI and containerized AI platform operations.
  • 3+ years of experience with GPU orchestration technologies, including RunAI and Kubernetes-based environments.
  • Experience deploying, managing, and supporting open-source Large Language Models in production environments.
  • Proven experience managing the Hugging Face model lifecycle, including onboarding, deployment, version management, and retirement.
  • Strong understanding of AI inference architectures, GPU resource management, workload optimization, and performance tuning.
  • Experience working with containerization technologies, Kubernetes, and cloud-native application platforms.
  • Strong troubleshooting, performance analysis, and problem‑solving skills.
  • Excellent communication and collaboration skills with the ability to work across infrastructure, platform, and AI engineering teams.
Preferred Qualifications
  • Experience supporting enterprise-scale Generative AI platforms and private AI infrastructure environments.
  • Experience optimizing large-scale LLM inference workloads in highly regulated or secure environments.
  • Knowledge of AI observability, monitoring, logging, and performance analytics tools.
  • Experience implementing infrastructure automation and operational tooling for AI platforms.
  • Familiarity with enterprise governance, security, and compliance practices for AI workloads.
  • Experience supporting multi-cluster Kubernetes or OpenShift environments.
  • Knowledge of emerging trends and best practices in LLM inference optimization and AI platform engineering.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Prem LLM Inference & GPU Systems Architect
On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
Onsite LLM Inference Architect for NVIDIA GPU Infra
Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
AI Infrastructure Engineer
AI Infrastructure Engineer

The Josef Group • Chantilly (VA)

On-site
USD 200,000 - 250,000
Generative AI Engineer
Generative AI Engineer

XPath Solutions • Charlotte (NC)

On-site
USD 140,000 - 190,000
On-Prem LLM Platform Engineer (OpenShift AI / GPU)
On-Prem LLM Platform Engineer (OpenShift AI / GPU)

Infosys Limited • Charlotte (NC)

On-site
USD 80,000 - 120,000
Long-term disability
Health reimbursement accounts
Insurance offerings
+1
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Motion Recruitment • Raleigh (NC)

On-site
USD 180,000 - 240,000