On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc.

Charlotte (NC)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Compunnel, Inc. in Charlotte, North Carolina is seeking an On-Premise LLM Inference & GPU Systems Engineer. This role involves building, optimizing, and supporting a large-scale enterprise Generative AI infrastructure utilizing NVIDIA H200 GPU clusters and OpenShift AI. Candidates should have significant experience in GPU runtime optimization, Kubernetes orchestration, and managing open-source LLMs.

Key qualifications include 5+ years of experience in relevant roles and hands-on expertise with NVIDIA GPU environments. This position offers a contract opportunity with significant responsibilities across enterprise AI workloads.

Qualifications

  • 5+ years as an LLM Systems Engineer or related role.
  • Hands-on experience with NVIDIA GPU environments.
  • Expertise in deploying and managing inference frameworks.

Responsibilities

  • Design and maintain large-scale LLM inference infrastructure.
  • Optimize performance of token generation pipelines.
  • Manage workload scheduling using Kubernetes.

Skills

NVIDIA GPU environments
Runtime optimization techniques
Token generation pipelines
Kubernetes
OpenShift AI

Tools

vLLM
TensorRT-LLM
RunAI

Job description

Compunnel, Inc. in Charlotte, North Carolina is seeking an On-Premise LLM Inference & GPU Systems Engineer. This role involves building, optimizing, and supporting a large-scale enterprise Generative AI infrastructure utilizing NVIDIA H200 GPU clusters and OpenShift AI. Candidates should have significant experience in GPU runtime optimization, Kubernetes orchestration, and managing open-source LLMs.

Key qualifications include 5+ years of experience in relevant roles and hands-on expertise with NVIDIA GPU environments. This position offers a contract opportunity with significant responsibilities across enterprise AI workloads.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

On-Prem LLM Inference & GPU Systems Architect
On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
Onsite LLM Inference and GPU Systems Engineer
Onsite LLM Inference and GPU Systems Engineer

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 190,000
Onsite LLM Inference Architect for NVIDIA GPU Infra
Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
LLM Inference & GPU Systems Consultant
LLM Inference & GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 190,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)
Senior LLM Inference Runtime Engineer (Kubernetes & GPU)

Hewlett Packard Enterprise • Durham (NC)

On-site
USD 140,000 - 180,000
Health And Wellbeing Benefits
Personal And Professional Development
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
LLM Inference Optimization Engineer - Frontier Performance
LLM Inference Optimization Engineer - Frontier Performance

GMI Cloud, Inc • San Francisco (CA)

On-site
USD 180,000 - 260,000