On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc.

Charlotte (NC)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Compunnel, Inc. in Charlotte, North Carolina is seeking an On-Premise LLM Inference & GPU Systems Engineer. This role involves building, optimizing, and supporting a large-scale enterprise Generative AI infrastructure utilizing NVIDIA H200 GPU clusters and OpenShift AI. Candidates should have significant experience in GPU runtime optimization, Kubernetes orchestration, and managing open-source LLMs.

Key qualifications include 5+ years of experience in relevant roles and hands-on expertise with NVIDIA GPU environments. This position offers a contract opportunity with significant responsibilities across enterprise AI workloads.

Qualifications

  • 5+ years as an LLM Systems Engineer or related role.
  • Hands-on experience with NVIDIA GPU environments.
  • Expertise in deploying and managing inference frameworks.

Responsibilities

  • Design and maintain large-scale LLM inference infrastructure.
  • Optimize performance of token generation pipelines.
  • Manage workload scheduling using Kubernetes.

Skills

NVIDIA GPU environments
Runtime optimization techniques
Token generation pipelines
Kubernetes
OpenShift AI

Tools

vLLM
TensorRT-LLM
RunAI

Job description

Compunnel, Inc. in Charlotte, North Carolina is seeking an On-Premise LLM Inference & GPU Systems Engineer. This role involves building, optimizing, and supporting a large-scale enterprise Generative AI infrastructure utilizing NVIDIA H200 GPU clusters and OpenShift AI. Candidates should have significant experience in GPU runtime optimization, Kubernetes orchestration, and managing open-source LLMs.

Key qualifications include 5+ years of experience in relevant roles and hands-on expertise with NVIDIA GPU environments. This position offers a contract opportunity with significant responsibilities across enterprise AI workloads.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Prem LLM Inference & GPU Systems Architect
On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
Onsite LLM Inference Architect for NVIDIA GPU Infra
Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
Senior AI Inference Systems Engineer (GPU & HPC)
Senior AI Inference Systems Engineer (GPU & HPC)

NVIDIA • Santa Clara (CA)

Hybrid
USD 184,000 - 357,000
Equity
Benefits
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)
Senior DL Inference Engineer — GPU-Optimized LLMs (Remote)

NVIDIA Corporation • Northern (KY)

Hybrid
USD 152,000 - 288,000
ML Engineer—LLM Inference & GPU Optimization (Equity)
ML Engineer—LLM Inference & GPU Optimization (Equity)

IC Resources • San Francisco (CA)

On-site
USD 200,000 - 290,000
401(k)
Unlimited PTO
Modern engineering workspace
+1
On-Prem LLM Platform Engineer (OpenShift AI & GPUs)
On-Prem LLM Platform Engineer (OpenShift AI & GPUs)

Infosys • Charlotte (NC)

On-site
USD 80,000 - 120,000
Medical/Dental/Vision Insurance
401(k) plan with contributions
Paid holidays plus Paid Time Off
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes
Senior AI Inference Engineer: GPU Kernels & LLM Runtimes

NVIDIA • Redmond (WA)

On-site
USD 184,000 - 288,000
Equity
Benefits