On-Premise LLM Inference & GPU Systems Engineer

NTT DATA North America

Charlotte (NC)

On-site

USD 120,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

NTT DATA North America is seeking an On-Premise LLM Inference & GPU Systems Engineer in Charlotte, North Carolina. The role focuses on building and maintaining large-scale on-prem LLM infrastructure using NVIDIA H200 GPU clusters.

Responsibilities include optimizing runtime efficiencies, managing inference serving, and overseeing the complete Hugging Face model lifecycle within the OpenShift AI ecosystem. Ideal candidates should have extensive experience with NVIDIA clusters and modern inference frameworks.

Qualifications

  • 5+ years expertise as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
  • 5+ years hands-on experience with NVIDIA H200 clusters and runtime optimization techniques.
  • 3+ years experience in OpenShift AI and GPU orchestration tools.

Responsibilities

  • Drive extreme runtime efficiency and optimization for the token generation pipeline.
  • Deploy and manage inference engines including vLLM and TensorRT-LLM.
  • Optimize GPU throughput tuning, batching strategies, and latency optimization.

Skills

NVIDIA H200 clusters
Inference frameworks (vLLM, TensorRT-LLM)
OpenShift AI
GPU orchestration techniques

Tools

Kubernetes
RunAI

Job description

Company Overview

NTT DATA strives to hire exceptional, innovative, and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.

We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States (US).

Job Description

We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.

Key Responsibilities
  • NVIDIA GPU Runtime Optimization: Drive extreme runtime efficiency and optimization for the token generation pipeline. Specifically manage prefill/decode optimization and KV cache management.
  • Inference Serving: Deploy and manage inference engines including vLLM and TensorRT-LLM.
  • Hardware Utilization: Optimize GPU throughput tuning, batching strategies, and latency optimization. Manage workload orchestration using RunAI and Kubernetes GPU orchestration.
  • Model Lifecycle Management: Oversee the complete Hugging Face model lifecycle, including model onboarding, deployment, and retirement.
  • Platform Operations: Operate and maintain the OpenShift AI ecosystem as the primary container platform for GenAI workloads.
Required Qualifications
  • 5+ years expertise as an LLM Systems Engineer or AI Infrastructure Runtime Engineer.
  • 5+ years hands‑on experience with NVIDIA H200 clusters and runtime optimization techniques (KV Cache, prefill/decode).
  • 3+ years experience in OpenShift AI and GPU orchestration tools like RunAI.
  • Strong experience with modern inference frameworks, specifically vLLM and TensorRT-LLM.
  • Proven track record managing the Hugging Face deployment lifecycle.
Equal Employment Opportunity Statement

NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For Pay Transparency information, please click here. If you’d like more information on your EEO rights under the law, please click here.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

On-Premise LLM Inference & GPU Systems Engineer
On-Premise LLM Inference & GPU Systems Engineer

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
LLM Inference GPU Systems Consultant
LLM Inference GPU Systems Consultant

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Prem LLM Inference & GPU Systems Architect
On-Prem LLM Inference & GPU Systems Architect

NTT DATA North America • Charlotte (NC)

On-site
USD 120,000 - 150,000
On-Prem LLM Inference Engineer: GPU & AI Infra
On-Prem LLM Inference Engineer: GPU & AI Infra

Compunnel, Inc. • Charlotte (NC)

On-site
USD 120,000 - 150,000
Onsite LLM Inference Architect for NVIDIA GPU Infra
Onsite LLM Inference Architect for NVIDIA GPU Infra

Delan Associates, Inc • Charlotte (NC)

On-site
USD 150,000 - 210,000
On-Prem LLM Platform Engineer (OpenShift AI / GPU)
On-Prem LLM Platform Engineer (OpenShift AI / GPU)

Infosys Limited • Charlotte (NC)

On-site
USD 80,000 - 120,000
Long-term disability
Health reimbursement accounts
Insurance offerings
+1
Manager, Large Language Model Inference
Manager, Large Language Model Inference

NVIDIA Corporation • Santa Clara (CA)

On-site
USD 184,000 - 287,500
Competitive salary
Equity options
Comprehensive benefits
Senior LLM Inference & Algorithms Engineer Remote, Equity
Senior LLM Inference & Algorithms Engineer Remote, Equity

NVIDIA • United States

On-site
USD 272,000 - 432,000
Equity
Benefits
Solutions Architect, LLM Model Builder
Solutions Architect, LLM Model Builder

NVIDIA Gruppe • Santa Clara (CA)

On-site
USD 152,000 - 241,500
Comprehensive benefits package
Equity options
Senior High-Performance LLM Training Engineer
Senior High-Performance LLM Training Engineer

NVIDIA • Santa Clara (CA)

On-site
USD 184,000 - 356,500
Equity
Benefits