Get more replies from employers
Send a job-specific resume in minutes.
NTT DATA North America is seeking an On-Premise LLM Inference & GPU Systems Engineer in Charlotte, North Carolina. The role focuses on building and maintaining large-scale on-prem LLM infrastructure using NVIDIA H200 GPU clusters.
Responsibilities include optimizing runtime efficiencies, managing inference serving, and overseeing the complete Hugging Face model lifecycle within the OpenShift AI ecosystem. Ideal candidates should have extensive experience with NVIDIA clusters and modern inference frameworks.
NTT DATA strives to hire exceptional, innovative, and passionate individuals who want to grow with us. If you want to be part of an inclusive, adaptable, and forward-thinking organization, apply now.
We are currently seeking a On-Premise LLM Inference & GPU Systems Engineer to join our team in Charlotte, North Carolina (US-NC), United States (US).
We are seeking an AI Infrastructure Runtime Engineer to build and maintain large-scale on-prem LLM infrastructure. This is an enterprise private GenAI environment running on NVIDIA H200 GPU clusters and an OpenShift AI deployment ecosystem. You will manage production inference internally, including self-hosting open-source LLMs like Llama. We are focused exclusively on inferencing; this role involves no model training infrastructure or fine-tuning pipelines.
NTT DATA is an equal opportunity employer. Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status. For Pay Transparency information, please click here. If you’d like more information on your EEO rights under the law, please click here.