AI Operations Platform Consultant

Quantum Technologies. LLC

Jersey City, Northern (NJ, KY)

Hybrid

USD 120,000 - 165,000

Full time

23 hours ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

QUANTUM TECHNOLOGIES LLC in the United States seeks an AI Operations Platform Consultant to deploy, operate, and troubleshoot containerized AI services at scale on Kubernetes (OpenShift). You will configure LLM inference pipelines using TensorRT-LLM and Triton Inference Server, implement robust deployment workflows, and build observability for GPU health and performance.

This role requires enterprise experience with MLOps/LLMOps, incident management, and scalable infrastructure, ensuring

Qualifications

  • Experience deploying and operating containerized AI services at scale.
  • Hands-on with TensorRT-LLM and Triton Inference Server in production.
  • Strong knowledge of MLOps/LLMOps pipelines and model versioning.
  • Experience with monitoring GPU health, latency, and throughput.
  • Security-conscious runtime controls and reliability focus.

Responsibilities

  • Deploy and manage production-grade LLM inference services on Kubernetes (OpenShift).
  • Tune and optimize models with mixed precision, quantization, and batching.
  • Design and implement observability dashboards and alerts for inference systems.
  • Ensure reliability, incident management, and scalable infrastructure.

Skills

GPU-accelerated AI platforms
Kubernetes
LLM inference
MLOps/LLMOps
Observability and telemetry
Containerized microservices

Tools

TensorRT-LLM
Triton Inference Server
OpenShift
Container orchestration

Job description

  • Brings extensive experience operating large-scale GPU-accelerated AI platforms, deploying and managing LLM inference systems on Kubernetes with strong expertise in Triton Inference Server and TensorRT-LLM.
  • They have repeatedly built and optimized production-grade LLM pipelines with GPU-aware scheduling, load balancing, and real-time performance tuning across multi-node clusters. Their background includes designing containerized microservices, implementing robust deployment workflows, and maintaining operational reliability in mission-critical environments.
  • They have led end-to-end LLMOps processes involving model versioning, engine builds, automated rollouts, and secure runtime controls.
  • The candidate has also developed comprehensive observability for inference systems, using telemetry and custom dashboards to track GPU health, latency, throughput, and service availability.
  • Their work consistently incorporates advanced optimization methods such as mixed precision, quantization, sharding, and batching to improve efficiency. Overall, they bring a strong blend of platform engineering, AI infrastructure, and hands-on operational experience running high-performance LLM systems in production
Basic Info:
  • AI Operations Platform Consultant
  • Experience deploying, managing, operating, and troubleshooting containerized services at scale on Kubernetes for mission-critical applications (OpenShift)
  • Experience with deploying, configuring, and tuning LLMs using TensorRT-LLM and Triton Inference server.
  • Managing MLOps/LLMOps pipelines, using TensorRT-LLM and Triton Inference server to deploy inference services in production
  • Setup and operation of AI inference service monitoring for performance and availability.
  • Experience deploying and troubleshooting LLM models on a containerized platform, monitoring, load balancing, etc.
  • Operation and support of MLOps/LLMOps pipelines, using TensorRT-LLM and Triton Inference server to deploy inference services in production
  • Experience deploying and troubleshooting LLM models on a containerized platform, monitoring, load balancing, etc.
  • Experience with standard processes for operation of a mission critical system – incident management, change management, event management, etc.
  • Managing scalable infrastructure for deploying and managing LLMs
  • Deploying models in production environments, including containerization, microservices, and API design
  • Triton Inference Server, including its architecture, configuration, and deployment.
  • Model Optimization techniques using Triton with TRTLLM
  • Model optimization techniques, including pruning, quantization, and knowledge distillation

QUANTUM TECHNOLOGIES LLC is an equal opportunity employer inclusive of female, minority, disability and veterans, (M/F/D/V). Hiring, promotion, transfer, compensation, benefits, discipline, termination and all other employment decisions are made without regard to race, color, religion, sex, sexual orientation, gender identity, age, disability, national origin, citizenship/immigration status, veteran status or any other protected status. QUANTUM TECHNOLOGIES LLC will not make any posting or employment decision that does not comply with applicable laws relating to labor and employment, equal opportunity, employment eligibility requirements or related matters. Nor will QUANTUM TECHNOLOGIES LLC require in a posting or otherwise U.S. citizenship or lawful permanent residency in the U.S. as a condition of employment except as necessary to comply with law, regulation, executive order, or federal, state, or local government contract

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior LLMOps Platform Engineer — GPU AI Infra
Senior LLMOps Platform Engineer — GPU AI Infra

Quantum Technologies. LLC • Jersey City (NJ), Northern (KY)

Hybrid
USD 120,000 - 165,000
AI Platform Engineer
AI Platform Engineer

Bright Vision Technologies • United States

Remote
USD 130,000 - 180,000
ML engineer with Gen AI
ML engineer with Gen AI

Quantum Technologies. LLC • Austin (TX), Northern (KY)

Hybrid
USD 180,000 - 260,000
AI Engineer - Python/Java ML/AI
AI Engineer - Python/Java ML/AI

Quantum Technologies. LLC • Austin (TX), Northern (KY)

Hybrid
USD 120,000 - 160,000
Equal opportunity employer
AI Engineer (Generative AI & LLM)
AI Engineer (Generative AI & LLM)

Quantum Technologies. LLC • Waukesha (WI)

On-site
USD 180,000 - 240,000
AI/ML Tech lead with Gen AI
AI/ML Tech lead with Gen AI

Quantum Technologies. LLC • Rosemead (CA)

Hybrid
USD 180,000 - 240,000
AI Infrastructure Engineer - Inference Platform
AI Infrastructure Engineer - Inference Platform

Hoonify Technologies Inc. • Albuquerque (NM)

On-site
USD 120,000 - 190,000
Sr. NLP Analyst
Sr. NLP Analyst

Quantum Technologies. LLC • Dallas (TX), Northern (KY)

Hybrid
USD 120,000 - 190,000
Senior Developer
Senior Developer

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000
Tech Lead Software Engineer - AI Compute Infrastructure
Tech Lead Software Engineer - AI Compute Infrastructure

ByteDance • Seattle (WA)

On-site
USD 232,560 - 427,500