AI Operations Platform Engineer - LLMs on Kubernetes

MACHINE LEARNING TECHNOLOGIES LLC

Jersey City (NJ)

On-site

USD 70,000 - 100,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

MACHINE LEARNING TECHNOLOGIES LLC is seeking an AI Operations Platform Consultant experienced in deploying and managing large-scale GPU-accelerated AI platforms. Responsibilities include leading LLMOps processes and optimizing LLM pipelines for performance across Kubernetes environments.

The ideal candidate must possess strong expertise in Triton Inference Server and TensorRT-LLM, and will be responsible for building production-grade LLM inference systems. This position offers competitive hourly rates and the opportunity for extension beyond 24 months.

Qualifications

  • Extensive experience in deploying and managing LLM inference systems.
  • Strong expertise in Triton Inference Server and TensorRT-LLM.
  • Experience in optimizing production-grade LLM pipelines.

Responsibilities

  • Lead end-to-end LLMOps processes with model versioning and automated rollouts.
  • Manage AI inference service monitoring for performance.
  • Optimize LLM models using techniques like quantization and pruning.

Skills

Large-scale GPU-accelerated AI platforms
LLM inference systems on Kubernetes
Triton Inference Server
TensorRT-LLM
MLOps/LLMOps pipelines
Containerized services

Job description

  • Home
  • About
  • Adaptive AI
  • Industries
  • Consulting
  • Jobs-Portal
  • Contact Us
  • Home
  • About
  • Adaptive AI
  • Industries
  • Consulting
  • Jobs-Portal
  • Contact Us

Job Description

Job Overview

  • Job ID:

J53022

  • Job Title:

AI Operations Platform Consultant

  • Location:

Jersey City, NJ

  • Duration:

24 Months + Extension

  • Hourly Rate:

Depending on Experience (DOE)

  • Work Authorization:

US Citizen, Green Card, OPT-EAD, CPT, H-1B,

H4-EAD, L2-EAD, GC-EAD

  • Client:

To Be Discussed Later

  • Employment Type:

W-2, 1099, C2C

View All Jobs

Similar Jobs

  • Oracle HCM Cloud Technical Consultant - Kansas, KY
  • SAP EWM, MM Consultant - Pennsylvania, PA
  • Supply Chain Planning Consultant - Wawatosa,, WI
  • SAP PP QM Consultant - Dallas, TX
  • Supply Chain Functional Consultant - Albany, NY

Job Description:

  • Brings extensive experience operating large-scale GPU-accelerated AI platforms, deploying and managing LLM inference systems on Kubernetes with strong expertise in Triton Inference Server and TensorRT-LLM.
  • They have repeatedly built and optimized production-grade LLM pipelines with GPU-aware scheduling, load balancing, and real-time performance tuning across multi-node clusters. Their background includes designing containerized microservices, implementing robust deployment workflows, and maintaining operational reliability in mission-critical environments.
  • They have led end-to-end LLMOps processes involving model versioning, engine builds, automated rollouts, and secure runtime controls.
  • The candidate has also developed comprehensive observability for inference systems, using telemetry and custom dashboards to track GPU health, latency, throughput, and service availability.
  • Their work consistently incorporates advanced optimization methods such as mixed precision, quantization, sharding, and batching to improve efficiency. Overall, they bring a strong blend of platform engineering, AI infrastructure, and hands-on operational experience running high-performance LLM systems in production

Basic Info:

  • AI Operations Platform Consultant
  • Experience deploying, managing, operating, and troubleshooting containerized services at scale on Kubernetes for mission-critical applications (OpenShift)
  • Experience with deploying, configuring, and tuning LLMs using TensorRT-LLM and Triton Inference server.
  • Managing MLOps/LLMOps pipelines, using TensorRT-LLM and Triton Inference server to deploy inference services in production
  • Setup and operation of AI inference service monitoring for performance and availability.
  • Experience deploying and troubleshooting LLM models on a containerized platform, monitoring, load balancing, etc.
  • Operation and support of MLOps/LLMOps pipelines, using TensorRT-LLM and Triton Inference server to deploy inference services in production
  • Experience deploying and troubleshooting LLM models on a containerized platform, monitoring, load balancing, etc.
  • Experience with standard processes for operation of a mission critical system – incident management, change management, event management, etc.
  • Managing scalable infrastructure for deploying and managing LLMs
  • Deploying models in production environments, including containerization, microservices, and API design
  • Triton Inference Server, including its architecture, configuration, and deployment.
  • Model Optimization techniques using Triton with TRTLLM
  • Model optimization techniques, including pruning, quantization, and knowledge distillation

Apply Now

Equal Opportunity Employer

MACHINE LEARNING TECHNOLOGIES LLC is an equal opportunity employer inclusive of female, minority, disability and veterans, (M/F/D/V). Hiring, promotion, transfer, compensation, benefits, discipline, termination and all other employment decisions are made without regard to race, color, religion, sex, sexual orientation, gender identity, age, disability, national origin, citizenship/immigration status, veteran status or any other protected status. MACHINE LEARNING TECHNOLOGIES LLC will not make any posting or employment decision that does not comply with applicable laws relating to labor and employment, equal opportunity, employment eligibility requirements or related matters. Nor will MACHINE LEARNING TECHNOLOGIES LLC require in a posting or otherwise U.S. citizenship or lawful permanent residency in the U.S. as a condition of employment except as necessary to comply with law, regulation, executive order, or federal, state, or local government contract

  • Home

|

  • About

|

  • Adaptive AI

|

  • Industries

|

  • Consulting

|

  • Contact Us

© 2008-2018, Machine Learning Technologies LLC. All Rights Reserved.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

AI Operations Platform Consultant
AI Operations Platform Consultant

MACHINE LEARNING TECHNOLOGIES LLC • Jersey City (NJ)

On-site
USD 70,000 - 100,000
AI Operations Platform Consultant
AI Operations Platform Consultant

Cloud Analytics Technologies, LLC • Jersey City (NJ)

On-site
USD 130,000 - 160,000
Senior AI/ML Engineer — Scalable ML & LLM Strategy
Senior AI/ML Engineer — Scalable ML & LLM Strategy

MACHINE LEARNING TECHNOLOGIES LLC • Alpharetta (GA)

On-site
USD 40,000 - 70,000
GenAI Tech Lead (GCP) - LLM & MLOps Architect
GenAI Tech Lead (GCP) - LLM & MLOps Architect

MACHINE LEARNING TECHNOLOGIES LLC • Rosemead (CA)

On-site
USD 120,000 - 150,000
Senior .NET Lead: Kubernetes, Docker & AWS
Senior .NET Lead: Kubernetes, Docker & AWS

MACHINE LEARNING TECHNOLOGIES LLC • Fort Mill (SC)

On-site
USD 85,000 - 110,000
On-Prem LLM Platform Engineer (OpenShift AI / GPU)
On-Prem LLM Platform Engineer (OpenShift AI / GPU)

Infosys Limited • Charlotte (NC)

On-site
USD 80,000 - 120,000
Long-term disability
Health reimbursement accounts
Insurance offerings
+1
Senior Developer
Senior Developer

ICE Clear Europe Limited • Atlanta (GA)

On-site
USD 150,000 - 210,000
GenAI Platform / LLM Inference Optimization Engineer (Cloud)
GenAI Platform / LLM Inference Optimization Engineer (Cloud)

Infosys • Charlotte (NC)

On-site
USD 90,000 - 120,000
Medical/Dental/Vision/Life Insurance
401(k) plan
Paid Time Off
AI Platform Engineer
AI Platform Engineer

Park Place Technologies in • Highland Heights (OH)

On-site
USD 90,000 - 130,000
AI Engineer
AI Engineer

AdventHealth • Town of Florida (NY)

On-site
USD 96,000 - 179,000
Benefits from Day One
Paid Time Off from Day One
403-B Retirement Plan
+5