A leading AI infrastructure company located in New Jersey is seeking an experienced AI Operations Platform Consultant to lead and optimize large-scale GPU-accelerated AI platforms. The ideal candidate will have a strong background in deploying and managing LLM inference systems on Kubernetes, with expertise in TensorRT-LLM and Triton Inference Server. Responsibilities include managing production-grade LLM pipelines and ensuring operational reliability. This position is part of a team committed to diversity and equal opportunity.
Qualifications
Extensive experience operating large-scale GPU-accelerated AI platforms.
Strong expertise in deploying and managing LLM inference systems.
Experience with AI inference service monitoring and performance optimization.
Responsibilities
Lead production-grade LLM pipelines with performance tuning.
Manage MLOps processes for deploying inference services.
Develop observability for inference systems using telemetry.
Skills
Kubernetes
TensorRT-LLM
Triton Inference Server
MLOps
AI operations
Containerization
Job description
A leading AI infrastructure company located in New Jersey is seeking an experienced AI Operations Platform Consultant to lead and optimize large-scale GPU-accelerated AI platforms. The ideal candidate will have a strong background in deploying and managing LLM inference systems on Kubernetes, with expertise in TensorRT-LLM and Triton Inference Server. Responsibilities include managing production-grade LLM pipelines and ensuring operational reliability. This position is part of a team committed to diversity and equal opportunity.