Remote ML Platform Engineer - Scalable Inference & Systems

Bright Vision Technologies

Edison (NJ)

Remote

USD 100,000 - 160,000

Full time

3 days ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an ML Platform Engineer to design, build, and operate high‑performance inference platforms for production ML models. This role emphasizes systems engineering for AI deployment, covering request routing, batching, caching, autoscaling, GPU utilization, and end‑to‑end observability across model workloads.

The ideal candidate has 10+ years in distributed systems or ML platforms, strong Python and Go/Rust/C++, and hands‑on experience with large‑model frameworks

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 10+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Experience operating high‑throughput, low‑latency services in production.
  • Hands‑on experience with LLM or large model inference frameworks such as vcLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
  • Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing.
  • Implement multi‑tenant routing, rate limiting, and quality‑of‑service policies across model endpoints.
  • Build autoscaling and capacity management systems that balance latency, throughput, and cost.
  • Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads.
  • Integrate model serving with API gateways, identity systems, and observability platforms.
  • Implement caching, prompt deduplication, and response reuse strategies where appropriate.
  • Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking.
  • Develop deployment workflows including canary releases, shadow testing, and automated rollback.
  • Operate incident response for high‑availability AI services and drive durable reliability improvements.
  • Collaborate with ML and product teams to support new model releases and capability rollouts.
  • Implement security controls including request signing, content filtering, and abuse detection at the serving layer.
  • Document operational procedures, performance characteristics, and tuning guidance for internal teams.
  • Stay current with AI serving research and translate advances into production capabilities.

Skills

Distributed systems
Python
Go
Rust
C++
Kubernetes
Observability
Performance engineering
Incident response
GPU inference
ML serving
API gateways
Batching
Caching

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

vcLLM
TensorRT-LLM

Job description

Bright Vision Technologies is seeking an ML Platform Engineer to design, build, and operate high‑performance inference platforms for production ML models. This role emphasizes systems engineering for AI deployment, covering request routing, batching, caching, autoscaling, GPU utilization, and end‑to‑end observability across model workloads.

The ideal candidate has 10+ years in distributed systems or ML platforms, strong Python and Go/Rust/C++, and hands‑on experience with large‑model frameworks

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Inference Platform Engineer — Remote
Senior ML Inference Platform Engineer — Remote

Bright Vision Technologies • Beaverton (OR)

Remote
USD 105,000 - 143,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Senior ML Model Serving Engineer - Remote
Senior ML Model Serving Engineer - Remote

Bright Vision Technologies • United States

Remote
USD 74,000 - 98,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Remote MLOps Engineer - Build Scalable AI Serving
Remote MLOps Engineer - Build Scalable AI Serving

Bright Vision Technologies • United States

Remote
USD 100,000 - 150,000
Remote AI Systems Engineer — Scalable GPU Infra
Remote AI Systems Engineer — Scalable GPU Infra

Bright Vision Technologies • United States

Remote
USD 90,000 - 100,000
Senior AI Platform Engineer - Remote
Senior AI Platform Engineer - Remote

Bright Vision Technologies • Edison (NJ)

Remote
USD 100,000 - 160,000
Remote MLOps Engineer - Scale AI Inference Platforms
Remote MLOps Engineer - Scale AI Inference Platforms

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Remote ML Data Engineer - Scale-Power Data Pipelines
Remote ML Data Engineer - Scale-Power Data Pipelines

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Remote ML Data Engineer — Scale AI Data Pipelines
Remote ML Data Engineer — Scale AI Data Pipelines

Bright Vision Technologies • Folsom (CA)

Remote
USD 80,000 - 100,000