Remote ML Platform Engineer — Scalable Inference

Bright Vision Technologies

Framingham (MA)

On-site

USD 100,000 - 160,000

Full time

3 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking an ML Platform Engineer for 100% remote work in the U.S. to design, build, and operate high-performance inference platforms for production ML models.

The role emphasizes systems engineering, including request routing, batching, caching, autoscaling, GPU utilization, and observability across diverse model workloads. The ideal candidate brings 10+ years of distributed systems experience, strong Python and systems language skills, and a track record shipping

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • 10+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Proficiency in Python and a systems language (Go, Rust, or C++).
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vcLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
  • Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing.
  • Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints.
  • Build autoscaling and capacity management systems that balance latency, throughput, and cost.
  • Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads.
  • Integrate model serving with API gateways, identity systems, and observability platforms.
  • Implement caching, prompt deduplication, and response reuse strategies where appropriate.
  • Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking.
  • Develop deployment workflows including canary releases, shadow testing, and automated rollback.
  • Operate incident response for high-availability AI services and drive durable reliability improvements.
  • Collaborate with ML and product teams to support new model releases and capability rollouts.
  • Implement security controls including request signing, content filtering, and abuse detection at the serving layer.
  • Document operational procedures, performance characteristics, and tuning guidance for internal teams.
  • Stay current with AI serving research and translate advances into production capabilities.

Skills

Python
Go
Rust
C++
Distributed systems
Performance engineering
Incident response
Capacity planning
Observability

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

vcLLM
TensorRT-LLM
Kubernetes

Job description

Bright Vision Technologies is seeking an ML Platform Engineer for 100% remote work in the U.S. to design, build, and operate high-performance inference platforms for production ML models.

The role emphasizes systems engineering, including request routing, batching, caching, autoscaling, GPU utilization, and observability across diverse model workloads. The ideal candidate brings 10+ years of distributed systems experience, strong Python and systems language skills, and a track record shipping

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Remote ML Systems Engineer — Scalable Inference & GPU Ops
Remote ML Systems Engineer — Scalable Inference & GPU Ops

Bright Vision Technologies • Euless (TX), Bedford (TX)

On-site
USD 145,000 - 165,000
Remote ML Systems Engineer: Scalable AI Inference
Remote ML Systems Engineer: Scalable AI Inference

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior ML Infrastructure Engineer — Remote
Senior ML Infrastructure Engineer — Remote

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 105,000 - 143,000
Senior ML Infra Engineer — Remote AI Serving
Senior ML Infra Engineer — Remote AI Serving

United States Digital Space LLC • United States

Remote
USD 90,000 - 150,000
Remote MLOps Engineer - Scalable AI Inference Platform
Remote MLOps Engineer - Scalable AI Inference Platform

Bright Vision Technologies • Sammamish (WA)

On-site
USD 100,000 - 150,000
Remote ML Infrastructure Engineer: GPU Scale & Platform
Remote ML Infrastructure Engineer: GPU Scale & Platform

Bright Vision Technologies • Sammamish (WA)

On-site
USD 100,000 - 150,000
Remote AI Systems Engineer – Scalable ML Infra
Remote AI Systems Engineer – Scalable ML Infra

Bright Vision Technologies • Mountain View (CA)

On-site
USD 90,000 - 100,000
Remote Model Serving Engineer for High-Scale ML
Remote Model Serving Engineer for High-Scale ML

Bright Vision Technologies • Canton Charter Township (MI)

On-site
USD 74,000 - 98,000
Remote MLOps Engineer: Scale AI Inference & Serving
Remote MLOps Engineer: Scale AI Inference & Serving

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000