Remote Model Serving Engineer for High-Scale ML

Bright Vision Technologies

Canton Charter Township (MI)

On-site

USD 74,000 - 98,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking a Model Serving Engineer to design, build, and operate scalable inference platforms for large ML models in production. You will focus on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and observability across diverse workloads.

The ideal candidate has 7+ years in distributed systems or ML platforms, strong Python and a systems language (Go/Rust/C++), and experience with vLLM or

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design, build, and operate high-performance inference platforms for ML models.
  • Focus on systems engineering aspects of AI deployment including routing, batching, and caching.
  • Ensure autoscaling, GPU utilization, and end-to-end observability across workloads.
  • Balance latency, throughput, and cost while maintaining reliability at scale.

Skills

Python
Go
Rust
C++
Distributed systems
Kubernetes
Cloud platforms
Observability
Performance engineering
Incident response

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

vLLM
TensorRT-LLM

Job description

Bright Vision Technologies is seeking a Model Serving Engineer to design, build, and operate scalable inference platforms for large ML models in production. You will focus on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and observability across diverse workloads.

The ideal candidate has 7+ years in distributed systems or ML platforms, strong Python and a systems language (Go/Rust/C++), and experience with vLLM or

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer — Remote
Senior ML Infrastructure Engineer — Remote

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 105,000 - 143,000
Senior ML Infra Engineer — Remote AI Serving
Senior ML Infra Engineer — Remote AI Serving

United States Digital Space LLC • United States

Remote
USD 90,000 - 150,000
Remote MLOps Engineer: Scale AI Inference & Serving
Remote MLOps Engineer: Scale AI Inference & Serving

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Remote ML Systems Engineer — Scalable Inference & GPU Ops
Remote ML Systems Engineer — Scalable Inference & GPU Ops

Bright Vision Technologies • Euless (TX), Bedford (TX)

On-site
USD 145,000 - 165,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Remote ML Systems Engineer: Scalable AI Inference
Remote ML Systems Engineer: Scalable AI Inference

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Remote MLOps Engineer - Scalable AI Inference Platform
Remote MLOps Engineer - Scalable AI Inference Platform

Bright Vision Technologies • Sammamish (WA)

On-site
USD 100,000 - 150,000
Remote AI Systems Engineer – Scalable ML Infra
Remote AI Systems Engineer – Scalable ML Infra

Bright Vision Technologies • Mountain View (CA)

On-site
USD 90,000 - 100,000
Senior Remote ML Infrastructure Engineer: GPU & Scale
Senior Remote ML Infrastructure Engineer: GPU & Scale

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Inference Performance Engineer: Optimize Model Serving
Inference Performance Engineer: Optimize Model Serving

Adaption • San Francisco (CA)

On-site
USD 180,000 - 240,000
Lunch stipend
Travel stipend (Adaption Passport)
Well-being benefits
+1