Remote Model Serving Engineer - Scale ML Inference

Bright Vision Technologies

Farmington Hills (MI)

On-site

USD 74,000 - 98,000

Full time

9 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking a Model Serving Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. The role focuses on the systems engineering side of AI deployment, including routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads.

The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale,

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design, build and operate high-performance inference platforms for serving large ML models.
  • Handle request routing, batching, caching, autoscaling, GPU utilization, and observability across workloads.
  • Optimize for latency, throughput, cost, and quality in production ML serving.
  • Collaborate with distributed teams to deploy, monitor, and maintain serving systems.

Skills

Distributed systems
Python
Performance engineering
Communication skills
Incident response
Cloud platforms
GPU optimization

Education

Bachelor’s or Master’s degree in Computer Science or related field

Tools

Kubernetes
Go
Rust
C++
TensorRT-LLM
vLLM
Kubernetes

Job description

Bright Vision Technologies is seeking a Model Serving Engineer to design, build, and operate high-performance inference platforms for serving large ML models in production. The role focuses on the systems engineering side of AI deployment, including routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads.

The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Senior Model Serving Engineer – Remote AI Infra
Senior Model Serving Engineer – Remote AI Infra

United States Digital Space LLC • United States

Remote
USD 74,000 - 98,000
Remote ML Infrastructure Engineer: Scale AI Inference
Remote ML Infrastructure Engineer: Scale AI Inference

JobCubby • Redwood City (CA)

Hybrid
USD 105,000 - 143,000
Remote ML Systems Engineer Scalable Inference Platforms
Remote ML Systems Engineer Scalable Inference Platforms

Bright Vision Technologies • Round Rock (TX)

On-site
USD 145,000 - 165,000
Senior ML Infrastructure Engineer – Remote
Senior ML Infrastructure Engineer – Remote

Bright Vision Technologies • Redwood City (CA), San Mateo (CA)

On-site
USD 105,000 - 143,000
Remote ML Platform Engineer — Scalable Inference Systems
Remote ML Platform Engineer — Scalable Inference Systems

Bright Vision Technologies • Nashua (NH)

On-site
USD 100,000 - 160,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Remote AI Platform Engineer - Scalable ML Inference
Remote AI Platform Engineer - Scalable ML Inference

Bright Vision Technologies • Charlotte (NC)

On-site
USD 100,000 - 150,000
Remote AI Systems Engineer — Scale GPU ML Infra
Remote AI Systems Engineer — Scale GPU ML Infra

Bright Vision Technologies • Palo Alto (CA)

On-site
USD 90,000 - 100,000
Model Serving Engineer
Model Serving Engineer

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000