Remote Model Serving Engineer: Scalable AI Inference

Socket.dev

Ann Arbor (MI)

On-site

USD 74,000 - 98,000

Full time

4 days ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking a Model Serving Engineer to design, build, and operate high-performance inference platforms for production machine learning models. This role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization and end-to-end observability across diverse model workloads.

The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust or C++.
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design, build, and operate high-performance inference platforms for production ML models.
  • Focus on systems engineering for AI deployment, including routing, batching, caching, and autoscaling.
  • Optimize GPU utilization and ensure end-to-end observability across diverse model workloads.
  • Collaborate across teams to balance latency, throughput, and cost in serving systems.

Skills

Distributed systems
Python
Go
Rust
C++
Low-latency systems
ML serving
LLM frameworks
GPU acceleration
Kubernetes
Autoscaling
Cloud platforms
Observability
Performance engineering
Capacity planning
Incident response
Communication

Education

Bachelor's or Master’s in Computer Science or related field

Tools

vLLM
TensorRT-LLM

Job description

Bright Vision Technologies is seeking a Model Serving Engineer to design, build, and operate high-performance inference platforms for production machine learning models. This role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization and end-to-end observability across diverse model workloads.

The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Model Serving Engineer – Remote AI Infra
Senior Model Serving Engineer – Remote AI Infra

United States Digital Space LLC • United States

Remote
USD 74,000 - 98,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Remote ML Platform Engineer for Scalable Inference
Remote ML Platform Engineer for Scalable Inference

Bright Vision Technologies • Cranberry Township

Remote
USD 100,000 - 160,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Model Serving Engineer
Model Serving Engineer

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Senior ML Infrastructure Engineer — Remote
Senior ML Infrastructure Engineer — Remote

Bright Vision Technologies • Hillsboro (OR)

Remote
USD 105,000 - 143,000
Senior AI Platform Engineer | Cloud-Native MLOps (Remote)
Senior AI Platform Engineer | Cloud-Native MLOps (Remote)

Bright Vision Technologies • United States

Remote
USD 130,000 - 180,000
Inference Systems Engineer — High-Performance AI Serving
Inference Systems Engineer — High-Performance AI Serving

River AI • Palo Alto (CA)

On-site
USD 200,000 - 420,000
Health insurance
Dental insurance
Vision insurance
+3
Senior AI Systems Engineer — Scalable Model Serving
Senior AI Systems Engineer — Scalable Model Serving

Apply • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Remote Senior ML Systems Engineer - Model Serving
Remote Senior ML Systems Engineer - Model Serving

Atlassian • Northern (KY)

Hybrid
USD 149,000 - 235,000
Benefits and bonuses
Equity