Senior MLOps Engineer for Scalable AI Serving (Remote)

Bright Vision Technologies

Redmond (WA)

On-site

USD 100,000 - 150,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking an MLOps Engineer to design, build, and operate high-performance inference platforms for production ML models. You will focus on the systems engineering aspects of AI deployment, including routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads.

The ideal candidate has strong distributed systems and performance engineering experience, has shipped serving systems at scale, and can balance latency,

Qualifications

  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Hands‑on experience with LLM or large model inference frameworks such as vLLM or TensorRT‑LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
  • Optimize inference performance using batching, paging, and caching strategies.
  • Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints.
  • Build autoscaling and capacity management systems balancing latency, throughput, and cost.
  • Tune GPU utilization and memory management for large-scale serving workloads.
  • Integrate model serving with API gateways, identity systems, and observability platforms.
  • Develop deployment workflows including canary releases, shadow testing, and automated rollback.
  • Operate incident response for high-availability AI services and drive reliability improvements.

Skills

Python
Go
Rust
C++
Distributed systems
Kubernetes
vLLM
TensorRT-LLM
GPU architecture
Observability
Incident response
Cloud platforms

Education

Bachelor's or Master's in Computer Science

Job description

Bright Vision Technologies is seeking an MLOps Engineer to design, build, and operate high-performance inference platforms for production ML models. You will focus on the systems engineering aspects of AI deployment, including routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse model workloads.

The ideal candidate has strong distributed systems and performance engineering experience, has shipped serving systems at scale, and can balance latency,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote MLOps Engineer: Scale AI Inference & Serving
Remote MLOps Engineer: Scale AI Inference & Serving

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior ML Infra Engineer — Remote AI Serving
Senior ML Infra Engineer — Remote AI Serving

United States Digital Space LLC • United States

Remote
USD 90,000 - 150,000
Remote AI Systems Engineer - Scalable ML Infra
Remote AI Systems Engineer - Scalable ML Infra

Bright Vision Technologies • Santa Clara (CA)

On-site
USD 90,000 - 100,000
Remote ML Systems Engineer — Scalable Inference & GPU Ops
Remote ML Systems Engineer — Scalable Inference & GPU Ops

Bright Vision Technologies • Euless (TX), Bedford (TX)

On-site
USD 145,000 - 165,000
Remote Model Serving Engineer — High-Performance AI
Remote Model Serving Engineer — High-Performance AI

Bright Vision Technologies • Novi (MI)

On-site
USD 74,000 - 98,000
Remote ML Systems Engineer: Scalable AI Inference
Remote ML Systems Engineer: Scalable AI Inference

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Senior Remote ML Infrastructure Engineer: GPU & Scale
Senior Remote ML Infrastructure Engineer: GPU & Scale

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Senior AI Infrastructure Engineer — Remote
Senior AI Infrastructure Engineer — Remote

Bright Vision Technologies • Sterling (VA)

On-site
USD 100,000 - 150,000
Remote AI Operations Engineer for Large-Scale ML Pipelines
Remote AI Operations Engineer for Large-Scale ML Pipelines

Bright Vision Technologies • Flower Mound (TX)

On-site
USD 150,000 - 165,000