Remote MLOps Engineer: Scale AI Inference & Serving

United States Digital Space LLC

United States

Remote

USD 100,000 - 150,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Bright Vision Technologies is seeking a MLOps Engineer to design, build, and operate high‑performance inference platforms for serving large ML models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end‑to‑end observability across diverse model workloads.

The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale,

Qualifications

  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
  • Optimize inference performance using continuous batching, paged attention, speculative decoding, and request multiplexing.
  • Implement multi-tenant routing, rate limiting, and quality-of-service policies across model endpoints.
  • Build autoscaling and capacity management systems that balance latency, throughput, and cost.
  • Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads.
  • Integrate model serving with API gateways, identity systems, and observability platforms.
  • Implement caching, prompt deduplication, and response reuse strategies where appropriate.
  • Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking.
  • Develop deployment workflows including canary releases, shadow testing, and automated rollback.
  • Operate incident response for high-availability AI services and drive durable reliability improvements.
  • Collaborate with ML and product teams to support new model releases and capability rollouts.
  • Implement security controls including request signing, content filtering, and abuse detection at the serving layer.
  • Document operational procedures, performance characteristics, and tuning guidance for internal teams.
  • Stay current with AI serving research and translate advances into production capabilities.

Skills

Distributed systems
Python
Incident response
Communication

Tools

Go
Rust
C++
Kubernetes
TensorRT-LLM
vLLM

Job description

Bright Vision Technologies is seeking a MLOps Engineer to design, build, and operate high‑performance inference platforms for serving large ML models in production. The role focuses on the systems engineering side of AI deployment, including request routing, batching, caching, autoscaling, GPU utilization, and end‑to‑end observability across diverse model workloads.

The ideal candidate brings strong distributed systems and performance engineering expertise, has shipped serving systems at scale,

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote AI Platform Engineer - Scale Production ML
Remote AI Platform Engineer - Scale Production ML

Bright Vision Technologies • Raleigh (NC), Concord (NC)

On-site
USD 100,000 - 180,000
Senior ML Infra Engineer — Remote AI Serving
Senior ML Infra Engineer — Remote AI Serving

United States Digital Space LLC • United States

Remote
USD 90,000 - 150,000
Remote ML Systems Engineer: Scale AI Inference
Remote ML Systems Engineer: Scale AI Inference

Bright Vision Technologies • Flower Mound (TX)

On-site
USD 145,000 - 165,000
Remote ML Systems Engineer: Scalable AI Inference
Remote ML Systems Engineer: Scalable AI Inference

United States Digital Space LLC • United States

Remote
USD 100,000 - 150,000
Remote ML Systems Engineer – High-Performance Inference
Remote ML Systems Engineer – High-Performance Inference

Triwill Group • United States

Remote
USD 145,000 - 165,000
Remote AI Systems Engineer - Scalable ML Infra
Remote AI Systems Engineer - Scalable ML Infra

Bright Vision Technologies • Santa Clara (CA)

On-site
USD 90,000 - 100,000
Senior Remote ML Infrastructure Engineer: GPU & Scale
Senior Remote ML Infrastructure Engineer: GPU & Scale

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Remote AI Operations Engineer for Large-Scale ML Pipelines
Remote AI Operations Engineer for Large-Scale ML Pipelines

Bright Vision Technologies • Flower Mound (TX)

On-site
USD 150,000 - 165,000
Remote ML Performance Engineer: Optimize Training Inference
Remote ML Performance Engineer: Optimize Training Inference

Bright Vision Technologies • Bellevue (WA)

On-site
USD 100,000 - 150,000
Remote MLOps Engineer: Scale AI Deployments & Pipelines
Remote MLOps Engineer: Scale AI Deployments & Pipelines

AgileEngine, LLC. • Town of Vermont (WI)

On-site
USD 140,000 - 190,000
Growth without limits
Competitive compensation
Remote flexibility: 100% remote with a
+3