Remote MLOps Engineer — Scalable AI Inference

Bright Vision Technologies

Maple Grove (MN)

Remote

USD 100,000 - 150,000

Full time

8 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an MLOps Engineer to design, build, and operate high-performance inference platforms for large ML models in production. This role emphasizes systems engineering, request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse workloads.

The ideal candidate has strong distributed systems experience, shipped serving systems at scale, and understands latency, throughput, and cost trade-offs in ML serving.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related field.
  • Six or more years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language such as Go, Rust, or C++.
  • Experience with high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
  • Optimize inference performance using batching, paging, and caching strategies.
  • Implement multi-tenant routing, rate limiting, and QoS policies across model endpoints.
  • Build autoscaling and capacity management systems balancing latency, throughput, and cost.
  • Tune GPU utilization and memory management for serving workloads.
  • Integrate model serving with API gateways, identity systems, and observability platforms.
  • Develop caching and response reuse strategies where appropriate.
  • Drive end-to-end observability including latency histograms and queue dynamics.
  • Develop deployment workflows with canary releases, shadow testing, and rollback.
  • Operate incident response for high-availability AI services and drive reliability improvements.
  • Collaborate with ML and product teams on new model releases and rollouts.
  • Implement security controls at the serving layer and document procedures.

Skills

Python
Distributed systems
Performance engineering
Incident response
ML model serving
Communication
System monitoring

Education

Bachelor’s or Master’s in CS or related

Tools

Kubernetes
TensorRT/LLM frameworks
vLLM
Go/Rust/C++
CUDA/GPU architecture

Job description

Bright Vision Technologies is seeking an MLOps Engineer to design, build, and operate high-performance inference platforms for large ML models in production. This role emphasizes systems engineering, request routing, batching, caching, autoscaling, GPU utilization, and end-to-end observability across diverse workloads.

The ideal candidate has strong distributed systems experience, shipped serving systems at scale, and understands latency, throughput, and cost trade-offs in ML serving.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior ML Inference Platform Engineer — Remote
Senior ML Inference Platform Engineer — Remote

Bright Vision Technologies • Beaverton (OR)

Remote
USD 105,000 - 143,000
Senior AI Platform Engineer | Cloud-Native MLOps (Remote)
Senior AI Platform Engineer | Cloud-Native MLOps (Remote)

Bright Vision Technologies • United States

Remote
USD 130,000 - 180,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Senior Model Serving Engineer – Remote AI Infra
Senior Model Serving Engineer – Remote AI Infra

United States Digital Space LLC • United States

Remote
USD 74,000 - 98,000
Remote AI Infra Engineer — Scale GPU Clusters
Remote AI Infra Engineer — Scale GPU Clusters

Bright Vision Technologies • Monroeville

Remote
USD 100,000 - 160,000
Remote ML Performance Engineer: AI Throughput Optimizer
Remote ML Performance Engineer: AI Throughput Optimizer

Bright Vision Technologies • Maple Grove (MN)

Remote
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Bright Vision Technologies • Maple Grove (MN)

Remote
USD 100,000 - 150,000
Senior ML Infra Engineer - Scale GPU Clusters, Remote
Senior ML Infra Engineer - Scale GPU Clusters, Remote

AI Breaking Wire • San Francisco (CA), Northern (KY)

Hybrid
USD 320,000 - 500,000
Equity
Medical/Dental/Vision coverage
Unlimited PTO
+1
AI Inference Engineer
AI Inference Engineer

Acceler8 Talent • San Francisco (CA)

On-site
USD 150,000 - 230,000