Lead MLOps Engineer: Remote Inference Platform

Bright Vision Technologies

Eden Prairie (MN)

Remote

USD 100,000 - 150,000

Full time

8 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an experienced MLOps Engineer to design, build, and operate high-performance inference platforms for production ML models. The role emphasizes distributed systems, scalability, and observability, with a remote-first structure across the United States.

The candidate should have 6+ years of experience in ML infrastructure and be proficient in Python and a systems language (Go, Rust, or C++), with exposure to LLM inference frameworks and GPU optimization.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related field.
  • 6+ years of experience in distributed systems or ML platform engineering.
  • Proficiency in Python and a systems language (Go, Rust, or C++).
  • Experience with high-throughput, low-latency production services.
  • Hands-on with LLM inference frameworks such as vLLM or TensorRT-LLM.
  • Understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks (metrics, tracing, logging).
  • Strong grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
  • Optimize inference performance using batching, paged attention, speculative decoding, and request multiplexing.
  • Implement multi-tenant routing, rate limiting, and QoS across model endpoints.
  • Build autoscaling and capacity management systems balancing latency, throughput, and cost.
  • Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads.
  • Integrate model serving with API gateways, identity systems, and observability platforms.
  • Implement caching, prompt deduplication, and response reuse strategies where appropriate.
  • Drive end-to-end observability including latency histograms, queue dynamics, GPU utilization, and error tracking.
  • Develop deployment workflows including canary releases, shadow testing, and automated rollback.
  • Operate incident response for high-availability AI services and drive durable reliability improvements.
  • Collaborate with ML and product teams to support new model releases and capability rollouts.
  • Implement security controls at the serving layer.
  • Document operational procedures, performance characteristics, and tuning guidance for internal teams.
  • Stay current with AI serving research and translate advances into production capabilities.

Skills

Python
Go
Rust
C++
Distributed systems
Low latency
High throughput
Performance engineering
Incident response
Observability
Kubernetes
Autoscaling
Cloud platforms
LLM inference

Education

Bachelor’s or Master’s in CS or related

Tools

vLLM
TensorRT-LLM

Job description

Bright Vision Technologies is seeking an experienced MLOps Engineer to design, build, and operate high-performance inference platforms for production ML models. The role emphasizes distributed systems, scalability, and observability, with a remote-first structure across the United States.

The candidate should have 6+ years of experience in ML infrastructure and be proficient in Python and a systems language (Go, Rust, or C++), with exposure to LLM inference frameworks and GPU optimization.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior ML Infrastructure Engineer — Remote
Senior ML Infrastructure Engineer — Remote

Bright Vision Technologies • Hillsboro (OR)

Remote
USD 105,000 - 143,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Senior AI Platform Engineer | Cloud-Native MLOps (Remote)
Senior AI Platform Engineer | Cloud-Native MLOps (Remote)

Bright Vision Technologies • United States

Remote
USD 130,000 - 180,000
Remote ML Platform Engineer for Scalable Inference
Remote ML Platform Engineer for Scalable Inference

Bright Vision Technologies • Cranberry Township

Remote
USD 100,000 - 160,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Remote ML Systems Engineer for Scalable Inference
Remote ML Systems Engineer for Scalable Inference

Bright Vision Technologies • Pflugerville (TX)

Remote
USD 145,000 - 165,000
MLOps Engineer
MLOps Engineer

Bright Vision Technologies • Eden Prairie (MN)

Remote
USD 100,000 - 150,000
ML Platform Engineer
ML Platform Engineer

Bright Vision Technologies • Cranberry Township

Remote
USD 100,000 - 160,000
Senior Remote ML Infrastructure Engineer
Senior Remote ML Infrastructure Engineer

Bright Vision Technologies • Kirkland (WA)

Remote
USD 100,000 - 150,000
Remote Data Platform Engineer for ML Serving at Scale
Remote Data Platform Engineer for ML Serving at Scale

Bright Vision Technologies • Reston (VA)

On-site
USD 100,000 - 150,000