Remote MLOps Engineer - Build Scalable AI Serving

Bright Vision Technologies

United States

Remote

USD 100,000 - 150,000

Full time

7 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Bright Vision Technologies is seeking an MLOps Engineer to design and operate high-performance inference platforms for large ML models in production. You will focus on distributed systems, latency, throughput, and observability across diverse workloads, collaborating with ML and product teams to roll out new models.

The role emphasizes systems engineering for AI deployment, including batching, caching, autoscaling, and security at the serving layer, with a strong emphasis on reliability and cost

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related field.
  • 6+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and a systems language (Go, Rust, or C++).
  • Deep experience operating high-throughput, low-latency services in production.
  • Hands-on experience with LLM or large model inference frameworks such as vLLM or TensorRT-LLM.
  • Strong understanding of GPU architecture, memory hierarchies, and accelerator utilization.
  • Familiarity with Kubernetes, autoscaling, and modern cloud platforms.
  • Experience with observability stacks including metrics, tracing, and structured logging.
  • Solid grounding in performance engineering and capacity planning.
  • Strong communication and incident response skills.

Responsibilities

  • Design and operate model serving platforms supporting diverse workloads including LLMs, vision models, and recommendation systems.
  • Optimize inference performance using batching, paged attention, speculative decoding, and request multiplexing.
  • Implement multi-tenant routing, rate limiting, and QoS policies across model endpoints.
  • Build autoscaling and capacity management systems balancing latency, throughput, and cost.
  • Tune GPU utilization, memory management, and KV cache strategies for LLM serving workloads.
  • Integrate model serving with API gateways, identity systems, and observability platforms.
  • Implement caching, prompt deduplication, and response reuse strategies.
  • Drive end-to-end observability including latency histograms and error tracking.
  • Develop deployment workflows including canary releases, shadow testing, and automated rollback.
  • Operate incident response for high-availability AI services and drive reliability improvements.
  • Collaborate with ML and product teams to support new model releases and capability rollouts.
  • Implement security controls including request signing and abuse detection at the serving layer.
  • Document operational procedures and performance characteristics for internal teams.
  • Stay current with AI serving research and translate advances into production capabilities.

Skills

Python
Go
Rust
C++
Kubernetes
observability
performance engineering

Education

Bachelor’s or Master’s degree in Computer Science

Tools

vLLM
TensorRT-LLM

Job description

Bright Vision Technologies is seeking an MLOps Engineer to design and operate high-performance inference platforms for large ML models in production. You will focus on distributed systems, latency, throughput, and observability across diverse workloads, collaborating with ML and product teams to roll out new models.

The role emphasizes systems engineering for AI deployment, including batching, caching, autoscaling, and security at the serving layer, with a strong emphasis on reliability and cost

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote MLOps Engineer — Scalable AI Inference
Remote MLOps Engineer — Scalable AI Inference

Bright Vision Technologies • Maple Grove (MN)

Remote
USD 100,000 - 150,000
Remote MLOps Engineer - Scale AI Inference Platforms
Remote MLOps Engineer - Scale AI Inference Platforms

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Senior ML Inference Platform Engineer — Remote
Senior ML Inference Platform Engineer — Remote

Bright Vision Technologies • Beaverton (OR)

Remote
USD 105,000 - 143,000
Senior ML Model Serving Engineer (Remote)
Senior ML Model Serving Engineer (Remote)

NEPSE Trading • Northern (KY)

Hybrid
USD 74,000 - 98,000
Remote ML Infra Engineer - High-Performance Inference
Remote ML Infra Engineer - High-Performance Inference

Bright Vision Technologies • Mountain View (CA)

On-site
USD 105,000 - 143,000
Remote MLOps Engineer - Automate & Scale ML Pipelines
Remote MLOps Engineer - Automate & Scale ML Pipelines

brightvisiontechnologies • Cranberry Township

Remote
USD 100,000 - 150,000
Senior Model Serving Engineer – Remote AI Infra
Senior Model Serving Engineer – Remote AI Infra

United States Digital Space LLC • United States

Remote
USD 74,000 - 98,000
MLOps Engineer — Scalable AI Infra & Deployment, Equity
MLOps Engineer — Scalable AI Infra & Deployment, Equity

Fundamental • United States

Remote
USD 180,000 - 260,000
Salary + equity
Health coverage for you and dependents
Parental leave for all
+2
MLOps Engineer
MLOps Engineer

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
Remote Staff ML Platform Engineer (MLOps) - Build AI Infra
Remote Staff ML Platform Engineer (MLOps) - Build AI Infra

futurefitai • United States

Remote
USD 172,000 - 215,000
Remote-friendly
Competitive compensation