Remote MLOps Engineer: LLM Serving & GPU Profiling

24 Mag

New York (NY)

Remote

USD 124,000 - 165,000

Full time

10 days ago
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

24 MAG is seeking an experienced MLOps Engineer for a remote, full-time engagement focused on LLM systems, GPU kernels, profiling, and distributed debugging.

You will develop reference solutions, evaluate model outputs, and help set evaluation standards across GPU kernels, profiling, and high‑throughput serving. The role emphasizes production infrastructure experience over pure modelling.

Qualifications

  • 2 years of hands-on professional experience in ML systems, ML infrastructure, model serving, or accelerator‑performance engineering
  • Strong practical experience in at least one of GPU kernel programming, performance profiling, distributed debugging, or high‑throughput inference serving
  • Production experience with JAX and/or PyTorch
  • Familiarity with CUDA, Triton, Pallas, or comparable accelerator‑programming technologies
  • Experience with profiling tools such as Kineto, torch.prof profiler, Nsight, XLA, or JAX profiler
  • Experience debugging distributed or accelerator‑bound workloads
  • Familiarity with vLLM, SGLang, TensorRT‑LLM, Ray Serve, KV cache, paged attention, or continuous batching
  • Framework‑level experience with custom operators, FSDP, DDP, DeepSpeed, Megatron, compiler, or graph‑level work is highly valuable
  • Familiarity with accelerators such as A100, H100, B200, or TPU
  • Ability to reason precisely about throughput, latency, memory, and compute trade‑offs
  • Demonstrable professional progression in ML infrastructure or systems engineering
  • Strong written communication and ability to explain complex technical decisions clearly

Responsibilities

  • Design GPU/accelerator workloads and kernel tasks.
  • Develop CUDA, Triton, or Pallas‑based solutions.
  • Evaluate kernel optimisation for correctness and efficiency.
  • Analyse memory, compute, and hardware utilisation trade‑offs.
  • Apply accelerator engineering judgement to model‑generated solutions.
  • Create tasks for performance profiling and trace interpretation.
  • Analyse outputs from Kineto, torch.profiler, Nsight, XLA, or JAX profilers.
  • Identify bottlenecks across compute, memory, and scheduling.
  • Evaluate throughput, latency, and utilisation characteristics.
  • Produce clear analyses explaining performance behaviour.
  • Design scenarios for distributed or accelerator‑bound ML workloads.
  • Diagnose failures across training and inference infra.
  • Evaluate FSDP, DDP, DeepSpeed, Megatron, and related systems.
  • Review framework‑level and distributed‑system troubleshooting.
  • Identify plausible but incorrect explanations or fixes.
  • Develop tasks for high‑throughput LLM serving.
  • Apply vLLM, SGLang, TensorRT‑LLM, Ray Serve.
  • Evaluate KV‑cache, paged‑attention, and continuous batching.
  • Analyse serving architectures for latency, memory, scalability.
  • Review production‑grade inference deployment approaches.
  • Evaluate MLOps and ML‑systems tasks and proposals.
  • Provide precise written feedback for technical reviews.
  • Develop rubrics and evaluation frameworks for systems work.
  • Help research/engineering teams close knowledge gaps.
  • Collaborate with SMEs to maintain data quality

Skills

ML systems engineering
GPU kernel programming
Performance profiling
Distributed debugging
High-throughput inference
Technical writing

Tools

CUDA
Triton
Pallas
Kineto
torch.profiler
Nsight
XLA
JAX
PyTorch
TensorRT-LLM
Ray Serve
vLLM

Job description

24 MAG is seeking an experienced MLOps Engineer for a remote, full-time engagement focused on LLM systems, GPU kernels, profiling, and distributed debugging.

You will develop reference solutions, evaluate model outputs, and help set evaluation standards across GPU kernels, profiling, and high‑throughput serving. The role emphasizes production infrastructure experience over pure modelling.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour
Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour

24 Mag • New York (NY)

Remote
USD 124,000 - 165,000
Remote MLOps Engineer: Scalable ML Systems & GPU Kernels
Remote MLOps Engineer: Scalable ML Systems & GPU Kernels

Remote Jobs • United States

Remote
USD 96,000 - 152,000
MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago
MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago

Benture • California (MO), Northern (KY)

Hybrid
USD 124,000 - 165,000
MLOps Engineer: JAX/PyTorch & GPU Kernel Expert — Remote
MLOps Engineer: JAX/PyTorch & GPU Kernel Expert — Remote

Weekday AI (YC W21) • United States

On-site
USD 96,000 - 152,000
Remote MLOps LLM Systems Reviewer (Code Quality)
Remote MLOps LLM Systems Reviewer (Code Quality)

AuraOne • United States

On-site
USD 124,000 - 165,000
Remote MLOps Engineer for Scalable LLM Systems
Remote MLOps Engineer for Scalable LLM Systems

Benture • California (MO), Northern (KY)

Hybrid
USD 124,000 - 165,000
Remote MLOps Engineer - 4-Day Week
Remote MLOps Engineer - 4-Day Week

MaziCTools • United States

Remote
USD 90,000 - 130,000
Remote-Friendly MLOps Engineer - CV & Edge GPU Production
Remote-Friendly MLOps Engineer - CV & Edge GPU Production

AgileEngine, LLC. • Houston (TX)

On-site
USD 120,000 - 190,000
Professional growth
Competitive compensation
A selection of exciting projects
+1
MLOps Engineer
MLOps Engineer

Blue Signal Search • Santa Clara (CA)

On-site
USD 140,000 - 190,000
Advanced GPU infra exposure
Collaborative engineering culture
Open source AI frameworks access
+2
MLOps Engineer — CV Pipelines, GPU, Edge/Cloud (Remote)
MLOps Engineer — CV Pipelines, GPU, Edge/Cloud (Remote)

AgileEngine, LLC. • Jacksonville (FL)

On-site
USD 120,000 - 180,000
Professional growth
Competitive compensation
A selection of exciting projects
+1