Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour

24 Mag

New York (NY)

Remote

USD 124,000 - 165,000

Full time

10 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

24 MAG is seeking an experienced MLOps Engineer for a remote, full-time engagement focused on LLM systems, GPU kernels, profiling, and distributed debugging.

You will develop reference solutions, evaluate model outputs, and help set evaluation standards across GPU kernels, profiling, and high‑throughput serving. The role emphasizes production infrastructure experience over pure modelling.

Qualifications

  • 2 years of hands-on professional experience in ML systems, ML infrastructure, model serving, or accelerator‑performance engineering
  • Strong practical experience in at least one of GPU kernel programming, performance profiling, distributed debugging, or high‑throughput inference serving
  • Production experience with JAX and/or PyTorch
  • Familiarity with CUDA, Triton, Pallas, or comparable accelerator‑programming technologies
  • Experience with profiling tools such as Kineto, torch.prof profiler, Nsight, XLA, or JAX profiler
  • Experience debugging distributed or accelerator‑bound workloads
  • Familiarity with vLLM, SGLang, TensorRT‑LLM, Ray Serve, KV cache, paged attention, or continuous batching
  • Framework‑level experience with custom operators, FSDP, DDP, DeepSpeed, Megatron, compiler, or graph‑level work is highly valuable
  • Familiarity with accelerators such as A100, H100, B200, or TPU
  • Ability to reason precisely about throughput, latency, memory, and compute trade‑offs
  • Demonstrable professional progression in ML infrastructure or systems engineering
  • Strong written communication and ability to explain complex technical decisions clearly

Responsibilities

  • Design GPU/accelerator workloads and kernel tasks.
  • Develop CUDA, Triton, or Pallas‑based solutions.
  • Evaluate kernel optimisation for correctness and efficiency.
  • Analyse memory, compute, and hardware utilisation trade‑offs.
  • Apply accelerator engineering judgement to model‑generated solutions.
  • Create tasks for performance profiling and trace interpretation.
  • Analyse outputs from Kineto, torch.profiler, Nsight, XLA, or JAX profilers.
  • Identify bottlenecks across compute, memory, and scheduling.
  • Evaluate throughput, latency, and utilisation characteristics.
  • Produce clear analyses explaining performance behaviour.
  • Design scenarios for distributed or accelerator‑bound ML workloads.
  • Diagnose failures across training and inference infra.
  • Evaluate FSDP, DDP, DeepSpeed, Megatron, and related systems.
  • Review framework‑level and distributed‑system troubleshooting.
  • Identify plausible but incorrect explanations or fixes.
  • Develop tasks for high‑throughput LLM serving.
  • Apply vLLM, SGLang, TensorRT‑LLM, Ray Serve.
  • Evaluate KV‑cache, paged‑attention, and continuous batching.
  • Analyse serving architectures for latency, memory, scalability.
  • Review production‑grade inference deployment approaches.
  • Evaluate MLOps and ML‑systems tasks and proposals.
  • Provide precise written feedback for technical reviews.
  • Develop rubrics and evaluation frameworks for systems work.
  • Help research/engineering teams close knowledge gaps.
  • Collaborate with SMEs to maintain data quality

Skills

ML systems engineering
GPU kernel programming
Performance profiling
Distributed debugging
High-throughput inference
Technical writing

Tools

CUDA
Triton
Pallas
Kineto
torch.profiler
Nsight
XLA
JAX
PyTorch
TensorRT-LLM
Ray Serve
vLLM

Job description

About the job Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour

We are sharing a full-time opportunity for experienced MLOps Engineers with hands‑on expertise in large language model infrastructure, GPU acceleration, performance profiling, distributed‑system debugging, and high‑throughput inference serving to contribute to advanced AI training and evaluation initiatives.

Selected professionals will develop challenging ML‑systems tasks, produce technically rigorous reference solutions, evaluate model‑generated outputs, and help establish evaluation standards across GPU kernels, profiling, debugging, and LLM serving. This is a hands‑on systems role intended for engineers with production infrastructure experience rather than primarily applied modelling or data‑science backgrounds.

Key Responsibilities

GPU Kernels & Accelerator Engineering

  • Design technically challenging tasks involving GPU and accelerator workloads
  • Develop solutions covering CUDA, Triton, Pallas, or comparable kernel technologies
  • Evaluate kernel‑level optimisation approaches for correctness and efficiency
  • Analyse memory, compute, and hardware‑utilisation trade‑offs
  • Apply practical accelerator engineering judgement to model‑generated solutions

Performance Profiling & Trace Analysis

  • Develop tasks involving performance profiling and trace interpretation
  • Analyse outputs from tools such as Kineto, torch.profiler, Nsight, XLA, or JAX profilers
  • Identify bottlenecks across compute, memory, communication, and scheduling
  • Evaluate throughput, latency, and utilisation characteristics
  • Produce clear reference analyses explaining observed performance behaviour

Distributed Systems & Workload Debugging

  • Design scenarios involving distributed or accelerator‑bound ML workloads
  • Diagnose failures across training and inference infrastructure
  • Evaluate reasoning around FSDP, DDP, DeepSpeed, Megatron, and related systems
  • Review framework‑level and distributed‑system troubleshooting approaches
  • Identify technically plausible but incorrect explanations or proposed fixes

LLM Inference & Serving

  • Develop and assess tasks involving high‑throughput LLM serving
  • Apply expertise with vLLM, SGLang, TensorRT-LLM, Ray Serve, or comparable platforms
  • Evaluate KV‑cache, paged‑attention, and continuous‑batching strategies
  • Analyse serving architectures for latency, throughput, memory, and scalability trade‑offs
  • Review production‑oriented approaches to large‑scale inference deployment

Technical Evaluation & Research Collaboration

  • Evaluate MLOps and ML‑systems tasks and proposed solutions
  • Provide precise written feedback that can withstand technical review
  • Develop detailed rubrics and evaluation frameworks for systems‑level work
  • Help research and engineering teams close technical knowledge gaps
  • Collaborate with subject‑matter experts to maintain consistent training‑data quality

Ideal Profile

  • 2 years of hands‑on professional experience in ML systems, ML infrastructure, model serving, or accelerator‑performance engineering
  • Strong practical experience in at least one of GPU kernel programming, performance profiling, distributed debugging, or high‑throughput inference serving
  • Production experience with JAX and/or PyTorch
  • Familiarity with CUDA, Triton, Pallas, or comparable accelerator‑programming technologies
  • Experience with profiling tools such as Kineto, torch.profiler, Nsight, XLA, or JAX profiler
  • Experience debugging distributed or accelerator‑bound workloads
  • Familiarity with vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching
  • Framework‑level experience with custom operators, FSDP, DDP, DeepSpeed, Megatron, compiler, or graph‑level work is highly valuable
  • Familiarity with accelerators such as A100, H100, B200, or TPU
  • Ability to reason precisely about throughput, latency, memory, and compute trade‑offs
  • Demonstrable professional progression in ML infrastructure or systems engineering
  • Strong written communication and ability to explain complex technical decisions clearly

Engagement Details

  • Full‑time 40‑hour‑per‑week engagement
  • Remote — Canada, United Kingdom, and United States
  • Compensation: $90–$120/hour
  • Reliable weekday availability is required
  • The engagement requires no conflicting or concurrent professional engagements
  • Work will involve ML‑systems task development, reference‑solution authoring, technical evaluation, rubric development, and research collaboration
  • Primary technical areas include GPU kernels, performance profiling, distributed debugging, and high‑throughput LLM inference
  • Assignments may involve PyTorch, JAX, CUDA, Triton, distributed‑training frameworks, modern accelerators, and production serving systems
  • Projects may be extended, shortened, or concluded depending on project needs and performance
  • H1‑B and STEM OPT candidates cannot currently be supported
  • Employment classification should be confirmed during onboarding because the source materials contain conflicting W‑2 and independent‑contractor language
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24‑MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project‑based workstreams.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago
MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago

Benture • California (MO), Northern (KY)

Hybrid
USD 124,000 - 165,000
Remote | ML Engineer — $100–$150/hour
Remote | ML Engineer — $100–$150/hour

24-Mag Llc • New York (NY), Northern (KY)

Hybrid
USD 138,000 - 207,000
ML Systems Engineer - Fully Remote | Upto $110/hr
ML Systems Engineer - Fully Remote | Upto $110/hr

Remote Jobs • United States

Remote
USD 96,000 - 152,000
MLOps Engineer - GPU Specialist
MLOps Engineer - GPU Specialist

Obsidian • San Francisco (CA)

On-site
USD 120,000 - 180,000
Remote MLOps Engineer: LLM Serving & GPU Profiling
Remote MLOps Engineer: LLM Serving & GPU Profiling

24 Mag • New York (NY)

Remote
USD 124,000 - 165,000
MLOps Engineer - GPU Specialist
MLOps Engineer - GPU Specialist

Mercor • San Francisco (CA)

On-site
USD 150,000 - 210,000
MLOps Engineer
MLOps Engineer

Bright Vision Technologies • United States

Remote
USD 100,000 - 150,000
MLOps Engineer
MLOps Engineer

Bright Vision Technologies • Plymouth (MN)

Remote
USD 100,000 - 150,000
ML Ops Engineer (EMEA Remote)
ML Ops Engineer (EMEA Remote)

Pragmatike • Town of Italy (NY)

On-site
USD 120,000 - 160,000
MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)
MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling)

HumanitApp • Northern (KY)

Hybrid
USD 257,887,000 - 343,849,000