About the job Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour
We are sharing a full-time opportunity for experienced MLOps Engineers with hands‑on expertise in large language model infrastructure, GPU acceleration, performance profiling, distributed‑system debugging, and high‑throughput inference serving to contribute to advanced AI training and evaluation initiatives.
Selected professionals will develop challenging ML‑systems tasks, produce technically rigorous reference solutions, evaluate model‑generated outputs, and help establish evaluation standards across GPU kernels, profiling, debugging, and LLM serving. This is a hands‑on systems role intended for engineers with production infrastructure experience rather than primarily applied modelling or data‑science backgrounds.
Key Responsibilities
GPU Kernels & Accelerator Engineering
- Design technically challenging tasks involving GPU and accelerator workloads
- Develop solutions covering CUDA, Triton, Pallas, or comparable kernel technologies
- Evaluate kernel‑level optimisation approaches for correctness and efficiency
- Analyse memory, compute, and hardware‑utilisation trade‑offs
- Apply practical accelerator engineering judgement to model‑generated solutions
Performance Profiling & Trace Analysis
- Develop tasks involving performance profiling and trace interpretation
- Analyse outputs from tools such as Kineto, torch.profiler, Nsight, XLA, or JAX profilers
- Identify bottlenecks across compute, memory, communication, and scheduling
- Evaluate throughput, latency, and utilisation characteristics
- Produce clear reference analyses explaining observed performance behaviour
Distributed Systems & Workload Debugging
- Design scenarios involving distributed or accelerator‑bound ML workloads
- Diagnose failures across training and inference infrastructure
- Evaluate reasoning around FSDP, DDP, DeepSpeed, Megatron, and related systems
- Review framework‑level and distributed‑system troubleshooting approaches
- Identify technically plausible but incorrect explanations or proposed fixes
LLM Inference & Serving
- Develop and assess tasks involving high‑throughput LLM serving
- Apply expertise with vLLM, SGLang, TensorRT-LLM, Ray Serve, or comparable platforms
- Evaluate KV‑cache, paged‑attention, and continuous‑batching strategies
- Analyse serving architectures for latency, throughput, memory, and scalability trade‑offs
- Review production‑oriented approaches to large‑scale inference deployment
Technical Evaluation & Research Collaboration
- Evaluate MLOps and ML‑systems tasks and proposed solutions
- Provide precise written feedback that can withstand technical review
- Develop detailed rubrics and evaluation frameworks for systems‑level work
- Help research and engineering teams close technical knowledge gaps
- Collaborate with subject‑matter experts to maintain consistent training‑data quality
Ideal Profile
- 2 years of hands‑on professional experience in ML systems, ML infrastructure, model serving, or accelerator‑performance engineering
- Strong practical experience in at least one of GPU kernel programming, performance profiling, distributed debugging, or high‑throughput inference serving
- Production experience with JAX and/or PyTorch
- Familiarity with CUDA, Triton, Pallas, or comparable accelerator‑programming technologies
- Experience with profiling tools such as Kineto, torch.profiler, Nsight, XLA, or JAX profiler
- Experience debugging distributed or accelerator‑bound workloads
- Familiarity with vLLM, SGLang, TensorRT-LLM, Ray Serve, KV cache, paged attention, or continuous batching
- Framework‑level experience with custom operators, FSDP, DDP, DeepSpeed, Megatron, compiler, or graph‑level work is highly valuable
- Familiarity with accelerators such as A100, H100, B200, or TPU
- Ability to reason precisely about throughput, latency, memory, and compute trade‑offs
- Demonstrable professional progression in ML infrastructure or systems engineering
- Strong written communication and ability to explain complex technical decisions clearly
Engagement Details
- Full‑time 40‑hour‑per‑week engagement
- Remote — Canada, United Kingdom, and United States
- Compensation: $90–$120/hour
- Reliable weekday availability is required
- The engagement requires no conflicting or concurrent professional engagements
- Work will involve ML‑systems task development, reference‑solution authoring, technical evaluation, rubric development, and research collaboration
- Primary technical areas include GPU kernels, performance profiling, distributed debugging, and high‑throughput LLM inference
- Assignments may involve PyTorch, JAX, CUDA, Triton, distributed‑training frameworks, modern accelerators, and production serving systems
- Projects may be extended, shortened, or concluded depending on project needs and performance
- H1‑B and STEM OPT candidates cannot currently be supported
- Employment classification should be confirmed during onboarding because the source materials contain conflicting W‑2 and independent‑contractor language
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform
This opportunity is available through 24‑MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project‑based workstreams.