RLEE - Low-Level Engineering & Kernel Inference Optimization

Open Data Science

San Francisco (CA)

Remote

USD 123,984 - 172,200

Full time

14 days+
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Job summary

A cutting-edge AI company is looking for Low-Level Engineers to design RL environments that optimize kernel development and systems programming. Candidates should have strong Python skills and a solid understanding of LLMs. This remote contractor role offers an hourly rate ranging from $90 to $125, based on expertise. Applicants will contribute to the creation of feedback loops for model training across various architectures. A minimum of 4 hours overlap with PST and proficiency in advanced English is mandatory.

Qualifications

  • Strong Python programming skills are essential.
  • Clear understanding of LLMs and their limitations is required.
  • Experience with memory hierarchies and performance implications.
  • Knowledge of threading models and concurrent programming.
  • Familiarity with compiler frameworks and GPU architectures.

Responsibilities

  • Design and build MLE/SWE environments for language models.
  • Target specific language models and ensure difficulty distribution.

Skills

Strong Python
Clear understanding of LLMs
Memory hierarchies knowledge
Threading models understanding
Cache coherence knowledge
AOT compilation expertise
Modern C++ proficiency
Assembly-level programming
Debugging GPU kernels
Custom PyTorch operators
GPU communication libraries knowledge
Mixed-precision kernels understanding

Job description

RLEE - Low-Level Engineering & Kernel Inference Optimization

RL Environments Kernel Optimization GPU/CUDA Compilers (LLVM/MLIR) PyTorch Extensions Distributed Inference (vLLM/NCCL)

Brief Description of the Role

We're hiring Low-Level Engineers to design and build RL environments that teach LLMs kernel development, hardware optimization, and systems programming. The goal is to create realistic feedback loops where models learn to write high-performance code across GPU and CPU architectures.

This is a remote contractor role with ≥4 hours overlap to PST and advanced English (C1/C2) required.

About the Company

Preference Model is building the next generation of training data to power the future of AI. Today's models are powerful but fail to reach their potential across diverse use cases because so many of the tasks that we want to use these models are out of distribution. Preference Model creates RL environments where models encounter research and engineering problems, iterate, and learn from realistic feedback loops.

Our founding team has previous experience on Anthropic's data team building data infrastructure, tokenizers, and datasets behind the Claude model. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.

The company is backed by Tier 1 Silicon Valley VC.

Responsibilities
  • Design and build MLE/SWE environments and diverse tasks.
  • Target a specified language model and satisfy the required difficulty distribution.
Requirements
Minimal Qualifications
  • Strong Python (engineering-quality, not notebook-only)
  • Clear understanding of LLMs, their current limitations
  • Ability to meet throughput expectations and respond quickly to feedback
  • You may be a good fit if one of the following applies
  • Deep understanding of memory hierarchies (registers, L1/L2/shared memory, HBM, system RAM) and their performance implications
  • Threading models, synchronization primitives, and concurrent programming (warps, thread blocks, barriers, atomics)
  • Cache coherence, memory access patterns, coalescing, and bank conflicts
  • AOT compilation and optimization passes (LLVM, MLIR, TVM)
  • Compiler and kernel frameworks such as CUTLASS, BitBLAS, or JAX/Pallas
  • Modern C++, including templates, concurrency, and build systems
  • Assembly-level programming and low-level optimization across GPU and CPU architectures (e.g., x86, ARM, NVIDIA Hopper, NVIDIA Blackwell)
  • Debugging and optimizing GPU kernels using CUDA and/or HIP/ROCm
  • Developing PyTorch custom operators, backend extensions, or dispatcher integrations (e.g., ATen, TorchScript, or custom backends)
  • Customizing, extending, or optimizing vLLM, including distributed inference workflows
  • GPU communication libraries and collectives, such as NVIDIA NCCL, AMD RCCL, MPI, or UCX
  • Mixed-precision and low-precision kernels (e.g., FP16, BF16, FP8, INT8), including numerical stability and performance trade-offs
Working conditions

Hourly contractor rate: 90- 125 USD/hour (dependent on the expertise level and quality of take-home assignment).

Contacts

Log In Only registered users can open employer contacts.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Reinforcement Learning Environment Engineer
Reinforcement Learning Environment Engineer

Open Data Science • San Francisco (CA)

On-site
USD 100,000 - 150,000
Remote Low-Level Engineer: Kernel & Inference Optimization
Remote Low-Level Engineer: Kernel & Inference Optimization

Open Data Science • San Francisco (CA)

Remote
USD 123,984 - 172,200
AI/ML Software Engineer (RL Environments) (Contract)
AI/ML Software Engineer (RL Environments) (Contract)

Careerflow.ai • United States

Remote
USD 83,000 - 131,000
Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour
Remote | MLOps Engineer, LLM Systems (Serving, GPU Kernels, Profiling) — $90–$120/hour

24 Mag • New York (NY)

Remote
USD 124,000 - 165,000
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA), Northern (KY)

On-site
USD 150,000 - 300,000
Remote option
Visa sponsorship
Relocation support
+2
Software Engineer, Inference Runtime
Software Engineer, Inference Runtime

EngRadar • New York (NY)

On-site
USD 150,000 - 230,000
Equity grants
Medical plan
Vision plan
+5
Member of Technical Staff, AI-Driven Compilation
Member of Technical Staff, AI-Driven Compilation

San Francisco Tensor Company • San Francisco (CA)

On-site
USD 275,000 - 315,000
Relocation assistance
Meaningful equity
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

On-site
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
Remote RL Environment Engineer for LLM Tasks
Remote RL Environment Engineer for LLM Tasks

Careerflow.ai • United States

Remote
USD 83,000 - 131,000
MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago
MLOps Engineer, LLM Systems · Mercor Mercor · 90-120/hr · remote in US, UK, CA · 2w ago 90-120/hr remote in US, UK, CA 2w ago

Benture • California (MO), Northern (KY)

Hybrid
USD 124,000 - 165,000