RLEE - Low-Level Engineering & Kernel Inference Optimization

Open Data Science

San Francisco (CA)

Remote

USD 123,984 - 172,200

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

A cutting-edge AI company is looking for Low-Level Engineers to design RL environments that optimize kernel development and systems programming. Candidates should have strong Python skills and a solid understanding of LLMs. This remote contractor role offers an hourly rate ranging from $90 to $125, based on expertise. Applicants will contribute to the creation of feedback loops for model training across various architectures. A minimum of 4 hours overlap with PST and proficiency in advanced English is mandatory.

Qualifications

  • Strong Python programming skills are essential.
  • Clear understanding of LLMs and their limitations is required.
  • Experience with memory hierarchies and performance implications.
  • Knowledge of threading models and concurrent programming.
  • Familiarity with compiler frameworks and GPU architectures.

Responsibilities

  • Design and build MLE/SWE environments for language models.
  • Target specific language models and ensure difficulty distribution.

Skills

Strong Python
Clear understanding of LLMs
Memory hierarchies knowledge
Threading models understanding
Cache coherence knowledge
AOT compilation expertise
Modern C++ proficiency
Assembly-level programming
Debugging GPU kernels
Custom PyTorch operators
GPU communication libraries knowledge
Mixed-precision kernels understanding

Job description

RLEE - Low-Level Engineering & Kernel Inference Optimization

RL Environments Kernel Optimization GPU/CUDA Compilers (LLVM/MLIR) PyTorch Extensions Distributed Inference (vLLM/NCCL)

Brief Description of the Role

We're hiring Low-Level Engineers to design and build RL environments that teach LLMs kernel development, hardware optimization, and systems programming. The goal is to create realistic feedback loops where models learn to write high-performance code across GPU and CPU architectures.

This is a remote contractor role with ≥4 hours overlap to PST and advanced English (C1/C2) required.

About the Company

Preference Model is building the next generation of training data to power the future of AI. Today's models are powerful but fail to reach their potential across diverse use cases because so many of the tasks that we want to use these models are out of distribution. Preference Model creates RL environments where models encounter research and engineering problems, iterate, and learn from realistic feedback loops.

Our founding team has previous experience on Anthropic's data team building data infrastructure, tokenizers, and datasets behind the Claude model. We are partnering with leading AI labs to push AI closer to achieving its transformative potential.

The company is backed by Tier 1 Silicon Valley VC.

Responsibilities
  • Design and build MLE/SWE environments and diverse tasks.
  • Target a specified language model and satisfy the required difficulty distribution.
Requirements
Minimal Qualifications
  • Strong Python (engineering-quality, not notebook-only)
  • Clear understanding of LLMs, their current limitations
  • Ability to meet throughput expectations and respond quickly to feedback
  • You may be a good fit if one of the following applies
  • Deep understanding of memory hierarchies (registers, L1/L2/shared memory, HBM, system RAM) and their performance implications
  • Threading models, synchronization primitives, and concurrent programming (warps, thread blocks, barriers, atomics)
  • Cache coherence, memory access patterns, coalescing, and bank conflicts
  • AOT compilation and optimization passes (LLVM, MLIR, TVM)
  • Compiler and kernel frameworks such as CUTLASS, BitBLAS, or JAX/Pallas
  • Modern C++, including templates, concurrency, and build systems
  • Assembly-level programming and low-level optimization across GPU and CPU architectures (e.g., x86, ARM, NVIDIA Hopper, NVIDIA Blackwell)
  • Debugging and optimizing GPU kernels using CUDA and/or HIP/ROCm
  • Developing PyTorch custom operators, backend extensions, or dispatcher integrations (e.g., ATen, TorchScript, or custom backends)
  • Customizing, extending, or optimizing vLLM, including distributed inference workflows
  • GPU communication libraries and collectives, such as NVIDIA NCCL, AMD RCCL, MPI, or UCX
  • Mixed-precision and low-precision kernels (e.g., FP16, BF16, FP8, INT8), including numerical stability and performance trade-offs
Working conditions

Hourly contractor rate: 90- 125 USD/hour (dependent on the expertise level and quality of take-home assignment).

Contacts

Log In Only registered users can open employer contacts.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff - Low Level & Kernels Capabilities
Member of Technical Staff - Low Level & Kernels Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Reinforcement Learning Environment Engineer
Reinforcement Learning Environment Engineer

Open Data Science • San Francisco (CA)

Remote
USD 100,000 - 150,000
Remote Low-Level Engineer: Kernel & Inference Optimization
Remote Low-Level Engineer: Kernel & Inference Optimization

Open Data Science • San Francisco (CA)

Remote
Member of Technical Staff - Inference
Member of Technical Staff - Inference

Prime Intellect • San Francisco (CA)

Hybrid
USD 150,000 - 300,000
Cash compensation range of $150-300k
Flexible work arrangement (remote or San Francisco office)
Full visa sponsorship and relocation support
+3
Senior Software Engineer, RL Post-Training Frameworks
Senior Software Engineer, RL Post-Training Frameworks

NVIDIA • Indiana (PA)

On-site
USD 150,000 - 190,000
Member of Technical Staff - Machine Learning Capabilities
Member of Technical Staff - Machine Learning Capabilities

Preference Model • Seattle (WA)

On-site
USD 200,000 - 350,000
Health, vision, dental
401K match
Lunch onsite
+3
Member of Technical Staff - RL Inference
Member of Technical Staff - RL Inference

Pantera Capital • Palo Alto (CA)

On-site
USD 150,000 - 230,000
Member of Technical Staff, Exceptional Generalist (Remote)
Member of Technical Staff, Exceptional Generalist (Remote)

Inferact • United States

Remote
USD 180,000 - 240,000
Competitive salary and equity
Visa sponsorship
Health coverage where applicable
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • San Francisco (CA)

On-site
USD 180,000 - 300,000
Cash and equity
Ownership & autonomy
Visa sponsorship
+5
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training
Member of Technical Staff - Machine Learning Infrastructure Engineer, Post-training

Preference Model • Seattle (WA)

On-site
USD 180,000 - 300,000
Health insurance
Vision insurance
Dental insurance
+3