Member of Technical Staff, Post-training

Hark

San Jose (CA)

On-site

USD 180,000 - 450,000

Full time

14 days+
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Hark in San Jose, California is seeking a Member of Technical Staff to lead the development of strategies for coding agents using reinforcement learning. The role involves designing experiments, building evaluation frameworks, and collaborating across teams to implement improvements driven by research insights.

Strong candidates should have a background in machine learning, particularly within reinforcement learning, with skills in Python and PyTorch, and an ability to work in fast-paced, research-oriented settings.

Qualifications

  • Strong background in machine learning with experience in large model training.
  • Deep understanding of reinforcement learning concepts and techniques.
  • Proficiency in Python and experience with PyTorch.
  • Ability to design rigorous experiments and diagnose issues.

Responsibilities

  • Design and implement post-training strategies for coding agents.
  • Build simulation and scaffolding environments for agentic RL.
  • Develop reward modeling pipelines for training agents.
  • Scale synthetic data generation and improve training efficiency.
  • Build evaluation frameworks to measure the progress of agents.
  • Collaborate to translate research insights into improvements.

Skills

Machine learning
Reinforcement learning
Python
PyTorch
Simulation environments

Job description

About the Role

We are looking for a Member of Technical Staff, Post-Training to lead the development of post-training strategies that define how our models acquire coding, computer use, and agentic capabilities at scale.

This role sits at the frontier of a rapidly emerging discipline — one where reinforcement learning, simulation, and large-scale model training converge to produce agents that can reason, plan, and act over long horizons. There is no established playbook here. We’re looking for researchers and engineers who can bring rigor and creativity from adjacent fields — RL, robotics, game-playing systems, compiler tooling, formal verification, or program synthesis — and apply them to the next generation of coding and agentic AI.

Responsibilities
  • Design and implement post-training strategies, primarily RL-based, to develop strong coding agents capable of multi-step reasoning, tool use, and long-horizon task completion.
  • Build and scale simulation and scaffolding environments for agentic RL: code execution sandboxes, computer use environments, tool-calling harnesses, and verifiable reward signals.
  • Develop reward modeling pipelines — including outcome-based, execution-based, and process-based reward signals — and iterate on them based on training dynamics.
  • Scale synthetic data generation and trajectory distillation pipelines that feed RL training and improve sample efficiency.
  • Design and run rigorous ablations to understand how algorithm choice, data mixture, reward shaping, and scale interact in the agentic setting.
  • Build evaluation frameworks grounded in real agent tasks — code correctness, execution success, multi-step tool use — to measure progress and guide iteration.
  • Collaborate with mid-training, infrastructure, and product teams to translate research insights into durable improvements on the model.
Requirements
  • Strong background in machine learning, with hands-on experience training or fine-tuning large models — LLMs, multimodal, or equivalent systems.
  • Deep understanding of reinforcement learning: policy optimization, reward design, exploration, and the interplay between environment design and agent behavior.
  • Experience building or working within simulation or execution environments (e.g., code interpreters, sandboxed execution, game environments, robotics simulators).
  • Proven ability to design and execute rigorous experiments, with strong intuition for diagnosing training failures and scaling bottlenecks.
  • Proficiency in Python and PyTorch; comfort working across research and systems code.
  • Ability to work in a fast-moving, research-forward environment where the right approach is often unknown at the outset.

We expect strong candidates to come from a range of backgrounds — RL research, robotics, competitive programming systems, compilers, formal methods, or large-scale ML — rather than post-training specifically. The field is new enough that directly relevant experience is rare; what matters is depth, rigor, and transferability.

Bonus Qualifications
  • Experience with RL algorithms applied to language or code: RLHF, DPO, GRPO, PPO, or similar paradigms in the LLM setting.
  • Familiarity with coding agent benchmarks and evaluation environments (e.g., SWE-bench, HumanEval, LiveCodeBench, competitive programming judges).
  • Background in reward modeling — outcome-based, process-based, or learned reward signals.
  • Experience with trajectory-based training, imitation learning, or data distillation from stronger models or human demonstrations.
  • Prior work on computer use, GUI agents, or tool-using LLMs (e.g., OSWorld, WebArena-style tasks).
  • Experience training or scaling models at 10B+ parameters, with attention to efficiency, stability, and GPU utilization.
  • Contributions to open-source ML projects or publications at top venues (NeurIPS, ICML, ICLR, EMNLP, COLM, etc.).
Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components and benefits depending on the specific role. This information will be shared if an employment offer is extended.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Post-training San Jose
Member of Technical Staff, Post-training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Machine Learning Engineer, LLM Post-Training
Machine Learning Engineer, LLM Post-Training

GoTo Meeting • Mountain View (CA)

On-site
USD 150,000 - 230,000
Health, dental, and vision care for you and your family
Top-tier 401(K) plan with company matching
Paid time off and paid holidays
+2
Member of Technical Staff, Mid-training
Member of Technical Staff, Mid-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Member of Technical Staff, Research — Early Career(PHD)
Member of Technical Staff, Research — Early Career(PHD)

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Meals and office benefits
Member of Technical Staff - Research & Post-training
Member of Technical Staff - Research & Post-training

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Research, Post-Training
Research, Post-Training

Thinking Machines Lab Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 350,000 - 475,000
Health benefits
Dental benefits
Vision benefits
+3
Member of Technical Staff, Mid-training San Jose
Member of Technical Staff, Mid-training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
RESEARCHER, POST-TRAINING
RESEARCHER, POST-TRAINING

MakerMaker.AI • San Francisco (CA)

On-site
USD 120,000 - 160,000
Reinforcement Learning Infrastructure Engineer
Reinforcement Learning Infrastructure Engineer

Elorian AI • San Francisco (CA)

On-site
USD 200,000 - 400,000
Health, dental, and vision benefits
Unlimited PTO
Paid parental leave
+1
Research Engineer, Code RL (Reinforcement Learning)
Research Engineer, Code RL (Reinforcement Learning)

Anthropic • New York (NY)

Hybrid
USD 500,000 - 850,000