Member of Technical Staff, Frontier RL Environments

Hark

San Jose (CA)

On-site

USD 180,000 - 450,000

Full time

14 days+
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Hark is seeking a Member of Technical Staff, Frontier RL Environments to build agent environments for proactive, personal intelligence. You will design RL environments and curricula, scale infrastructure, and collaborate with research, engineering, product, and safety teams to advance training runs.

This role sits at the frontier of environment design, evaluation science, and large-scale model training, requiring hands-on engineering and strong collaboration across teams to push multimodal

Qualifications

  • Strong background in machine learning with hands-on experience training or fine-tuning large models (LLMs, multimodal, or equivalent systems).
  • Hands-on experience building RL environments, simulators, or task suites (Gym-style interfaces, game engines, browser/OS automation sandboxes, robotics/hardware simulators).
  • Direct work with LLMs and post-training techniques: RL, RLHF/RLAIF, reward modeling and graders, synthetic data pipelines, or agentic/tool-using systems.
  • Track record on fuzzy problems with noisy data and ambiguous objectives requiring experimentation.
  • Strong point of view on making a personal assistant useful, trustworthy, and pleasant to rely on daily.

Responsibilities

  • Design and build RL environments and tasks spanning long-horizon planning and multimodal interaction.
  • Architect reward functions, graders, and curricula for environments.
  • Scale environment infrastructure for thousands of tasks across simulations and real sessions.
  • Shape decisions on large training runs and assess intelligence capabilities.
  • Create self-improvement loops where models generate and refine training environments.

Skills

ML fundamentals
RL environments
LLMs / multimodal
Python / tooling
Cross-team collaboration

Education

Advanced degree in ML or CS

Tools

Gym-style interfaces
Simulation tools
Robotics simulators

Job description

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.


We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.


To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.


About the Team

We are looking for a Member of Technical Staff, Frontier RL Environments to build the north-star agent environments that drive progress toward truly personal, proactive intelligence. You'll decide which skills and behaviors our agents need next, build the evaluations and RL environments that reveal whether they have them, and help set the research agenda behind our highest-stakes training runs.


This role sits at the frontier of a rapidly emerging discipline, where environment design, evaluation science, and large-scale model training converge to produce agents that reason, plan, remember, and act on a person's behalf with proactivity across apps, devices, and the physical world through our hardware. You'll partner daily with research, engineering, product, infrastructure, and safety teams to decide what belongs in the next major training run, confirm whether it actually worked, and get the resulting improvements into the hands of people who rely on Hark every day.


Responsibilities


  • Design and build RL environments and tasks, spanning long-horizon planning, tool and computer use, multimodal interaction, and our hardware, that push models toward the capabilities a proactive personal assistant actually needs.

  • Architect the reward functions, graders, and task curricula that determine what a given environment actually teaches or reveals.

  • Scale environment infrastructure so thousands of tasks and rollouts can run in parallel across simulated settings and real agentic/hardware sessions.

  • Shape decisions on our biggest training runs and get an early look at what Hark's intelligence can do next.

  • Build self-improvement loops where the model helps generate, grade, and refine its own training environments, cutting the time from idea to validated result.


Requirements


  • Strong background in machine learning, with hands-on experience training or fine-tuning large models - LLMs, multimodal, or equivalent systems.

  • Hands-on experience building RL environments, simulators, or task suites, such as Gym-style interfaces, game engines, browser or OS-level automation sandboxes, or robotics/hardware simulators, and using them to train or evaluate models.

  • Direct, hands-on work with LLMs and modern post-training techniques: RL, RLHF/RLAIF, reward modeling and graders, synthetic data pipelines, or building agentic/tool-using systems.

  • A track record on fuzzy problems, where the objective is loosely defined, the data is noisy, and getting to a good answer takes both judgment and hands-on engineering.

  • A strong point of view on what makes a personal assistant genuinely useful, trustworthy, and pleasant to rely on daily, not just a leaderboard number moving up.

  • Skill at turning a vague behavioral concern into a testable experiment: form the hypothesis, stand up the pipeline, run it, read the results, and decide the next move.

  • Ease operating across research, product, infrastructure, data, hardware, and safety teams, and translating clearly between each.


Bonus Qualifications


  • Experience with RL algorithms applied to language, code, or agentic settings: RLHF, DPO, GRPO, PPO, or similar paradigms.

  • Familiarity with agent benchmarks and evaluation environments (e.g., OSWorld 1.0/2.0, Toolathlon, GDPval etc).

  • Research- or publication-level work on reward model design, e.g., comparing outcome-based vs. process-based rewards, learned reward signals, or reward-hacking mitigations.

  • Experience with trajectory-based training, imitation learning, or data distillation from stronger models or human demonstrations.

  • Prior work on computer use, GUI agents, or multimodal tool-using systems.

  • Experience training or scaling models at 100B+ parameters, with attention to efficiency, stability, and GPU utilization.

  • Contributions to open-source ML projects or publications at top venues (NeurIPS, ICML, ICLR, EMNLP, COLM, etc.).


Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.


The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components and benefits depending on the specific role. This information will be shared if an employment offer is extended.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Mid-training San Jose
Member of Technical Staff, Mid-training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Member of Technical Staff, Mid-training
Member of Technical Staff, Mid-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Member of Technical Staff, Post-training San Jose
Member of Technical Staff, Post-training San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 180,000 - 450,000
Full-Stack Engineer
Full-Stack Engineer

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000
Backend Engineer
Backend Engineer

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000
Platform Engineer
Platform Engineer

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000
Full-Stack Engineer San Jose
Full-Stack Engineer San Jose

Hark, Inc. • San Jose (CA)

On-site
USD 170,000 - 400,000
Member of Technical Staff, Multimodal Post-train/RL
Member of Technical Staff, Multimodal Post-train/RL

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000
Data Engineering Lead
Data Engineering Lead

Hark • San Jose (CA)

On-site
USD 170,000 - 450,000
Frontend Engineer
Frontend Engineer

Hark • San Jose (CA)

On-site
USD 170,000 - 400,000