Member of Technical Staff, Post-Training

Goaly

Menlo Park, Northern (CA, KY)

Hybrid

USD 150,000 - 230,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Meals and office benefits
Visa sponsorship

Job summary

Goaly is seeking a research-engineering role owner for the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning experiments, diagnose model behavior, and convert results into better recipes and production models.

This is a research-engineering role; you will work with RL systems, training, inference, product, and domain experts, improving infrastructure and pipelines as

Qualifications

  • Strong Python and software-engineering skills, ability to turn ambiguous ideas into reliable experimental systems.
  • Hands-on experience training, fine-tuning, or evaluating modern language models, or a closely related ML research area.
  • Solid understanding of deep learning and optimization, plus reinforcement-learning intuition.
  • Excellent experimental judgment: controls, data inspection, metrics, reproducibility.
  • Ability to debug across model behavior, data, code, and distributed infrastructure.

Responsibilities

  • Design and run post-training experiments for agentic capabilities, including tool use, coding, reasoning, planning, long-horizon task completion, and recovery from failure.
  • Prepare high-quality training and evaluation data: define task distributions, curate and filter examples, control contamination, balance difficulty, and build reproducible data-generation pipelines.
  • Build realistic RL environments and task harnesses with clear interfaces, reliable resets, isolated execution, useful telemetry, and reward signals that are hard to game.
  • Develop evaluations that measure both capability and reliability. Create regression suites, behavioral slices, error taxonomies, and dashboards that connect metrics to failures.
  • Iterate on training recipes, including supervised warm starts, sampling strategies, reward design, verifiers, curricula, optimization choices, and reinforcement fine-tuning methods.
  • Analyze trajectories and model behavior to find reward hacking, shortcut learning, and other failure modes; turn findings into targeted experiments.
  • Improve the research workflow through better experiment configuration, rollout inspection, reproducibility, checkpoint evaluation, and automated comparison of runs.
  • Partner with systems engineers to debug cross-layer problems in rollout inference, environment execution, distributed training.
  • Translate successful research ideas into stable pipelines and help set the team's longer-term post-training roadmap.

Skills

Python
Software engineering
RL basics

Tools

PyTorch
JAX
DeepSpeed
Ray

Job description

About us

We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.


About the role

You will own the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning and related post-training experiments, diagnose model behavior, and convert results into better recipes and production models.


This is a research-engineering role. The best candidates are equally comfortable forming hypotheses, writing high-quality code, operating training pipelines, and investigating why a model or metric moved. You will work closely with RL systems, training, inference, product, and domain experts; when infrastructure slows the science, you will help improve the infrastructure rather than treating it as someone else's problem.


What you'll do


  • Design and run post-training experiments for agentic capabilities, including tool use, coding, reasoning, planning, long-horizon task completion, and recovery from failure.


  • Prepare high-quality training and evaluation data: define task distributions, curate and filter examples, control contamination, balance difficulty, and build reproducible data-generation pipelines.


  • Build realistic RL environments and task harnesses with clear interfaces, reliable resets, isolated execution, useful telemetry, and reward signals that are hard to game.


  • Develop evaluations that measure both capability and reliability. Create regression suites, behavioral slices, error taxonomies, and dashboards that connect aggregate metrics to concrete model failures.


  • Iterate on training recipes, including supervised warm starts, sampling strategies, reward design, verifiers, curricula, optimization choices, and reinforcement fine-tuning methods.


  • Analyze trajectories and model behavior to find reward hacking, shortcut learning, mode collapse, distribution gaps, and other failure modes; turn those findings into targeted experiments.


  • Improve the research workflow through better experiment configuration, rollout inspection, reproducibility, checkpoint evaluation, and automated comparison of runs.


  • Partner with systems engineers to debug cross-layer problems in rollout inference, environment execution, distributed training, and data movement.


  • Translate successful research ideas into stable, repeatable pipelines and help set the team's longer-term post-training roadmap.



You may be a good fit if you have


  • Strong Python and software-engineering skills, including the ability to turn ambiguous research ideas into reliable experimental systems.


  • Hands-on experience training, fine-tuning, or evaluating modern language models, or an exceptional record in a closely related ML research area.


  • Solid understanding of deep learning and optimization, plus enough reinforcement-learning intuition to reason about policies, rewards, sampling, credit assignment, and evaluation bias.


  • Excellent experimental judgment: you define controls, inspect data, validate metrics, keep results reproducible, and distinguish a real improvement from noise or leakage.


  • Ability to debug across model behavior, data, code, and distributed infrastructure without losing sight of the user-facing capability being improved.


  • Clear written and verbal communication and a track record of productive collaboration across research and engineering.



Strong pluses


  • Experience with RLHF, reinforcement fine-tuning, preference optimization, reward or verifier modeling, or large-scale online sampling.


  • Experience building agent environments, secure sandboxes, coding benchmarks, tool-use tasks, or long-horizon evaluations.


  • Familiarity with PyTorch or JAX and distributed ML systems; experience with frameworks such as FSDP, Megatron, DeepSpeed, Ray, veRL, or related stacks.


  • A record of influential research, open-source contributions, technically ambitious independent projects, or production model launches.



How we work


  • Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.


  • High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.


  • Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.


  • Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.


  • Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.


  • Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.



Location, visa sponsorship & benefits



  • Location-based hybrid policy. This is a location-based hybrid role. We currently expect all staff to work from one of our offices at least three days per week. Exact office options will be confirmed during the recruiting process.


  • Visa sponsorship. We do sponsor visas. However, we cannot successfully sponsor a visa for every role and every candidate. If we make you an offer, we will make every reasonable effort to secure the necessary visa, and we retain immigration counsel to support the process.


  • Meals and office benefits. We provide complimentary lunch and dinner in our offices, along with snacks and beverages.



A note on qualifications. We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.


Equal opportunity

We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Member of Technical Staff, Research — Early Career(PHD)
Member of Technical Staff, Research — Early Career(PHD)

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Meals and office benefits
Member of Technical Staff, New Grad
Member of Technical Staff, New Grad

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 120,000 - 170,000
Meals and office benefits
Member of Technical Staff, AI Infrastructure
Member of Technical Staff, AI Infrastructure

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 180,000
RL Environments Engineer
RL Environments Engineer

Bespoke Labs Inc. • Mountain View (CA), Northern (KY)

Hybrid
USD 250,000 - 300,000
Health, dental, and vision coverage
401(k)
Daily onsite lunch provided
+2
RL Environments Engineer
RL Environments Engineer

Bespoke-Labs • Mountain View (CA)

On-site
USD 250,000 - 300,000
Health, dental, and vision
401(k)
Daily onsite lunch
+3
Agent Post-Training Research
Agent Post-Training Research

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer, Post-Training
Research Engineer, Post-Training

cognition • San Francisco (CA)

On-site
USD 150,000 - 210,000
Research Engineer
Research Engineer

Cerebras • San Jose (CA)

On-site
USD 180,000 - 240,000
Research Engineer, ML Infrastructure
Research Engineer, ML Infrastructure

cognition • San Francisco (CA)

On-site
USD 180,000 - 250,000
Applied Scientist
Applied Scientist

Vecna AI • Chicago (IL)

On-site
USD 140,000 - 210,000