Staff AI Research Engineer - Post-Training Agents

Goaly

Menlo Park, Northern (CA, KY)

Hybrid

USD 150,000 - 230,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Meals and office benefits
Visa sponsorship

Job summary

Goaly is seeking a research-engineering role owner for the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning experiments, diagnose model behavior, and convert results into better recipes and production models.

This is a research-engineering role; you will work with RL systems, training, inference, product, and domain experts, improving infrastructure and pipelines as

Qualifications

  • Strong Python and software-engineering skills, ability to turn ambiguous ideas into reliable experimental systems.
  • Hands-on experience training, fine-tuning, or evaluating modern language models, or a closely related ML research area.
  • Solid understanding of deep learning and optimization, plus reinforcement-learning intuition.
  • Excellent experimental judgment: controls, data inspection, metrics, reproducibility.
  • Ability to debug across model behavior, data, code, and distributed infrastructure.

Responsibilities

  • Design and run post-training experiments for agentic capabilities, including tool use, coding, reasoning, planning, long-horizon task completion, and recovery from failure.
  • Prepare high-quality training and evaluation data: define task distributions, curate and filter examples, control contamination, balance difficulty, and build reproducible data-generation pipelines.
  • Build realistic RL environments and task harnesses with clear interfaces, reliable resets, isolated execution, useful telemetry, and reward signals that are hard to game.
  • Develop evaluations that measure both capability and reliability. Create regression suites, behavioral slices, error taxonomies, and dashboards that connect metrics to failures.
  • Iterate on training recipes, including supervised warm starts, sampling strategies, reward design, verifiers, curricula, optimization choices, and reinforcement fine-tuning methods.
  • Analyze trajectories and model behavior to find reward hacking, shortcut learning, and other failure modes; turn findings into targeted experiments.
  • Improve the research workflow through better experiment configuration, rollout inspection, reproducibility, checkpoint evaluation, and automated comparison of runs.
  • Partner with systems engineers to debug cross-layer problems in rollout inference, environment execution, distributed training.
  • Translate successful research ideas into stable pipelines and help set the team's longer-term post-training roadmap.

Skills

Python
Software engineering
RL basics

Tools

PyTorch
JAX
DeepSpeed
Ray

Job description

Goaly is seeking a research-engineering role owner for the experimental loop that turns a capable base model into a useful agent. You will design tasks and environments, prepare training and evaluation data, run reinforcement-learning experiments, diagnose model behavior, and convert results into better recipes and production models.

This is a research-engineering role; you will work with RL systems, training, inference, product, and domain experts, improving infrastructure and pipelines as

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Research Scientist, Agentic RL & Scalable AI
Research Scientist, Agentic RL & Scalable AI

Goaly • Menlo Park (CA), Northern (KY)

Hybrid
USD 150,000 - 210,000
Meals and office benefits
AI Research Engineer: Post-Training & Agentic RL
AI Research Engineer: Post-Training & Agentic RL

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Agent Post-Training Researcher
Agent Post-Training Researcher

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Founding Backend Engineer, AI Agent Platform
Founding Backend Engineer, AI Agent Platform

RiseMe • Palo Alto (CA)

Hybrid
USD 180,000 - 240,000
Hybrid in Palo Alto
Visa sponsorship available
Meals & perks: lunch, dinner, snacks,
Pioneering AI Infrastructure Engineer
Pioneering AI Infrastructure Engineer

Goaly • Palo Alto (CA)

On-site
USD 140,000 - 210,000
Agent Post-Training Research
Agent Post-Training Research

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Research Engineer: RL & Post-Training LLM Systems
Research Engineer: RL & Post-Training LLM Systems

Preference Model • San Francisco (CA)

On-site
USD 120,000 - 150,000
Competitive cash and equity compensation (>90th percentile)
Health, vision, dental benefits
401K match
+2
Research Engineer
Research Engineer

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Research Engineer, RL Environments and Infrastructure
Research Engineer, RL Environments and Infrastructure

Hyphen Connect • San Francisco (CA)

Hybrid
USD 180,000 - 240,000
GenAI Enterprise Agent ML Research Engineer
GenAI Enterprise Agent ML Research Engineer

United States Digital Space LLC • San Francisco (CA), New York (NY)

On-site
USD 265,000 - 331,000
Health, dental and vision coverage
Equity-based compensation
Retirement benefits