Applied Scientist: Agentic RL & Post-Training

Vecna AI

Chicago (IL)

On-site

USD 140,000 - 210,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Vecna AI seeks a researcher to own the end-to-end post-training and alignment loop for long-horizon, tool-using agents. You will design and execute SFT/DPO/GRPO with trajectory rewards, training models that remain coherent across hundreds of tool calls in unseen environments.

You will publish and contribute to the community. You will work directly with founders, shaping how models are trained and evaluated, not a static benchmark.

Qualifications

  • PhD or equivalent research depth in ML with published work in post-training, RL, alignment, evaluation, or agentic systems.
  • Owned a post-training pipeline end to end: data curation through post training to evaluation and deployment.
  • Experience building evaluation infrastructure for LLMs with automated benchmarks and distribution-shift analysis.

Responsibilities

  • Post-training and alignment: SFT, DPO, GRPO, and trajectory-level rewards over real agent runs.
  • Agentic RL: live tool interactions and trajectory outcomes with reward shaping for coherence.
  • Procedural environment generation: domain randomization, non-stationary dynamics, adversarial perturbations.
  • Agent memory as model substrate: read/write memory, retrieval at inference, memory-aware prompts.
  • Graph-based reasoning: path traversal, link prediction, multi-step trajectory planning.
  • Evaluation infrastructure: offline benchmarks, red-teaming, regression testing.
  • Inference optimization and serving: quantization, KV-cache management, deployment for long-context workloads.

Education

PhD or equivalent ML research depth

Tools

PyTorch
Distributed training
Serving frameworks

Job description

Vecna AI seeks a researcher to own the end-to-end post-training and alignment loop for long-horizon, tool-using agents. You will design and execute SFT/DPO/GRPO with trajectory rewards, training models that remain coherent across hundreds of tool calls in unseen environments.

You will publish and contribute to the community. You will work directly with founders, shaping how models are trained and evaluated, not a static benchmark.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Research Engineer — Agentic Post-Training
Remote AI Research Engineer — Agentic Post-Training

Tether.io • Town of Italy (NY)

On-site
Confidential
AI Research Engineer: Post-Training & Agentic RL
AI Research Engineer: Post-Training & Agentic RL

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Research Engineer
Research Engineer

Oho Group • San Francisco (CA)

On-site
USD 160,000 - 230,000
Lead RL & Agentic AI Research—On-Device & Tools
Lead RL & Agentic AI Research—On-Device & Tools

Socket.dev • Cupertino (CA)

On-site
USD 220,000 - 320,000
Research Engineer, Post-Training
Research Engineer, Post-Training

cognition • San Francisco (CA)

On-site
USD 150,000 - 210,000
Agent Post-Training Researcher
Agent Post-Training Researcher

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Remote AI Research Engineer: Agentic Tooling
Remote AI Research Engineer: Agentic Tooling

Tether Operations Limited • United States

Remote
USD 140,000 - 240,000
Senior LLM Researcher - Post-Training & Alignment (Remote)
Senior LLM Researcher - Post-Training & Alignment (Remote)

techire ai • San Francisco (CA)

On-site
USD 350,000 - 500,000
Stock options
Remote work worldwide
Competitive compensation
Agent Post-Training Research
Agent Post-Training Research

OpenAI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Member of Technical Staff, Post-training
Member of Technical Staff, Post-training

Hark • San Jose (CA)

On-site
USD 180,000 - 450,000