Mountain View, USA Founding Engineer - Reinforcement Learning

S27a

San Mateo (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Deccan AI is seeking a founding engineer to design and build reinforcement-learning environments, evaluation harnesses, and feedback systems that support frontier AI labs and sophisticated customers. You will bridge research ambiguity to concrete environments, writing robust, scalable code and technical docs.

The role emphasizes moving from prototype to production, enabling repeatable RL workflows, and shaping Deccan's RL practice with high-ambition, hands-on engineering ownership.

Qualifications

  • Proven experience with reinforcement learning and evaluation workflows.
  • Strong programming and systems design skills.
  • Ability to work with researchers and customers to translate requirements into deployable systems.

Responsibilities

  • Design and build RL environments, task suites, evaluation harnesses, and feedback systems for frontier AI workflows.
  • Translate ambiguous research or customer needs into concrete technical specs and data requirements.
  • Prototype quickly, then harden reusable infrastructure for production use.

Skills

Reinforcement learning
Python
ML systems
Documentation

Education

BS in CS or related

Tools

PyTorch
JAX
TensorFlow
Git

Job description

Founding Engineer - Reinforcement Learning

Location: Bay Area / San Mateo, CA
Employment Type: Full-Time
Department: Engineering

About Us

Deccan AI is a model training and evaluation startup headquartered in the Bay Area, with a delivery office in Hyderabad. We are founded by IIT Bombay, IIM Ahmedabad, and ex-Google alumni, and we work with leading AI labs and enterprise teams on high-quality human data, evaluations, and AI-first scaled operations.

We believe frontier AI systems are only as good as the data, evaluations, and human judgment behind them. Our work sits at the intersection of customer needs, research ambition, and operational execution. That makes the work high-touch, technical, ambiguous, and execution-heavy.

We are looking for people who can operate with urgency, write clearly, build trust, and turn ambiguity into outcomes.

About the Role

We are hiring a Founding Engineer - Reinforcement Learning to serve as the SME for Deccan’s RL practice.

The market is moving beyond generic training data and simple evaluations. Frontier AI teams increasingly need high-quality environments, tasks, verifiers, feedback loops, and human-in-the-loop systems that can support RL, RLHF, agentic workflows, and model improvement. This work is technically hard, operationally messy, and only valuable when it reaches a bar that serious AI teams will actually use.

This is a hands‑on founding engineering role. You will design and build RL environments, evaluation harnesses, reward / verifier systems, task pipelines, and tooling that help Deccan serve frontier AI labs and technically sophisticated customers. You should be comfortable moving from research ambiguity to working systems, and from customer or researcher requirements to concrete environments that can be tested, improved, and scaled.

What You’ll Do
  • Design and build RL environments, task suites, evaluation harnesses, and feedback systems for frontier AI workflows.
  • Translate ambiguous research or customer needs into concrete technical specs: task definitions, environment behavior, reward signals, verifier logic, grading criteria, data requirements, and failure modes.
  • Prototype quickly, then harden the pieces that need to become reusable infrastructure.
  • Design human‑in‑the‑loop workflows for RLHF, preference data, expert grading, model behavior evaluation, and quality improvement.
  • Build tools to measure environment quality, task difficulty, model failure patterns, grader consistency, and production readiness.
  • Work across engineering, ML, research, delivery, quality, and GTM so RL opportunities are technically credible before Deccan commits to them.
  • Help define what frontier level means for Deccan’s RL environments, then raise the bar through real implementation.
  • Write clear technical docs, playbooks, and customer‑facing explanations so the team can reuse what you build.
  • Stay close to the market: agentic systems, coding environments, tool‑use tasks, verifiers, reward modeling, RLHF, and emerging post‑training workflows.
  • Create enough structure that Deccan can move from one‑off RL experiments to a repeatable practice.
What You Will Own
  • Technical architecture for Deccan’s RL environments and supporting systems.
  • Environment and task design for RL / RLHF workflows.
  • Verifier, reward, grading, and quality‑measurement logic.
  • Prototype‑to‑production path for RL tooling and environments.
  • Technical feasibility assessments for RL opportunities before Deccan makes commitments.
  • Internal engineering standards for reliability, instrumentation, reproducibility, and documentation.
  • A reusable foundation that helps Deccan compete for serious RL and post‑training work.
What We’re Looking For

We are looking for a strong engineer with deep ML judgment and a builder’s bias.

  • technical expertise to design and implement RL environments, task pipelines, evaluation systems, and feedback loops;
  • systems‑minded to build reusable tools instead of one‑off scripts;
  • comfortable with ambiguity, incomplete specs, and fast‑changing customer requirements;
  • rigorous about measurement, failure modes, quality bars, and reproducibility;
  • high‑agency to create the first version without waiting for a mature team around you;
  • clear in writing to explain technical tradeoffs to engineers, researchers, and customer‑facing teams;
  • direct to push back when an RL opportunity is underspecified, technically weak, or not worth pursuing yet.
Preferred Qualifications
  • 2+ years of experience building ML, AI, data, infra, or product systems; exceptional earlier‑career candidates with strong evidence of ability will also be considered.
  • Hands‑on experience with reinforcement learning, RLHF, preference modeling, reward modeling, agent environments, LLM post‑training, model evaluation, or related workflows.
  • Strong programming ability in Python and modern ML / data tooling; experience with PyTorch, JAX, TensorFlow, or similar frameworks is a plus.
  • Experience building evaluation harnesses, automated graders, verifiers, simulators, task‑generation systems, or data‑quality pipelines.
  • Strong understanding of modern LLM and agent workflows: tool use, coding tasks, multi‑step reasoning, rubric‑based grading, and failure analysis.
  • Ability to move between prototype code, production‑quality systems, and written technical specs.
  • Experience working with researchers, applied ML teams, or technically sophisticated customers.
  • Startup or zero‑to‑one experience where the job was to create structure, not wait for it.
  • Bay Area presence or willingness to work closely with the San Mateo team.
Why Join Deccan AI

At Deccan AI, you will work close to the frontier of AI model training, evaluation, and scaled human‑data operations. This role puts you at the center of one of the hardest and fastest‑moving parts of the market: building the environments and feedback systems that help AI systems improve.

This is a high‑agency role for someone who wants to build, not just advise. You will help define Deccan’s RL practice, create the first durable technical foundations, and work on problems where the bar is set by serious AI teams.

If you are excited by RL environments, agentic workflows, technical ambiguity, and the chance to build a practice from zero to one, we would like to meet you.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Mountain View, USA Founding Engineer - Robotics
Mountain View, USA Founding Engineer - Robotics

S27a • San Mateo (CA)

On-site
USD 150,000 - 210,000
Founding RL Engineer — Build Frontier AI Environments
Founding RL Engineer — Build Frontier AI Environments

S27a • San Mateo (CA)

On-site
USD 150,000 - 210,000
Mountain View, USA Strategic Project Lead
Mountain View, USA Strategic Project Lead

S27a • San Mateo (CA)

On-site
USD 150,000 - 210,000
Mountain View, USA Machine Learning Engineer - Text and Evals
Mountain View, USA Machine Learning Engineer - Text and Evals

S27a • San Mateo (CA)

On-site
USD 140,000 - 190,000
Senior AI Forward Deployed Engineer
Senior AI Forward Deployed Engineer

Handshake • San Francisco (CA)

Hybrid
USD 120,000 - 160,000
Equity in a fast-growing company
401(k) match
Paid parental leave
+2
Member of Technical Staff, Platform Engineering
Member of Technical Staff, Platform Engineering

David Joseph & Company • San Francisco (CA)

On-site
USD 200,000 - 250,000
Healthcare
Relocation support
401k with 4% match
+3
Founding ML Research Engineer — RL for AI Chips (Equity)
Founding ML Research Engineer — RL for AI Chips (Equity)

Architect • Palo Alto (CA)

On-site
Research Engineer (Mountain View) - 17813
Research Engineer (Mountain View) - 17813

somewhere • Mountain View (CA)

On-site
USD 120,000 - 160,000
Health coverage
Ownership upside
Collaboration with leading AI research organizations
Member of Technical Staff - Machine Learning Capabilities, New Graduates
Member of Technical Staff - Machine Learning Capabilities, New Graduates

Preference Model • San Francisco (CA)

On-site
USD 100,000 - 130,000
Competitive cash and equity compensation
Health, vision, dental benefits
401K match
+2
Member of Technical Staff - ML Research
Member of Technical Staff - ML Research

Architect • Palo Alto (CA)

On-site
USD 120,000 - 160,000
Competitive salary and equity
Fast-paced environment with autonomy
Cutting-edge challenges in chip design