ML Researcher — Agentic Reinforcement Learning

KAISHI PARTNERS PTE. LTD.

Singapore

On-site

SGD 120,000 - 180,000

Full time

7 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

KAISHI PARTNERS PTE. LTD. in Singapore is seeking researchers to advance a personal shopping AI agent. You will design RL experiments, build evaluation frameworks, and translate research into actionable product behaviours that improve recommendations and user satisfaction.

You will work closely with engineers to shape learning objectives, collect usable data, and rigorously test new methods in a production-like environment.

Qualifications

  • Deep knowledge of reinforcement learning and hands-on research or implementation experience.
  • Strong Python and ML engineering skills.
  • Experience designing experiments and evaluating results.
  • Ability to translate an open-ended product problem into a tractable research question.
  • Interest in working closely with engineers and learning from real product behaviour.
  • Evidence of research depth through papers, experiments, implementations, or deployed systems.

Responsibilities

  • Investigate reward design and credit assignment for delayed outcomes such as satisfaction, returns, and repeat use.
  • Develop approaches to user modelling, memory, and adaptation as preferences change.
  • Explore policies for when an agent should ask, recommend, act, or wait.
  • Build reproducible experiments, meaningful baselines, and ablation studies.
  • Evaluate policy improvements critically, including uncertainty and unintended behaviours.
  • Collaborate with engineers on data and instrumentation for research.
  • Translate findings into product behaviours and testing.

Skills

Reinforcement learning
Python
Experiment design
Critical evaluation
Research depth

Tools

Python

Job description

About the company

Our client is an early-stage, venture-backed AI startup building a personal shopping agent that understands users’ preferences, helps them discover relevant products, and supports them in taking action. The team brings together AI agents, personalisation, and commerce to build a service that becomes more useful through ongoing interactions and real customer outcomes.


They are hiring in Singapore, with an opportunity for early team members to shape the research agenda and learning systems behind the product.


The opportunity

You will investigate how an AI agent can make better decisions for a user over time. In commerce, an immediate action is an incomplete measure of success: a purchase may later be returned, preferences can change, and the most useful recommendation may be to buy nothing.


Working closely with engineers, you will help define learning objectives, build rigorous evaluations, and develop methods for improving agent behaviour as usable data becomes available.


What you’ll do


  • Investigate reward design and credit assignment for delayed outcomes such as satisfaction, returns, and repeat use.

  • Develop approaches to user modelling, memory, and adaptation as preferences and needs change.

  • Explore policies for deciding when an agent should ask, recommend, act, or wait.

  • Build reproducible experiments, meaningful baselines, and ablation studies.

  • Evaluate policy improvements critically, including uncertainty, misleading proxies, and unintended behaviours.

  • Collaborate with engineers on the data and instrumentation needed for research.

  • Translate promising findings into behaviours that can be tested in the product, introducing more sophisticated learning methods when the evidence supports them.


What you’ll bring


  • Deep knowledge of reinforcement learning and hands‑on research or implementation experience.

  • Strong Python and machine learning engineering skills.

  • Experience designing experiments and evaluating results critically.

  • The ability to translate an open-ended product problem into a tractable research question.

  • An interest in working closely with engineers and learning from real product behaviour.

  • Evidence of research depth through papers, experiments, implementations, or deployed systems.


Useful additional experience


  • Sequential decision-making, contextual bandits, or offline reinforcement learning.

  • Recommendation systems, personalisation, or long‑term user modelling.

  • Agent evaluation, delayed feedback, or learning from logged interactions.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Product Manager Consumer AI
Product Manager Consumer AI

kaishi partners pte. ltd. • Singapore

On-site
SGD 90,000 - 150,000
AI Research Engineer
AI Research Engineer

MOZAT PTE LTD • Singapore

On-site
SGD 80,000 - 120,000
Competitive salary
Equity
Flexible working arrangements
+1
Agentic Commerce AI Agent Algorithm Engineer
Agentic Commerce AI Agent Algorithm Engineer

Shopee • Singapore

On-site
SGD 120,000 - 180,000
Agentic RL Researcher for Personal Shopping AI
Agentic RL Researcher for Personal Shopping AI

KAISHI PARTNERS PTE. LTD. • Singapore

On-site
SGD 120,000 - 180,000
Full Stack Engineer — AI Products
Full Stack Engineer — AI Products

KAISHI PARTNERS PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
Senior AI Engineer
Senior AI Engineer

OCBC (Singapore) • Singapore

On-site
SGD 120,000 - 180,000
Competitive salary
Extensive learning opportunities
Senior AI Engineer
Senior AI Engineer

OCBC • Singapore

On-site
SGD 180,000 - 240,000
AI Researcher (Data Science & LLMs)
AI Researcher (Data Science & LLMs)

JAC RECRUITMENT PTE. LTD. • Singapore

On-site
SGD 90,000 - 130,000
AI Engineer Engineering and Technology Singapore Sea Corporate Lab
AI Engineer Engineering and Technology Singapore Sea Corporate Lab

SEA Singapore • Singapore

On-site
SGD 70,000 - 100,000
AI Researcher
AI Researcher

JAC Recruitment Pte Ltd • Singapore

On-site
SGD 60,000 - 90,000