AI/ML Software Engineer (RL Environments) (Contract)

Careerflow.ai

United States

Remote

USD 83,000 - 131,000

Part time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Careerflow.ai is seeking experienced ML/Software Engineers to design and build RL training environments for LLM agents, delivering diverse tasks with minimal supervision.

This remote contractor role requires 40 hours per week and overlap with PST hours. Strong Python skills, deep ML background, and hands-on experience with large language models are essential for delivering high-quality outputs and reliable task generation.

Qualifications

  • Strong ML background demonstrated by coursework, prior work, or personal projects.
  • Proficient Python coding with clean, efficient style.
  • Extensive daily use of LLMs and understanding model capabilities and failure modes.
  • Self-directed with a track record of generating novel ML task ideas.
  • High responsibility and integrity with consistent delivery and deadlines.

Responsibilities

  • Design and build tasks for ML domains targeting language models and specific difficulty distributions.
  • Iterate rapidly on task designs based on customer feedback with 24-hour turnaround times.
  • Create diverse, challenging scenarios that test language model capabilities and expose limitations.
  • Hit the ground running with minimal onboarding and rapid contribution.

Skills

Strong ML background
Python fluency
Heavy LLM user
Self-directed and creative
High responsibility and integrity
PST overlap availability

Job description

About the Role

We're seeking experienced Machine Learning Engineers and Software Engineers with ML experience to design and build high-quality RL training environments for LLM agents. As an RL Environment Engineer, you'll create diverse machine learning tasks that challenge and improve language models, working with minimal supervision to deliver consistent, quality outputs.

What You'll Do
  • Design and build tasks for machine learning domains that target specific language models and difficulty distributions

  • Iterate rapidly on task designs based on customer feedback, with 24-hour turnaround times

  • Create diverse, challenging scenarios that test language model capabilities and expose their limitations

  • Hit the ground running with minimal onboarding time

What We're Looking For
  • Strong machine learning background through coursework, previous work experience, or personal projects

  • Python fluency: you write clean, efficient Python code regularly

  • Heavy LLM user who understands current model capabilities and failure modes through daily hands-on experience

  • Self-directed and creative. You can generate novel ML task ideas in your domain without constant guidance

  • High responsibility and integrity. You deliver quality work consistently and meet deadlines

  • Availability overlap with PST 9am-5pm (minimum 3 hours required)

Work Details
  • Location: Remote

  • Type: Contractor

Time Commitment: 40 hours a week. Must have at least 3 hours of overlap with PST business hours (9am-5pm)

Selection Process:
  1. Screening

  2. Hacker rank assessment

  3. 1 Week paid task

  4. Full time

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Reinforcement Learning Environment Engineer
Reinforcement Learning Environment Engineer

Open Data Science • San Francisco (CA)

On-site
USD 100,000 - 150,000
Remote RL Environment Engineer for LLM Tasks
Remote RL Environment Engineer for LLM Tasks

Careerflow.ai • United States

Remote
USD 83,000 - 131,000
Machine Learning Engineers (Environment Design)
Machine Learning Engineers (Environment Design)

AIChamp Inc. • United States

Remote
USD 69,000 - 227,000
Senior AI Engineer (LLM Training & RLHF) - Remote
Senior AI Engineer (LLM Training & RLHF) - Remote

Prolific • Mesa (AZ)

On-site
USD 140,000 - 230,000
RLEE - Low-Level Engineering & Kernel Inference Optimization
RLEE - Low-Level Engineering & Kernel Inference Optimization

Open Data Science • San Francisco (CA)

On-site
USD 123,984 - 172,200
AI Engineer/ML Engineer - Senior Developers - AI Training - Mesa, US
AI Engineer/ML Engineer - Senior Developers - AI Training - Mesa, US

Prolific • Mesa (AZ)

On-site
USD 140,000 - 230,000
Work from home
Flexible hours
Competitive pay
RL Environment Software Engineer
RL Environment Software Engineer

talentpluto • San Francisco (CA)

On-site
USD 180,000 - 220,000
AI Engineer/ML Engineer - Senior Developers - AI Training - Louisville, US
AI Engineer/ML Engineer - Senior Developers - AI Training - Louisville, US

Prolific • Louisville (KY)

On-site
USD 100,000 - 150,000
Competitive pay rates
Flexible hours
Ability to work from home
SWE (RL Environments) "Reinforcement Learning"
SWE (RL Environments) "Reinforcement Learning"

AI Talent Now • San Francisco (CA)

On-site
USD 150,000 - 250,000
Remote ML Environment Design Engineer (Python Fluent)
Remote ML Environment Design Engineer (Python Fluent)

AIChamp Inc. • United States

Remote
USD 69,000 - 227,000