AI Training and Prompt Engineering Specialist

Yonder Media Mobile Inc.

Poland

Remote

USD 120,000 - 180,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Flexible working hours
Corporate equipment for work
Competitive salary
Real opportunity for personal and.prof

Job summary

House of YO is seeking a prompt and evaluation specialist for YOlanda, the AI concierge. You will own the system prompts, ensure smooth handoffs to the compute cooperative, and build evaluation sets with measurable scores.

You will also develop data and training pipelines, test robustness against adversarial prompts, and maintain versioned prompts in git. The role is remote, global, full-time, requiring strong English writing and bilingual capabilities, plus Python and LangChain/ LangGraph

Qualifications

  • 2+ years building with large language models in production.
  • Proficient Python for shipping scripts, API clients, data handling, tests.
  • Experience with LangChain and LangGraph; ability to read and modify a graph you didn’t write.
  • Experience with retrieval in production: chunking, embeddings, reranking, grounding.
  • Exceptional written English and bilingual Spanish and English.
  • Methodology: iteratively test one variable, record results, and move forward.

Responsibilities

  • Own YOlanda's system prompt — persona, tone, safety rules, refusals — and prompts behind every user action.
  • Decide what stays in YOlanda's frame of reference and what is handed off to YOnC, ensuring a seamless handoff.
  • Build a few-shot library and manage context so the model doesn’t start from a blank prompt or repeat answered questions.
  • Create an evaluation set with tasks, pass criteria and release-traceable scores.
  • Rank and critique outputs (RLHF), rewrite weak answers into fine-tuning data, and fact-check plan prices, balances and YOYO$ maths.
  • Run tests against adversarial prompts, jailbreak attempts, and abuse of top-up/rewards flows.
  • Write Python to run and score evaluations at volume, work with the API (streaming, function calling, structured output), and version prompts in git.

Skills

LLMs in production
Python
LangChain
LangGraph
Retrieval in production
English writing
English/Spanish bilingual
Git versioning

Tools

Git

Job description

ABOUT THE ROLE

House of YO is a connectivity and distribution business with compute layered on top. YOlanda is our AI concierge.

On V3 of the YO platform, YOlanda is the interface. Users talk to her instead of tapping through menus. She answers questions about plans, data, top-ups and YOYO$ balances, and she takes the action on the user’s account. When she is wrong, the user is stuck and the product has failed.

Her knowledge is the YO ecosystem and nothing else. When a question falls outside it she hands off to YOnC, our compute cooperative, which routes it to a specialist model. Knowing where that line sits, and holding it, is central to the job.

She works to one golden rule: within five responses she has put a product or service in front of the user. Not five responses to be helpful. Five to arrive somewhere real.

This role makes her right. You write the prompts that govern how she thinks. You build the evaluations that prove she works. You write the training data that fixes her when she does not. You ship the code that runs all of it.

YOlanda runs on our own LangChain and LangGraph build. You will open pull requests against that graph alongside the two engineers who own it.

You will not manage anyone. You will be in the model every day.

Who this is for

You break things on purpose and then explain exactly why they broke. You can read a bad answer and say in one sentence what is wrong with it. You would rather run fifty tests than win one argument about which prompt is better. You have opinions about tone and you can defend them. You write well enough that people notice.

You are relentless. The first version is never the one you ship.

Location:

Remote, global

Work format:

Remote, full time

WHAT YOU'LL DO
  • Own YOlanda's system prompt — persona, tone, safety rules, what she refuses — and the prompts behind every user action: plan changes, top-ups, balance queries, offers, support.
  • Own the boundary: decide what's inside YOlanda's frame of reference and what hands off to YOnC, and make the handoff seamless — same conversation, same voice.
  • Build the few-shot library and manage context, so she never opens on a blank prompt or repeats a question she already answered.
  • Build the evaluation set: every task gets test cases, a pass condition and a score you can track release over release. Nothing ships if the score went down.
  • Rank and critique outputs (RLHF), rewrite weak answers into reference/fine-tuning data, and fact-check plan prices, balances and YOYO$ maths — the hallucination target on money is zero.
  • Red team her: adversarial prompts, prompt injection, jailbreaks and abuse of the top-up and rewards flows.
  • Write Python to run and score evaluations at volume, work in the API (streaming, function calling, structured output), and keep prompts versioned in git.
WHAT YOU NEED
  • 2+ years building with large language models in production. Using a chat window every day is not this.
  • Python you can ship: scripts, API clients, data handling, tests.
  • LangChain and LangGraph — you can read a graph you didn't write and change it without breaking it.
  • Retrieval in production: chunking, embeddings, reranking, grounding, and knowing when retrieval is the wrong tool.
  • Exceptional written English, and bilingual Spanish and English to the same standard — YOlanda has one voice that has to survive the crossing.
  • Method: change one variable, test it, record it, move on.

Nice to have: fine-tuning (SFT, DPO, LoRA); evaluation frameworks (promptfoo, Braintrust, LangSmith, DeepEval); mobile, telecoms or fintech; editorial or linguistics training.

WHAT WE OFFER
  • Top-notch products disrupting the world of the traditional media-services industry
  • A brilliant, highly collaborative team with a shared drive to achieve goals
  • Flexible working hours
  • Corporate equipment for work
  • Competitive salary
  • Real opportunity for personal and professional growth
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Remote AI Prompt Architect & Evaluation Engineer
Remote AI Prompt Architect & Evaluation Engineer

Yonder Media Mobile Inc. • Poland

Remote
USD 120,000 - 180,000
Flexible working hours
Corporate equipment for work
Competitive salary
+1
Software Engineer with AI
Software Engineer with AI

Codilime • Warszawa

On-site
PLN 240,000 - 360,000
AI Prompt Engineer
AI Prompt Engineer

SoftSnow AI • Poland

On-site
PLN 180,000 - 240,000
Comprehensive Training
Growth Opportunities
Collaborative Environment
+3
Staff AI Engineer, Payments Intelligence
Staff AI Engineer, Payments Intelligence

Yuno • Kraków

On-site
PLN 320,000 - 460,000
Competitive pay
Remote work
Home office bonus
+5
Data Engineer
Data Engineer

Viktor • Warszawa

On-site
PLN 240,000 - 360,000
Senior RAG Engineer
Senior RAG Engineer

Newcode.ai • Warszawa

On-site
PLN 387,930 - 517,240
AI Engineer
AI Engineer

Delvedeeper • Warszawa

On-site
PLN 250,000 - 400,000
Hybrid working model
Competitive salary
Medicover private medical care
+5
Backend Engineer
Backend Engineer

Viktor • Warszawa

On-site
PLN 230,000 - 420,000
Equity
Lead FDE Production & Consulting
Lead FDE Production & Consulting

Newpage Solutions • Poland

On-site
PLN 180,000 - 300,000
AI Enablement Analyst (Prompt engineering, AI workflows)
AI Enablement Analyst (Prompt engineering, AI workflows)

Oliver Wyman • Warszawa

On-site
PLN 60,000 - 90,000
Private health care
Sport card
Lunch card
+6