Data Scientist – Quality Assurance, Gen-AI-Agents, Python

Jobtailor

Kraków

On-site

PLN 150,000 - 220,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Jobtailor in Kraków, Poland, seeks an engineer to develop automated evaluation pipelines for agentic AI. You will craft datasets, benchmarks, and synthetic tests, while defining measurable client acceptance criteria and reporting results in German.

The role requires strong Python, stats background, and hands-on CI/CD experience with modern tools. German proficiency is essential for client communications.

Qualifications

  • Excellent Python skills; production-ready code using pytest, asyncio, pandas, and pydantic.
  • Hands-on experience with agentic AI tools and LLM SDKs.
  • Strong statistical background including hypothesis testing and sampling design.
  • CI/CD experience with GitHub Actions or GitLab CI; basic IT security understanding.

Responsibilities

  • Develop automated evaluation pipelines and test harnesses in Python.
  • Build golden datasets and task-based benchmarks.
  • Generate synthetic test cases and define acceptance criteria with clients.
  • Evaluate results and trajectories of agents’ problem-solving processes.
  • Assess tool selection, parameters, and solution paths for efficiency.
  • Design LLM-as-a-Judge approaches and calibrate against human evals.
  • Ensure statistical reliability of evaluation results.
  • Operationalize quality metrics (Task Success Rate, Tool-Call Accuracy, Groundedness).
  • Set up observability with Langfuse/LangSmith/OpenTelemetry.
  • Integrate eval suites into CI/CD; conduct regression and red-team tests.

Skills

Python programming
Pytest
Asyncio
Pydantic
Statistical analysis
Hypothesis testing
Sampling design
Significance testing
CI/CD
Agent frameworks

Tools

Langfuse
LangSmith
OpenTelemetry
GitHub Actions
GitLab CI
KNIME
Power BI
Claude Code
Anthropic SDK
OpenAI SDK

Job description

  • • Develop automated evaluation pipelines and test harnesses in Python using pytest, asyncio, and pydantic
  • • Build golden datasets and task-based benchmarks
  • • Generate synthetic test cases
  • • Work with clients to define measurable acceptance criteria
  • • Evaluate final results and agents’ full problem-solving processes through trajectory testing
  • • Assess correct tool selection, valid parameters, and efficient solution paths
  • • Design and calibrate LLM-as-a-Judge approaches against human evaluations
  • • Ensure statistical reliability of evaluation results
  • • Operationalize quality metrics including Task Success Rate, Tool-Call Accuracy, Groundedness, and Robustness
  • • Calculate confidence intervals and establish sound sampling designs
  • • Set up observability and tracing solutions using tools such as Langfuse, LangSmith, or OpenTelemetry
  • • Develop error taxonomies and monitor quality in production environments
  • • Integrate evaluation suites as quality gates into CI/CD pipelines
  • • Conduct regression testing when prompts, tools, or models change
  • • Conduct red-team testing, including prompt injection and guardrail testing
  • • Document testing procedures with regulatory requirements such as the EU AI Act in mind
  • • Prepare and communicate findings in German to clients’ specialist departments and management teams
  • • Write reports and present results clearly, precisely, and in a decision-supporting manner
Requirements
  • Excellent Python skills; production-ready code using pytest, asyncio, pandas, and pydantic
  • Native-level German proficiency (C2)
  • Strong proficiency with agentic AI development tools such as Claude Code
  • Hands-on experience using agentic AI development tools productively in day-to-day development
  • Understanding of the strengths and limitations of agentic AI development tools
  • Experience with LLM SDKs such as Anthropic and OpenAI
  • Experience with agent frameworks such as LangGraph
  • Solid foundation in statistics, including hypothesis testing, sampling design, and significance testing
  • Hands-on CI/CD experience with GitHub Actions or GitLab CI
  • Basic understanding of IT security
  • Experience with KNIME or Power BI
Core Competencies

Demonstrates expertise in developing automated evaluation pipelines and test harnesses using Python, with a strong focus on statistical reliability and quality metrics. Proficient in communicating findings in German and integrating evaluation suites into CI/CD pipelines.

Highest-signal resume keywords
  • Python Programming
  • Agentic AI Development Tools
  • Statistical Analysis
  • CI/CD Experience
  • Native-Level German Proficiency
Hard Skills
  • Python
  • Pytest
  • Asyncio
  • Pydantic
  • Statistical Analysis
  • Hypothesis Testing
  • Sampling Design
  • Significance Testing
  • CI/CD
  • Agent Frameworks
Soft Skills
  • Communication
  • Documentation
  • Problem-Solving
Industry Keywords
  • LLM-as-a-Judge
  • Task Success Rate
  • Tool-Call Accuracy
  • Groundedness
  • Robustness
  • Red-Team Testing
  • Prompt Injection
  • EU AI Act
Tools & Technologies
  • Langfuse
  • LangSmith
  • OpenTelemetry
  • GitHub Actions
  • GitLab CI
  • KNIME
  • Power BI
  • Claude Code
  • Anthropic SDK
  • OpenAI SDK
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist – Quality Assurance for Gen-AI-Agents, Python
Data Scientist – Quality Assurance for Gen-AI-Agents, Python

Jobtailor • Kraków

On-site
PLN 120,000 - 200,000
Senior AI Engineer
Senior AI Engineer

Bayer CropScience Limited • Warszawa

On-site
PLN 100,000 - 140,000
Python AI Developer
Python AI Developer

Britenet • Warszawa

On-site
PLN 240,000 - 360,000
Senior QA / ML Tester
Senior QA / ML Tester

EPAM Systems • Poland

Hybrid
PLN 180,000 - 260,000
Hybrid work design
Work remotely within Poland
Opportunity to work abroad up to 60/90
Principal Engineer – GenAI
Principal Engineer – GenAI

Jobtailor • Kraków

On-site
PLN 320,000 - 520,000
AI Developer Experience Engineer
AI Developer Experience Engineer

Jobtailor • Warszawa

On-site
PLN 120,000 - 170,000
QA Lead
QA Lead

Jobtailor • Gliwice

On-site
PLN 180,000 - 275,000
Senior AI Engineer
Senior AI Engineer

Intellias • Poland

On-site
PLN 240,000 - 360,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Poland

On-site
PLN 180,000 - 240,000
Technical Lead
Technical Lead

SmartDev • Województwo kujawsko-pomorskie

Hybrid
PLN 420,000 - 640,000