Data Scientist – Quality Assurance for Gen-AI-Agents, Python

Jobtailor

Kraków

On-site

PLN 120,000 - 200,000

Full time

4 days ago
Be an early applicant
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Job summary

Jobtailor in Kraków, Poland, seeks a GenAI evaluation engineer to develop automated evaluation pipelines and Python test harnesses for reproducibly assessing agent behavior, build golden datasets, and generate synthetic test cases.

You will work with clients to define measurable acceptance criteria, evaluate problem-solving trajectories, and communicate findings in German for specialists and management, while integrating quality gates into CI/CD workflows.

Qualifications

  • Excellent Python skills and production-ready code experience with pytest, asyncio, pandas, and pydantic.
  • Native-level German proficiency (C2).
  • Strong proficiency with agentic AI development tools such as Claude Code.
  • Hands-on experience with LLM SDKs such as Anthropic and OpenAI.
  • Experience with agent frameworks such as LangGraph.
  • Solid foundation in statistics including hypothesis testing, sampling design, and significance testing.
  • Hands-on CI/CD experience with GitHub Actions or GitLab CI.
  • Basic understanding of IT security.
  • Experience with KNIME or Power BI.

Responsibilities

  • Develop automated evaluation pipelines and Python test harnesses for reproducibly assessing GenAI agent behavior.
  • Build golden datasets and task-based benchmarks.
  • Generate synthetic test cases.
  • Work with clients to define measurable acceptance criteria.
  • Evaluate agent problem-solving trajectories, including tool selection, parameter validity, and solution-path efficiency.
  • Design and calibrate LLM-as-a-Judge approaches against human evaluations.
  • Ensure statistical reliability of evaluation results.
  • Operationalize quality metrics including Task Success Rate, Tool-Call Accuracy, Groundedness, and Robustness.
  • Include confidence intervals and sound sampling designs in quality measurement.
  • Set up observability and tracing solutions using tools such as Langfuse, LangSmith, or OpenTelemetry.
  • Develop error taxonomies and monitor quality in production environments.
  • Integrate evaluation suites as quality gates into CI/CD pipelines.
  • Conduct regression testing when prompts, tools, or models change.
  • Conduct red-team testing, including prompt injection and guardrail testing.
  • Document testing procedures with EU AI Act regulatory requirements in mind.
  • Prepare and communicate findings in German for clients’ specialist departments and management teams.

Skills

Python programming
German proficiency
Agentic AI tools
LLM SDKs
LangGraph
Statistics
CI/CD
IT security basics
Power BI

Tools

Claude Code
Anthropic SDK
OpenAI SDK
Langfuse
OpenTelemetry
GitHub Actions
GitLab CI
KNIME
Power BI

Job description

  • Develop automated evaluation pipelines and Python test harnesses for reproducibly assessing GenAI agent behavior
  • Build golden datasets and task-based benchmarks
  • Generate synthetic test cases
  • Work with clients to define measurable acceptance criteria
  • Evaluate agent problem-solving trajectories, including tool selection, parameter validity, and solution-path efficiency
  • Design and calibrate LLM-as-a-Judge approaches against human evaluations
  • Ensure statistical reliability of evaluation results
  • Operationalize quality metrics including Task Success Rate, Tool-Call Accuracy, Groundedness, and Robustness
  • Include confidence intervals and sound sampling designs in quality measurement
  • Set up observability and tracing solutions using tools such as Langfuse, LangSmith, or OpenTelemetry
  • Develop error taxonomies and monitor quality in production environments
  • Integrate evaluation suites as quality gates into CI/CD pipelines
  • Conduct regression testing when prompts, tools, or models change
  • Conduct red-team testing, including prompt injection and guardrail testing
  • Document testing procedures with EU AI Act regulatory requirements in mind
  • Prepare and communicate findings in German for clients’ specialist departments and management teams
Requirements
  • Excellent Python skills; production-ready code using pytest, asyncio, pandas, and pydantic
  • Native-level German proficiency (C2)
  • Strong proficiency with agentic AI development tools such as Claude Code
  • Hands-on experience with LLM SDKs such as Anthropic and OpenAI
  • Experience with agent frameworks such as LangGraph
  • Solid foundation in statistics, including hypothesis testing, sampling design, and significance testing
  • Hands-on CI/CD experience with GitHub Actions or GitLab CI
  • Basic understanding of IT security
  • Experience with KNIME or Power BI
Core Competencies

Demonstrates expertise in developing automated evaluation pipelines and Python test harnesses for GenAI agent behavior assessment, with a strong focus on statistical reliability and quality metrics. Proficient in communicating findings in German and integrating evaluation suites into CI/CD pipelines.

Highest-signal resume keywords
  • Python Programming
  • Agentic AI Development Tools
  • LLM SDK Experience
  • Statistical Analysis
  • CI/CD Experience
Hard Skills
  • Python
  • Pytest
  • Asyncio
  • Pandas
  • Pydantic
  • Statistical Analysis
  • Hypothesis Testing
  • Sampling Design
  • Significance Testing
  • German Proficiency (C2)
Industry Keywords
  • GenAI
  • Evaluation Pipelines
  • Quality Metrics
  • Task Success Rate
  • Tool-Call Accuracy
  • Groundedness
  • Robustness
  • Red-Team Testing
  • EU AI Act
  • Error Taxonomies
Tools & Technologies
  • Claude Code
  • Anthropic SDK
  • OpenAI SDK
  • LangGraph
  • GitHub Actions
  • GitLab CI
  • KNIME
  • Power BI
  • Langfuse
  • OpenTelemetry
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Data Scientist – Quality Assurance, Gen-AI-Agents, Python
Data Scientist – Quality Assurance, Gen-AI-Agents, Python

Jobtailor • Kraków

On-site
PLN 150,000 - 220,000
Python AI Developer
Python AI Developer

Britenet • Warszawa

On-site
PLN 240,000 - 360,000
Senior AI Engineer
Senior AI Engineer

Bayer CropScience Limited • Warszawa

On-site
PLN 100,000 - 140,000
Senior ML / Evaluation Engineer
Senior ML / Evaluation Engineer

Intellias • Poland

On-site
PLN 180,000 - 240,000
Senior QA / ML Tester
Senior QA / ML Tester

EPAM Systems • Poland

Hybrid
PLN 180,000 - 260,000
Hybrid work design
Work remotely within Poland
Opportunity to work abroad up to 60/90
Gen-AI QA Data Scientist | Python, LLM Evaluation, German
Gen-AI QA Data Scientist | Python, LLM Evaluation, German

Jobtailor • Kraków

On-site
PLN 150,000 - 220,000
Principal Engineer – GenAI
Principal Engineer – GenAI

Jobtailor • Kraków

On-site
PLN 320,000 - 520,000
Senior AI Platform Engineer
Senior AI Platform Engineer

EPAM Systems • Poland

Hybrid
PLN 180,000 - 260,000
Hybrid by design
Remote work within Poland
Relocation opportunities
+7
Fullstack Senior Data Scientist
Fullstack Senior Data Scientist

Lingarogroup • Poland

On-site
PLN 180,000 - 280,000
Technical Lead
Technical Lead

SmartDev • Województwo kujawsko-pomorskie

Hybrid
PLN 420,000 - 640,000