Senior Software Test Engineer

Jobtailor

Deutschland

Remote

EUR 90.000 - 120.000

Vollzeit

Vor 3 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Mach aus dieser Rolle ein Vorstellungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor is seeking an experienced quality engineer to build and maintain the eval platform, including harnesses, datasets, graders, and CI integration. You will measure agent behavior, drift, latency, and cost while conducting exploratory end-to-end tests based on real-world workflows used by members, operations teams, and providers.

You will work with stakeholders to validate the product, develop domain expertise in Curative’s business workflows, and use traces and dashboards to translate

Qualifikationen

  • 5+ years in software quality, testing, or engineering with real range.
  • Experience writing automation and performing serious exploratory testing.
  • Fluency with LLM evaluation, including golden datasets, LLM-as-judge, programmatic graders, and regression suites for prompts and agent loops.
  • Hands-on experience with AI agents and failure modes including silent drift, tool misuse, and compounding errors.
  • Debugging depth with logs, traces, queries, and production forensics.

Aufgaben

  • Build and maintain the eval platform with harnesses, datasets, graders, and CI integration.
  • Measure agent behavior through task success, tool-use correctness, drift, latency, and cost.
  • Perform hands-on exploratory end-to-end testing based on how members, operations teams, and providers use the product.
  • Work with business stakeholders to understand real workflows and validate that the right product was built.
  • Develop domain expertise in Curative’s business workflows.
  • Use traces, failure taxonomies, and dashboards to turn agent issues into specific findings.
  • Reproduce production failures, identify root causes, fix issues, or provide precise diagnoses.
  • Advise on quality concerns and make product behavior or requirements decisions when needed.
  • Prioritize testing and quality work according to risk and business impact.
  • Review AI-generated code changes and maintain quality standards

Kenntnisse

Automation Testing
Exploratory Testing
Python
TypeScript
LLM Evaluation
Debugging

Tools

Claude Code
Cursor
LLM-as-Judge
LangSmith
Braintrust
Arize

Jobbeschreibung

• Build and maintain the eval platform, including harnesses, datasets, graders, and CI integration
• Measure agent behavior through task success, tool-use correctness, drift, latency, and cost
• Perform hands‑on exploratory end‑to‑end testing based on how members, operations teams, and providers use the product
• Work with business stakeholders to understand real workflows and validate that the right product was built
• Develop domain expertise in Curative’s business workflows
• Use traces, failure taxonomies, and dashboards to turn agent issues into specific findings
• Reproduce production failures, identify root causes, fix issues, or provide precise diagnoses
• Advise on quality concerns and make product behavior or requirements decisions when needed
• Prioritize testing and quality work according to risk and business impact
• Review AI-generated code changes and maintain quality standards

Requirements
  • 5+ years in software quality, testing, or engineering with real range
  • Experience writing automation and performing serious exploratory testing
  • Fluency with LLM evaluation, including golden datasets, LLM-as-judge, programmatic graders, and regression suites for prompts and agent loops
  • Hands‑on experience with AI agents and failure modes including silent drift, tool misuse, and compounding errors
  • Debugging depth with logs, traces, queries, and production forensics
  • Ability to fix bugs in Python or TypeScript
  • Ability to work with non‑technical business stakeholders and operations leads
  • Comfort with ambiguity, speed, and working without mature processes or sprint boundaries
  • Pragmatic quality judgment
  • Sharp written communication
  • Claude Code, Cursor, or equivalent as a primary development tool
  • Uses AI to accelerate quality work while recognizing the limits of AI judgment
  • Reviews every AI‑generated diff and does not merge on vibes
  • Strongly preferred: experience standing up eval or observability infrastructure from zero
  • Strongly preferred: LLM observability tooling such as LangSmith, Braintrust, Arize, or homegrown tools
  • Strongly preferred: defining unsupervised agent capabilities and building guardrails
  • Strongly preferred: product requirements and behavior definition experience
  • Strongly preferred: building deep domain expertise in a complex operational business
Core Competencies

Demonstrates expertise in software quality assurance, including automation, exploratory testing, and debugging in Python or TypeScript. Possesses strong communication skills to collaborate with business stakeholders and develop domain knowledge in complex operational workflows.

Highest‑signal resume keywords
  • Software Quality Assurance
  • Automation Testing
  • LLM Evaluation
  • Debugging in Python
  • AI Agent Experience
ATS Optimization Keywords
Hard Skills
  • Software Quality
  • Exploratory Testing
  • Automation Writing
  • Debugging Depth
  • Production Forensics
  • Failure Modes Analysis
  • Root Cause Identification
  • Quality Judgment
  • Agent Behavior Measurement
  • Product Requirements Definition
Soft Skills
  • Sharp Written Communication
  • Comfort with Ambiguity
  • Pragmatic Quality Judgment
Industry Keywords
  • Eval Platform
  • Observability Infrastructure
  • AI‑Generated Code Review
  • Business Workflows
  • Agent Capabilities
Tools & Technologies
  • Claude Code
  • Cursor
  • LLM-as-Judge
  • LangSmith
  • Braintrust
  • Arize
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Senior Software Engineer – Applied AI
Senior Software Engineer – Applied AI

Jobtailor • Deutschland

Remote
EUR 120.000 - 180.000
AI Engineering Strategist
AI Engineering Strategist

Jobtailor • Deutschland

Remote
EUR 70.000 - 100.000
Manager, Machine Learning Engineer
Manager, Machine Learning Engineer

Jobtailor • Deutschland

Remote
EUR 120.000 - 180.000
Junior Data Scientist
Junior Data Scientist

Jobtailor • Hamburg

Vor Ort
EUR 70.000 - 100.000
Principal Quality Engineer
Principal Quality Engineer

Jobtailor • Deutschland

Remote
EUR 110.000 - 150.000
Senior AI Engineer – Location AI
Senior AI Engineer – Location AI

Jobtailor • Deutschland

Remote
EUR 110.000 - 170.000
AI Research Engineer
AI Research Engineer

Jobtailor • Berlin

Vor Ort
EUR 90.000 - 130.000
Prompt Engineer
Prompt Engineer

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Director of Engineering – AI Apps, Agent Studio
Director of Engineering – AI Apps, Agent Studio

Jobtailor • Deutschland

Remote
EUR 180.000 - 240.000
Staff AI Test Automation Engineer
Staff AI Test Automation Engineer

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000
Health Insurance
Benefits Administration
Regulated-Industry Experience