Senior Applied AI Engineer, Agent Quality & Evaluations

vibehackers

Northern (KY)

Hybrid

USD 150,000 - 250,000

Full time

5 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Health insurance
Parental leave
Unlimited flexible time off
401(k) match
Learning stipend

Job summary

Flodesk is seeking a Senior Applied AI Engineer to own agent quality and evaluation for member-facing AI features. You will build prompting, agent instructions, evaluation infrastructure, monitoring, and tooling needed to measure and improve production LLM/agent behavior across the product.

Responsibilities include defining how context, member data, tools, and product state should be used; prototyping agent loops, tools, and structured interfaces; and building scalable evaluation harnesses with

Qualifications

  • Proven track record building or meaningfully improving production LLM/agentic products.
  • Strong software engineering skills in Python, TypeScript, or similar languages to build internal tools, evaluation systems, and prototypes.
  • Clear understanding of interactions among prompts, context, tools, agent loops, orchestration, and product state.

Responsibilities

  • Manage member-facing AI behavior across prompt-to-create, agentic editing, and future jobs (segmentation, scheduling, analytics, recommendations).
  • Create, test, version, and document prompts and agent instructions with a clear change history and measurable impact.
  • Define how context, member data, tools, and product state should be used; prototype agent loops, tools, and structured interfaces.
  • Own AI evaluation practice: representative datasets, behavioral scenarios, scoring rubrics, regression suites, and human review.
  • Build and operate evaluation harnesses, monitoring, and quality dashboards to enable scalable measurement.
  • Diagnose failures across prompts, context, orchestration, models, tools, data, and product code; fix issues or partner with owners.
  • Set quality baselines and release gates for prompt, model, and agent changes; convert production failures and feedback into regression cases.
  • Partner with product, design, marketing, and copy experts to encode product point-of-view while contributing product judgment.

Skills

Systems thinking
Software engineering
Prompt engineering
Evaluation & observability
Experimental design
Monitoring and diagnostics
Product sense
Prototyping
Cross-functional collaboration
Documentation and versioning
Quality assurance

Tools

Python
TypeScript
Langfuse
Braintrust
Observability platforms

Job description

Senior Applied AI Engineer, Agent Quality & Evaluations

Building AI dev tools and evaluation for LLMs and agents; focuses on prompts, agent loops, and observability.

About the Role

Flodesk is hiring a Senior Applied AI Engineer to own agent quality and evaluation for member-facing AI features. You will build the prompting, agent instructions, evaluation infrastructure, monitoring, and tooling needed to measure and improve production LLM/agent behavior across the product.

Job Description
Role

Flodesk is seeking a Senior Applied AI Engineer (Agent Quality & Evaluations) to report to the Head of Product, AI Systems and Core Experience. The role focuses on shaping AI-driven member experiences by owning prompts, agent instructions, quality definitions, evaluation criteria, and the continuous-improvement loop through technical tooling and evaluation infrastructure.

Key Responsibilities
  • Manage member-facing AI behavior across prompt-to-create, agentic editing, and future jobs (segmentation, scheduling, analytics, recommendations).
  • Create, test, version, and document prompts and agent instructions with a clear change history and measurable impact.
  • Define how context, member data, tools, and product state should be used; prototype agent loops, tools, and structured interfaces.
  • Own AI evaluation practice: representative datasets, behavioral scenarios, scoring rubrics, regression suites, and human review.
  • Build and operate evaluation harnesses, monitoring, and quality dashboards to enable scalable measurement.
  • Diagnose failures across prompts, context, orchestration, models, tools, data, and product code; fix issues or partner with owners.
  • Set quality baselines and release gates for prompt, model, and agent changes; convert production failures and feedback into regression cases.
  • Partner with product, design, marketing, and copy experts to encode product point-of-view while contributing product judgment.
Requirements
  • Proven track record of building or meaningfully improving production LLM or agentic products (not just prototypes).
  • Strong software engineering skills in Python, TypeScript, or similar languages to build internal tools, evaluation systems, and prototypes.
  • Clear understanding of interactions among prompts, context, tools, agent loops, orchestration, and product state.
  • Direct experience with LLM evaluation, observability, and experimental design; ability to combine automated checks with human judgment.
  • Strong product sense to translate ambiguous member feedback into testable quality definitions and concrete system changes.
  • Systems thinking and emphasis on reusable patterns to avoid short-term assumptions that hinder future surfaces.
  • Clear communication across technical and nontechnical disciplines; able to explain tradeoffs involving quality, latency, reliability, and cost.
Bonus / Nice to have
  • Experience building content-generation, creative, or marketing tools.
  • Experience building multi-turn agents that use tools and preserve user intent across edits.
  • Experience with Langfuse, Braintrust, or comparable evaluation and observability platforms.
  • Experience working on products for small businesses, creators, or marketers.
  • Base salary range: $150,000 - $250,000 (tiers noted: Tier 1 cities $165,000 - $250,000; other US locations $150,000 - $235,000).
  • Fully paid health insurance for individual coverage.
  • Parental leave: 16 weeks paid for non-birthing parents; 22 weeks paid maternity leave for birthing parents.
  • Unlimited flexible time off.
  • 401(k) match (US employees only).
  • $1,000 annual stipend for learning and development.
  • Remote-first company with globally distributed team and in-person hubs.

Python TypeScript LLMs Langfuse Braintrust

Skills

Systems thinking Software engineering Prompt engineering Evaluation & observability Experimental design Monitoring and diagnostics Product sense Prototyping Cross-functional collaboration Documentation and versioning Quality assurance

Experience Level

Senior

USD 150,000 - 250,000/year

Employment Type

Full Time

  • Fully paid health insurance (individual coverage)
  • 16 weeks paid parental leave (non-birthing parents)
  • 22 weeks paid maternity leave (birthing parents)
  • Unlimited flexible time off
  • 401(k) match (US employees only)
  • $1,000 annual learning and development stipend
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Sr. Product Manager, AI
Sr. Product Manager, AI

Flodesk • United States

Remote
USD 150,000 - 200,000
Fully paid health insurance for single
16 weeks paid parental leave for non‑b
Unlimited flexible time off
+2
Sr. Product Manager, AI New
Sr. Product Manager, AI New

Flodesk, Inc. • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 200,000
Health insurance
Parental leave (16-22 weeks)
Unlimited time off
+2
AI Engineer
AI Engineer

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
US$1,000 per month food and commuting allowance
Laptop of choice
ESOP available
AI Engineer
AI Engineer

Vibehackers • Chicago (IL), Northern (KY)

On-site
USD 180,000 - 220,000
Equity options
Health, dental, and vision
Parental leave
+4
Software Engineer, AI Platform
Software Engineer, AI Platform

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
Base salary range of US$180,000 to US$250,000
ESOP available
US$1,000 per month food and commuting allowance
+1
Applied AI Engineer
Applied AI Engineer

SherlockTalent • Miami (FL)

On-site
USD 120,000 - 140,000
Solid Benefits
Referral bonus of $2,500
Prompt and Evaluation Engineer
Prompt and Evaluation Engineer

Pop-Up Talent • United States

On-site
USD 140,000 - 180,000
Health insurance
401K with employer matching
Discretionary Time off
+1
Junior Full Stack Automation Engineer
Junior Full Stack Automation Engineer

GigaBrands • United States

Remote
USD 28,000 - 39,000
PTO after probation
Full-time remote
High-impact ownership
Staff Engineer, AI/LLM Platform
Staff Engineer, AI/LLM Platform

Simulations Plus • Northern (KY)

On-site
USD 100,000 - 120,000
Fully remote work
Flexible schedules
Generous vacation policy
+2
Senior / Staff Software Engineer, AI Systems
Senior / Staff Software Engineer, AI Systems

Lighthouse • New York (NY)

On-site
USD 180,000 - 240,000
Visa sponsorship
Competitive base and equity
Medical, dental, and vision benefits
+3