Lead Engineer, AI Quality

Circle

Northern (KY)

Hybrid

USD 150,000 - 190,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Equity
Remote-friendly
35 days PTO

Job summary

Circle is seeking a Lead Engineer for AI Quality to own and evolve the evaluation frameworks, observability tooling, and diagnostic infrastructure for Circle's AI-powered features. You will lead a growing team while remaining hands-on in code, ensuring reliable, efficient AI systems at scale.

The role emphasizes building evaluation pipelines, datasets, and CI/CD workflows, with collaboration across AI core engineering. Strong Ruby on Rails or Python skills and experience in ML evals are required.

Qualifications

  • 7+ years of experience building and shipping production software, ideally including LLM-powered agents that take real actions in a product.
  • Comfortable in Ruby on Rails / Python or ready to pick them up quickly.
  • Experience building evaluation or observability infrastructure for ML/AI systems.
  • Experience designing datasets, annotation workflows, or labeling pipelines for ML/AI evaluation.
  • Learn fast, experiment aggressively, and thrive in a highly dynamic environment.
  • Comfortable in a fast-paced environment with ambiguity.
  • Strong alignment with our company values. You are proficient in English (CEFR Level C2 / ILR Level 5).

Responsibilities

  • Build and own our evaluation infrastructure. Design CI/CD pipelines, scorers, and datasets that determine AI agents performance.
  • Diagnose exactly where quality breaks down across the agent pipeline.
  • Grow the datasets with AI-generated and simulated conversations for faster coverage.
  • Run structured experiments across prompts, models, and the agent harness; evaluate models and optimize costs.
  • Lead and grow the AI Quality engineering team; set technical direction while staying hands-on.
  • Partner with AI Core engineering to ensure changes improve quality, not just ship features.

Skills

7+ years experience
Ruby on Rails
Python
Evaluation/observability infra
Dataset annotation workflows
Experimentation across prompts/models
Fast learner / adaptable
English proficiency (C2/ILR5)

Tools

Braintrust
LangSmith

Job description

# Lead Engineer, AI QualityLocationEurope, Middle East, Africa, Asia-Pacific (EMEA, APAC)TeamSpecial ProjectsCompensation$170K a yearApply## **About Us**Circle is building the world’s leading AI-powered, all-in-one platform for digital businesses. We make it possible for creators, coaches, educators, and businesses to bring together their audience with engaging discussions, live streams, events, chat, courses, and payments — all in one place, all under their own brand.We’re proud to be a fully remote company of around 270 (and growing!) team members from 30+ countries around the world. We seek exceptional individuals around the world, set them up to do the best work of their lives, and in turn, create a meaningful impact in their own lives. We don't track hours, but we do manage for high expectations very closely. We collaborate across time zones, are highly async, and like to document a lot.Twice a year, we bring the whole company together in beautiful places around the world for our company offsites. So far, we’ve hosted offsites in Turkey, Portugal, Mexico, Thailand, Colombia, Italy, Ireland, and more, with still more to come!Check out our **Careers** page for more about working at Circle.## About the roleThe AI Quality engineering team at Circle owns the foundation for measuring, diagnosing, and improving the quality of Circle's AI-powered features. This team focuses on building the infrastructure to measure, diagnose, and improve production AI systems, rather than ML research or model training.We're looking for a Lead Engineer to help us build out the evaluation frameworks, observability tooling, and diagnostic infrastructure that tell us whether our AI Agents are working well, where to improve them, and how to make them faster and more cost-efficient.This is a hands-on player-coach role. You’ll also lead and manage the AI Quality engineering team, with an ambitious and growing roadmap.If you're excited about making AI systems work reliably, efficiently, and at scale, this is for you.## What you'll be doing* **Build and own our evaluation infrastructure.** Design the CI/CD pipelines, scorers, and datasets that tell us whether Circle's AI agents (planners, tool-callers, and sub-agents) are actually working, from a single tool call to a full multi-turn conversation.* **Diagnose exactly where quality breaks down.** Trace failures across the agent pipeline including plan creation vs. execution, tool selection, tool trajectory in complex areas like workflows, site builder, and analytics. Turn what you find into prototypes that solve the issue or identify priorities for AI core engineering can help.* **Grow the datasets that make evaluation possible.** Stand up annotation workflows and build out AI-generated and simulated conversations so we can cover more of the product faster than manual labeling alone.* **Run structured experiments across prompts, models, and the agent harness.** Evaluate new and open-source models against our production baseline, build the framework we use to decide when to shift models, and chase cost and latency wins through model swaps, caching, and routing by plan complexity.* **Lead and grow the AI Quality engineering team.** Set technical direction and manage day-to-day priorities, all while staying hands-on in the code yourself.* **Partner closely with AI Core engineering.** Work with the engineers building Circle's AI products so they have real confidence that their changes are actually improving quality, not just shipping.## What you'll need to be successful* **7+ years of experience building and shipping production software, ideally including LLM-powered agents that take real actions in a product.** You've worked on complex, tool-using systems (multiple tools, planning or orchestration, sub-agents) not just simple, single-turn assistants, and you can walk us through something you shipped and how you knew it was actually working.* **Comfortable in Ruby on Rails / Python or ready to pick them up quickly.** Ruby on Rails is our production system and the foundation that Circle’s AI Agents are built on and proficiency in Python is a strong plus, especially for the data and evaluation side of the work.* **Experience building evaluation or observability infrastructure for ML/AI systems.** You've built eval pipelines, scorers, dashboards, or CI/CD for evals before and have experience with evaluation frameworks like Braintrust, LangSmith, or similar.* **Experience designing datasets, annotation workflows, or labeling pipelines for ML/AI evaluation.** You know how to turn raw examples into a dataset you can trust through human annotation, AI-generated data, or simulation.* **Learn fast, experiment aggressively, and thrive in a highly dynamic environment.** You’ll need to be comfortable running a lot of experiments and letting the results settle the argument, including the ones that don't work (which will be many of them in the beginning).* **Comfortable in a fast-paced environment with ambiguity.** You’ll need to learn and pick up new technologies when projects require it.* **Strong alignment with** **our company values****.*** You are proficient in English (spoken, written, and reading) at a **CEFR Level C2** / **ILR Level 5**.## Compensation & benefitsCircle offers U.S.-benchmarked compensation globally, equity in the company with ongoing refresh grants, and 35 days of paid time off each year.We’re a remote-only team that comes together twice a year for company retreats in incredible destinations around the world. Alongside incredible flexibility and autonomy, we offer a benefits package that supports health, wellbeing, and professional growth. Learn more in our Candidate Hub.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Lead Engineer, AI Platform
Lead Engineer, AI Platform

Circle • United States

Remote
USD 180,000 - 230,000
Equity with refresh grants
35 days paid time off
Remote‑only team
+2
Lead Engineer, AI Platform
Lead Engineer, AI Platform

CIRCLE • New York (NY)

Remote
USD 180,000 - 240,000
Equity in the company
35 days paid time off
Remote-only team
+1
Lead AI Quality & Platform Engineer
Lead AI Quality & Platform Engineer

CIRCLE • New York (NY)

Remote
USD 180,000 - 240,000
Equity in the company
35 days paid time off
Remote-only team
+1
Customer Support Specialist, APAC
Customer Support Specialist, APAC

Circle • Pacific (MO), Northern (KY)

Hybrid
USD 43,000 - 58,000
Equity
35 days of paid time off
Fully remote team
Quality Systems Lead
Quality Systems Lead

Encord • London (KY)

On-site
USD 119,000 - 158,000
Competitive salary
London office culture
25 days annual leave
+4
Lead AI Quality Engineer — Observability & Evaluation
Lead AI Quality Engineer — Observability & Evaluation

Circle • Northern (KY)

Hybrid
USD 150,000 - 190,000
Equity
Remote-friendly
35 days PTO
Senior SRE: AI-Driven Platform & Cloud Infra
Senior SRE: AI-Driven Platform & Cloud Infra

RADEMACHER GERÄTE-ELEKTRONIK GmbH • San Francisco (CA)

On-site
USD 153,000 - 205,000
Data Scientist — Agent Evaluations & Quality
Data Scientist — Agent Evaluations & Quality

Clera • United States

Remote
USD 120,000 - 190,000
REMOTE Senior AI/ML Engineer - SaaS
REMOTE Senior AI/ML Engineer - SaaS

CyberCoders, Inc. • United States

Remote
USD 140,000 - 240,000
Remote work
Unlimited vacation
Comprehensive benefits
+3
Principal AI Engineer
Principal AI Engineer

WeHireYou • Paris (IN)

On-site
USD 180,000 - 240,000
Competitive salary
Equity