Quality Systems Lead

Encord

London (KY)

On-site

USD 119,000 - 158,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Competitive salary
London office culture
25 days annual leave
Learning & development budget
Travel for UK & Europe
Company lunches
Team offsites

Job summary

Encord is seeking a Quality Systems Lead to own how we measure the quality of the human data we deliver to frontier AI labs and enterprises. You will build automated evaluation, manage a dedicated audit team in India, and define the standard across data types.

You’ll drive Python/SQL powered analysis, design sampling and agreement metrics, and bring quality measurement into Encord platform in collaboration with Product and Engineering.

Qualifications

  • 4+ years owning technical and operational outcomes in a data/AI environment with measurable quality.
  • Hands-on Python and SQL; build analysis and tooling rather than specify for others.
  • Experience applying models and sampling design to quality/evaluation problems.
  • Experience hiring, training and managing a distributed team (audit/review/QA).
  • Bonus: STEM degree or background in data science or research engineering.

Responsibilities

  • Build automated dataset quality evaluation and root-cause detection at scale.
  • Hire, train, calibrate and manage an audit team based in India.
  • Turn audit output into labeled ground truth to train the automated layer.
  • Own the quality standard for every data type delivered, with rubrics and acceptance criteria.
  • Build scoring systems for annotator/reviewer performance and operational decisions.
  • Set pass thresholds for certification gates and monitor quality KPIs.
  • Report quality metrics to leadership and customers.
  • Collaborate with Product and Engineering to embed quality measurement in the platform.

Skills

Python
SQL
Statistical analysis
Quality assurance
Auditing

Education

Bachelor's in STEM

Tools

Pandas
Airflow

Job description

About us

Encord is the universal data layer for AI that helps 300+ AI teams train and run models on the right data. Our platform indexes, curates, annotates, and evaluates data across the full AI lifecycle, from development through production.

Trusted by Woven by Toyota, AXA, UiPath, Zipline, and more. We're an ambitious team of 100+ working at the frontier of AI and have raised $60M in Series C funding from Wellington Management, CRV, Next47 and Y Combinator.

The role

We're hiring a Quality Systems Lead to own how Encord measures the quality of the human data we deliver to frontier AI labs, physical AI companies and enterprise AI teams — the standard itself, the systems that evaluate against it, and the audit function that produces the ground truth behind both. Data quality is what our customers buy. As we scale across data types — image and video, document, medical, LLM evaluation, robot teleoperation, egocentric capture — quality coverage cannot scale linearly with headcount. So this role has two halves that make each other work. You will build automated evaluation: model-assisted and LLM-based screening, agreement analysis at scale, anomaly and drift detection across annotation output. And you will build and run a dedicated audit team of around ten specialists in India, whose judgements become the labelled ground truth that trains and calibrates that automated layer. As coverage automates, the audit team moves up to the cases models can't judge and to generating gold sets for each new data type we take on. It is an unusual combination — engineering and consistent QC operations in one person — and it is the combination the job needs. You will also have an advantage your counterparts elsewhere in the industry don't: Encord owns the platform this work runs on, so the measurement you build can become native capability in the product rather than internal tooling.

What you'll do
  • Build automated dataset quality evaluation and root-cause detection — model-assisted and LLM-as-judge screening, agreement analysis at scale, anomaly and drift detection across annotation output

  • Hire, train, calibrate and manage a dedicated audit team of around ten specialists based in our India operation, held to inter-rater agreement and catch rate rather than volume audited

  • Turn audit output into labelled ground truth that trains and validates the automated layer, and manage the ratio of automated to manual coverage deliberately over time

  • Own the quality standard for every data type we deliver — written rubrics with worked edge cases, golden sets, and acceptance criteria agreed with the customer, alongside the Special Projects lead, before the first batch ships

  • Build scoring systems that rank annotator and reviewer performance and feed routing, staffing and offboarding decisions

  • Set the pass thresholds that certification gates on, so nobody works a project queue without having demonstrated they meet the standard

  • Report quality KPIs to leadership, and into the reporting our Special Projects leads take to customers: accuracy against client spec, inter-annotator agreement, rework rate, cost of rework, and coverage

  • Work with Project Management on remediation — you produce the measurement and the diagnosis, production owns fixing the project, and the standard stays independent of the people being measured

  • Own the unit economics of quality: cost per audited unit, and the coverage you buy per pound spent

  • Partner with Product and Engineering to bring quality measurement into Encord platform as native capability

Who we're looking for
  • You build and you operate. You'll write the evaluation pipeline, and you'll also run the weekly calibration session with ten auditors in another country

  • Your instinct on a coverage problem is to automate it — you reach for a model, a heuristic or a better sampling design before you reach for more auditors

  • Statistically literate in a practical way: sampling design, agreement statistics, and the judgement to spot a metric being optimised against rather than met

  • You can hold a calibrated standard across a distributed team you don't sit with — you know that ten uncalibrated auditors produce ten standards

  • A strong writer. Much of this job is producing rubrics a distributed workforce can follow without you in the room

  • You hold a standard under commercial pressure, and you bring the evidence that makes it stick with a delivery team or a client

  • Systems thinker: as interested in why a failure recurs across projects as in this project's defect rate

Experience requirements
  • 4+ years owning both technical and operational outcomes in a data, AI or service delivery environment where quality was measured rather than asserted

  • Hands-on Python and SQL. You build the analysis and the tooling rather than specify it for someone else

  • Practical experience applying models to a quality or evaluation problem — LLM-as-judge, model-assisted QA, automated evaluation, anomaly detection or classifier-based screening

  • Sampling methodology and agreement statistics (Cohen's and Fleiss' kappa, F1 against ground truth) applied to real production data

  • Experience hiring, training and managing a team, ideally an audit, review or QA team, and ideally distributed

  • Track record of building a quality framework or function, including the reporting leadership and customers run on

  • Bonus: direct experience of annotation, evaluation or model-training workflows, and of what frontier AI labs accept as evidence of quality

  • Bonus: a STEM degree, or a background in data science or research engineering

  • Bonus: multilingual delivery and linguistic quality assessment

Why Encord
  • Competitive salary, commission, and equity in a high-growth startup

  • Strong in-person culture — most of the team works from our London office 4+ days/week

  • 25 days annual leave + UK public holidays

  • Annual learning & development budget

  • Travel for customer visits, events, and conferences across the UK and Europe

  • Company lunches twice a week

  • Monthly socials & bi-annual team offsites

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Technical Program Manager
Technical Program Manager

Encord • London (KY)

Hybrid
USD 119,000 - 172,000
Competitive salary
London office 4+ days/week
25 days annual leave
+4
Learning & Development Specialist
Learning & Development Specialist

Encord • London (KY)

On-site
USD 66,000 - 92,000
Competitive compensation
London office-based role
Annual learning budget
+1
Strategic Projects Lead
Strategic Projects Lead

Encord • San Francisco (CA), Northern (KY)

Hybrid
USD 150,000 - 190,000
Competitive salary
Growth opportunities
4 days/week in-person
+5
Solutions Engineer
Solutions Engineer

Encord • New York (NY)

On-site
USD 120,000 - 170,000
Competitive compensation
Equity option
Learning budget
+2
Solutions Engineer
Solutions Engineer

The Consensus • New York (NY)

On-site
USD 140,000 - 190,000
Competitive salary and equity
Strong in-person culture: 4 days/week
Flexible PTO
+4
DevOps Engineer
DevOps Engineer

Encord • London (KY)

On-site
USD 119,000 - 172,000
Competitive salary
Equity
London office + hybrid work
+2
Product Lead
Product Lead

Encord • London (KY)

On-site
USD 119,000 - 172,000
Competitive equity
London office with strong in-person文化
Annual learning & development budget
+1
Senior Product Engineer
Senior Product Engineer

Crane Venture Partners • San Francisco (CA)

On-site
USD 120,000 - 190,000
Salary and equity
Growth opportunities
Loft office North Beach
+8
Senior Product Engineer
Senior Product Engineer

Encord • San Francisco (CA), Northern (KY)

Hybrid
USD 140,000 - 210,000
Competitive salary
Equity
Office in North Beach (SF)
+6
Senior Engineer - High Performance Data Systems
Senior Engineer - High Performance Data Systems

Encord • London (KY)

On-site
USD 119,000 - 198,000
Competitive salary and equity
In‑person culture — London office
25 days annual leave + UK holidays
+4