Machine Learning Engineer (Remote)

Ocho People Ltd.

Belfast City District

Remote

GBP 90,000 - 140,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Benefits offered by this job

Equity
Remote work in UK

Job summary

Ocho People Ltd. is seeking a Machine Learning Engineer to build and own the evaluation function for production LLM and agent systems.

You’ll work across Python, NLP, and GCP, crafting benchmarks and automated scoring pipelines, with a hands-on role in shaping how AI quality is measured before release. The role is fully remote from Northern Ireland, offering a substantial base salary (£90,000–£140,000) plus equity, and will involve close collaboration with data science and engineering teams to

Qualifications

  • 5+ years in ML engineering, NLP or applied data science.
  • Hands-on experience with LLMs or agent-based systems.
  • Strong Python and data-pipeline construction for evaluation datasets.
  • Solid understanding of NLP and LLM capabilities including prompting techniques and retrieval.
  • Strong statistics foundation including sampling and variance.
  • Experience with evaluating mechanisms and benchmark systems for ML/LLM outputs.

Responsibilities

  • Design evaluation frameworks and metrics for accuracy, safety, latency and cost.
  • Build benchmark eval sets reflecting real customer scenarios and edge cases.
  • Develop automated scoring pipelines using rubric-based grading and LLM-as-judge techniques.
  • Calibrate automated judges against human review to ensure trustworthy numbers.
  • Stand up regression suites to catch quality drops before release.
  • Create dashboards tracking model and agent quality over time and releases.
  • Investigate failure modes in multi-step agent behavior and prioritize fixes with engineering.
  • Collaborate with senior data scientist on deeper statistical analyses.
  • Balance evaluation coverage against compute and API cost.

Skills

ML engineering
NLP experience
Python pipelines
Statistical analysis

Education

Degree in CS/ML or related

Tools

GCP Vertex AI
BigQuery

Job description

Machine Learning Engineer | LLM & Agent Evaluation | Python | Remote UK

At-a-Glance

  • Own evaluation for production LLM and agent systems at a funded AI start-up
  • Mid to senior individual contributor role across Python, NLP, LLMs and GCP (Vertex AI, BigQuery)
  • Fully remote from Northern Ireland
  • £90,000 to £140,000 base plus equity
  • A new role: you will build the evaluation function from the ground up

About the Company

Our client is a fast-growing AI company building an agentic quality assurance platform that autonomously creates, runs and maintains software tests, so engineering teams can ship with confidence as AI writes more of their code. Its platform runs on custom machine learning models built specifically for testing, and it is trusted by globally recognised enterprise brands across technology, healthcare and professional services. The business is raising its next funding round and is on track for significant growth, with a small, high-calibre Belfast team now expanding into Data Science and Machine Learning.

The Role

This is a new role and the first dedicated evaluation hire, so you will define how the business measures the quality of its AI. The product is an AI agent, and every change to a model, prompt or piece of agent logic can quietly make it better or worse. Your job is to know which, before customers do.

You will work closely with the engineers building agent capabilities and with the team's senior data scientist, and report into the Head of Data Science. The team is small, so breadth matters as much as depth, and you will regularly pick up work outside your core specialism.

The people who thrive here are hands‑on and curious about the business itself, not just the technology. You will be comfortable in the data, confident learning on the job, calm when priorities shift, and you will know when an evaluation is good enough to ship and when it needs more work. Cost is a real constraint in a start‑up, so you will think about what each LLM call and pipeline costs as naturally as how accurate it is.

Key Responsibilities

  • Design evaluation frameworks and metrics covering accuracy, safety, latency and cost across agent and LLM systems
  • Build benchmark eval sets that reflect real customer scenarios and edge cases
  • Develop automated scoring pipelines using rubric‑based grading and LLM‑as‑judge techniques
  • Calibrate automated judges against human review so the team can trust the numbers
  • Stand up regression suites that catch quality drops from model, prompt or agent‑logic changes before release
  • Create dashboards that track model and agent quality over time and across releases
  • Investigate failure modes in multi‑step agent behaviour and prioritise fixes with the engineering team
  • Partner with the senior data scientist on deeper statistical analysis of results
  • Balance evaluation coverage against compute and API cost

What You'll Need

Essential:

  • 5+ years in ML engineering, NLP or applied data science, with hands‑on experience of LLM or agent‑based systems
  • Practical experience building or operating evaluation frameworks, automated scoring or benchmark systems for ML or LLM outputs
  • Strong Python and experience building data pipelines for evaluation datasets
  • Solid understanding of NLP and modern LLM capabilities, including prompting techniques, agentic workflows and retrieval
  • A strong statistics foundation, including sampling, variance and judging whether a change in results is meaningful
  • Ability to design metrics that represent real‑world quality, not just benchmark scores
  • Working experience with GCP (Vertex AI, BigQuery) or equivalent hands‑on experience with another major cloud ML platform
  • A degree in Computer Science, Machine Learning, Statistics or a related field, or equivalent practical experience
  • Right to work in the UK

Desirable / Nice to Have:

  • Experience with LLM‑as‑judge techniques, rubric design or human‑in‑the‑loop evaluation programmes
  • Familiarity with agent architectures and the failure modes specific to multi‑step agentic systems
  • Experience operating evaluation systems at scale in production

Why Apply?

  • £90,000 to £140,000 base salary plus equity in a business heading into its next funding round
  • Fully remote, must be based in Northern Ireland
  • Async‑first culture that trusts you to manage your own time and deliver
  • Build a function from scratch, with direct influence over how the product's AI is measured and released
  • Real breadth: a small team where you will work across engineering, data and ML rather than in a narrow lane
  • A clear, transparent five‑stage interview process, with AI tools actively encouraged in the take‑home exercise
  • A long‑term career home in a company that is investing seriously in its data and ML function
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Machine Learning Engineer
Machine Learning Engineer

Oscar Technology • City Of London

On-site
GBP 95,000 - 120,000
Machine Learning Engineer
Machine Learning Engineer

Oscar Technology • Greater London

On-site
GBP 95,000 - 120,000
AI Quality Engineer
AI Quality Engineer

SoCode Recruitment • Greater London

On-site
GBP 70,000 - 90,000
Employee shares
Private health insurance
Unlimited holidays
+1
Applied AI Engineer
Applied AI Engineer

WeDoTech • Greater London

Hybrid
GBP 84,000 - 140,000
AI Engineer - LLM & Agentic Systems
AI Engineer - LLM & Agentic Systems

Intec Select • Greater London

On-site
GBP 111,000 - 120,000
Applied AI Engineer
Applied AI Engineer

Artificial Intelligence Jobs • Greater London

On-site
GBP 120,000 - 140,000
Remote Senior Machine Learning Engineer
Remote Senior Machine Learning Engineer

NLP PEOPLE • Stirling

Remote
GBP 85,000 - 135,000
Monthly Health & Wellness budget
Learning & Development budget
Flexible working environment
+7
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

Callosum Technologies Ltd. • Greater London

On-site
GBP 75,000 - 120,000
Visa sponsorship
Relocation benefits
Research Engineer, Benchmarking - Member of Technical Staff
Research Engineer, Benchmarking - Member of Technical Staff

AI Startups UK • Greater London

On-site
GBP 120,000 - 180,000
Competitive salary
Equity & ownership
Private healthcare
+2
AI Solutions Engineer - Hyrbid - London
AI Solutions Engineer - Hyrbid - London

Initi8 Recruitment • Greater London

Hybrid
GBP 85,000 - 110,000