Senior Applied Machine Learning Engineer (Eval)

Carrerlift

Bengaluru

On-site

INR 1,900,000 - 3,200,000

Full time

3 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Cardboard, Bengaluru-based, is seeking a Senior Applied Machine Learning Engineer to own the quality feedback loop for its agentic video editor. The role focuses on turning production failures into trusted evaluations and measurable improvements.

You will define quality standards, build evaluation datasets, and implement offline/online evaluation systems. Strong ownership and experience with LLMs are essential for success.

Qualifications

  • Shipping and operating an LLM or agent system used by real customers.
  • Strong software skills in TypeScript or Python and ability to work across both languages.
  • Experience building evaluations, datasets, experiments or AI quality systems.

Responsibilities

  • Define quality standards for Cardboard's AI agent.
  • Build reliable evaluation datasets using data from real product usage.
  • Create offline and online evaluation systems using automated checks, model-based graders and human review.
  • Analyse real agent runs to identify repeated failure patterns.
  • Improve agent quality through better datasets, evaluation approaches, model selection and fine-tuning.
  • Build regression checks and release gates for important agent updates.
  • Monitor AI quality together with latency and cost.
  • Partner with product and engineering teams to ship measurable quality improvements.

Skills

LLM systems
Python
TypeScript
Evaluation design
Datasets
Model fine-tuning
Quality assurance
Production monitoring

Job description

About the role:
  • Cardboard is hiring a Senior Applied Machine Learning Engineer focused on evaluations and AI quality.
  • The role is based onsite in Bengaluru, India.
  • The engineer will own the quality feedback loop for Cardboard's agentic video editor.
  • The work involves turning failures observed in production into trusted evaluations and measurable improvements.
What you'll work on:
  • Define quality standards for Cardboard's AI agent.
  • Build reliable evaluation datasets using data from real product usage.
  • Create offline and online evaluation systems using automated checks, model-based graders and human review.
  • Analyse real agent runs to identify repeated failure patterns.
  • Improve agent quality through better datasets, evaluation approaches, model selection and fine-tuning.
  • Build regression checks and release gates for important agent updates.
  • Monitor AI quality together with latency and cost.
  • Partner with product and engineering teams to ship measurable quality improvements.
What we're looking for:
  • Candidates should have experience shipping and operating an LLM or agent system used by real customers.
  • Strong software engineering skills in TypeScript or Python are required, with the ability to work across both languages.
  • Experience building evaluations, datasets, experiments or AI quality systems is expected.
  • Applicants should have strong product judgement and be able to convert vague AI-quality issues into measurable problems.
  • Candidates should be comfortable working across data, evaluation methods, model selection and fine-tuning.
  • Strong ownership as a senior individual contributor is required.
Good to have:
  • Experience with multimodal AI, video, media or creative software is considered a plus.
  • Experience with human labeling, model graders or fine-tuning is also beneficial.
  • Strong understanding of experiment design and statistics is another bonus qualification.
About Cardboard:
  • Cardboard is an AI-first video editor building agentic tools that interpret user requests, work with media and make actual edits on a timeline.
  • The company is backed by a Tier-1 global fund, YC and founders of billion-dollar companies.
Interview process:
  • The process includes Recruiter Screen, Technical Interview, ML & Evaluation Deep Dive, Product & Engineering Interview, and Final Interview.

Note: ClanX is the recruitment partner helping Cardboard hire for this role.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Applied Machine Learning Engineer (Eval) - Bengaluru
Senior Applied Machine Learning Engineer (Eval) - Bengaluru

CLANS.AI PTE. LTD. • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Senior Applied ML Engineer, Evals & Data
Senior Applied ML Engineer, Evals & Data

Engg • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Founding-team equity
Unlimited AI model tokens
Budget for AI tools
Senior Applied ML Engineer, Evals & Data
Senior Applied ML Engineer, Evals & Data

Cardboard Inc. • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Founding-team equity
Senior Applied ML Engineer, Evals & Data
Senior Applied ML Engineer, Evals & Data

Cardboard • Bengaluru

On-site
INR 1,800,000 - 2,600,000
Equity
Unlimited AI tokens
AI tools budget
Fullstack Engineer
Fullstack Engineer

BookMyMentor • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Competitive salary
Founding-team equity
Unlimited AI tokens
+1
AI Evaluation Engineer - Proofline
AI Evaluation Engineer - Proofline

Fermi AI • Bengaluru

On-site
INR 1,500,000 - 2,100,000
Senior Applied AI Engineer (Fine Tuning)
Senior Applied AI Engineer (Fine Tuning)

Carrerlift • India

Remote
INR 3,600,000 - 4,800,000
Evals Engineer
Evals Engineer

PingAura AI Technologies Private Limited. • Mumbai

On-site
INR 1,000,000 - 1,500,000
Founding-tier ownership
Direct collaboration with AI experts
Competitive salary
Lead Engineer, AI Platform
Lead Engineer, AI Platform

Lever, Inc. • India

Remote
INR 14,451,000 - 18,304,000
Annual equity refresh
Fully remote work environment
Health & wellbeing benefits
+2
Software Engineer, Evals Bangalore, India
Software Engineer, Evals Bangalore, India

Startups • Bengaluru

On-site
INR 4,000,000 - 7,000,000