Senior Applied ML Engineer, Evals & Data

Cardboard

Bengaluru

On-site

INR 1,800,000 - 2,600,000

Full time

9 days ago
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Equity
Unlimited AI tokens
AI tools budget

Job summary

Cardboard is building a browser-based video editor with a cloud pipeline and AI agent. You will own how we measure and improve the AI agent's quality, studying real runs and turning failures into evaluation cases. You will collaborate with product and engineering to ship improvements this week, not next quarter.

You work with a small, craft-focused team and own problems end-to-end, balancing quality, latency, and cost while shipping useful features rapidly.

Qualifications

  • Experience shipping and operating an LLM or agent system used by customers.
  • Strong software engineering skills in TypeScript or Python across both.
  • Experience building evaluations, datasets, experiments, or AI quality systems.
  • Ability to translate unclear quality problems into measurable improvements.

Responsibilities

  • Define quality standards and build evaluation datasets from real product usage.
  • Create offline and online evaluations, including automated checks and human review.
  • Analyze model and agent failure patterns and improve quality through data, methods, and tuning.
  • Add regression checks and release gates while tracking quality, latency, and cost.
  • Solve these hard problems and ship meaningful improvements.

Skills

LLM/agent systems
Product judgment
Quality evaluation
Cross-functional collaboration

Tools

TypeScript
Python

Job description

Engineering
  • We're building the future of storytelling and video editing.
  • We're a small team that moves fast and builds things we're proud of.
  • We care obsessively about taste: in design, in product, in every detail.
  • We're solving these hard problems.
  • We're backed by a Tier-1 global fund, YC, and founders of billion dollar companies.
About
  • We're building the future of storytelling and video editing.
  • We're a small team that moves fast and builds things we're proud of.
  • We care obsessively about taste: in design, in product, in every detail.
  • We're solving these hard problems.
  • We're backed by a Tier-1 global fund, YC, and founders of billion dollar companies.
Engineering

Video is the most powerful way humans tell stories. It always has been. But creating it today is still painfully hard. Fragmented tools, steep learning curves, and workflows that get in the way of the actual creative work. We're building Cardboard to change that. Cardboard runs a real video editor in the browser, backed by a serious cloud media pipeline and an AI agent that actually understands footage. You will own how we measure and improve the quality of Cardboard's AI agent. You will study real agent runs, turn important failures into evaluation cases, and measure whether changes make the product better. You will also work with product and engineering to ship those improvements. This is not a research-only, prompt-only, or QA role. You'll be working alongside a team of engineers who all care deeply about craft, including the founders. You like owning problems end to end, and you'd rather ship something great this week than something perfect next quarter

What you'll actually do
  • Define quality standards and build trusted evaluation datasets from real product usage.
  • Build offline and online evaluations, including automated checks and human review.
  • Analyze model and agent failure patterns, then improve quality through better data, evaluation methods, model selection, and, where useful, fine-tuning.
  • Add regression checks and release gates while tracking quality, latency, and cost.
  • Solve these hard problems.
What we are looking for
  • Experience shipping and operating an LLM or agent system used by real customers.
  • Strong software engineering skills in TypeScript or Python, with the ability to work across both.
  • Experience building evaluations, datasets, experiments, or AI quality systems.
  • Strong product judgment and the ability to turn unclear quality problems into measurable improvements.

You do not need a PhD or experience training foundation models. Evidence of building reliable AI products matters more than formal credentials or knowledge of a specific framework.

Nice to have
  • Experience with multimodal AI, video, media, or creative software.
  • Experience with human labeling, model graders, or fine-tuning.
  • Good knowledge of experiment design and statistics.
Within your first six months:
  • We have a trusted quality baseline for our main agent workflows.
  • Production failures regularly become new evaluation cases.
  • Important agent changes pass clear regression checks before release.
  • We can show measurable improvements in key editing workflows.
What you get

You'd be surrounded by people who are absurdly good at what they do. One started coding at 11 and shipped an app with 6M+ downloads in high school. One got into CS engineering at 14 and has been working on distributed systems for 8+ years. One's an ex-founder who took a company to 1.2M users and $300M+ in transactions. That's the team. We're looking for someone who'll raise the bar on technical craftsmanship and creative product quality. Apart from that you'd get:

  • Competitive salary and founding-team equity.
  • Unlimited tokens across every AI model. Use whatever you want, as much as you want.
  • A healthy budget for AI tools and any peripherals you need to do your best work.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Senior Applied ML Engineer, Evals & Data
Senior Applied ML Engineer, Evals & Data

Cardboard Inc. • Bengaluru

On-site
INR 2,500,000 - 5,000,000
Founding-team equity
Senior Applied ML Engineer, Evals & Data
Senior Applied ML Engineer, Evals & Data

Engg • Bengaluru

On-site
INR 4,000,000 - 6,000,000
Founding-team equity
Unlimited AI model tokens
Budget for AI tools
Fullstack Engineer
Fullstack Engineer

Cardboard • Bengaluru

On-site
INR 900,000 - 1,300,000
Competitive salary
Unlimited tokens across AI models
Budget for AI tools
Fullstack Engineer
Fullstack Engineer

BookMyMentor • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Competitive salary
Founding-team equity
Unlimited AI tokens
+1
Senior Applied Machine Learning Engineer (Eval)
Senior Applied Machine Learning Engineer (Eval)

Carrerlift • Bengaluru

On-site
INR 1,900,000 - 3,200,000
Backend Engineer
Backend Engineer

Cardboard, Inc • Bengaluru

On-site
INR 1,500,000 - 3,000,000
Competitive salary and founding-team
Unlimited tokens across all AI models
Budget for AI tools and peripherals
Senior Applied Machine Learning Engineer (Eval) - Bengaluru
Senior Applied Machine Learning Engineer (Eval) - Bengaluru

CLANS.AI PTE. LTD. • Bengaluru

On-site
INR 1,800,000 - 3,000,000
Applied AI Engineer
Applied AI Engineer

Adam • Bengaluru

On-site
INR 2,500,000 - 4,200,000
AI Engineering Intern: Applied ML & Agentic Systems
AI Engineering Intern: Applied ML & Agentic Systems

Yodaplus Technologies Private Limited • India

Hybrid
INR 391,000 - 781,000
Stipend
PPO opportunity
Certificate
+1
Newpage - Full stack AI engineer - Python/React.js
Newpage - Full stack AI engineer - Python/React.js

Newpage Solutions • Maharashtra

On-site
INR 2,400,000 - 4,200,000