Senior Applied ML Engineer, Evals & Data

Cardboard, Inc.

India

On-site

INR 3,200,000 - 5,200,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Unlimited tokens across AI models
AI tools budget

Job summary

Cardboard, Inc. is hiring a Senior Applied ML Engineer in India to own how we measure and improve the quality of our AI agent. You will define what is considered good, build trusted evaluation datasets, and implement offline and online evaluation pipelines across product usage.

You will study agent runs, identify failure patterns, and push for data-driven improvements through model selection, fine-tuning, and regression gates. Strong TS or Python skills are essential.

Qualifications

  • Shipped and operated an LLM or agent system used by real customers.
  • Strong software engineering in TypeScript or Python.

Responsibilities

  • Define what "good" means for our agent and build trusted evaluation datasets.
  • Build offline and online evaluations including automated checks and model graders.
  • Study real agent runs to identify failure patterns and drive improvements.
  • Add regression checks and release gates with quality and cost considerations.
  • Collaborate with product and engineering to ship measurable AI quality improvements.

Skills

TypeScript
Python
AI quality systems
Evaluation datasets
Model evaluation

Job description

About

We're building the future of storytelling and video editing.

We're a small team that moves fast and builds things we're proud of.

We care obsessively about taste: in design, in product, in every detail.

We're backed by a Tier-1 global fund, YC, and founders of billion dollar companies.

Video is the most powerful way humans tell stories. It always has been. But creating it today is still painfully hard.Fragmented tools, steep learning curves, and workflows that get in the way of the actual creative work. We're buildingCardboard to change that.

Cardboard is an AI-first video editor. Our agent understands a user's request, works with their media, and makes realedits on the timeline. When it gets an edit right, it feels like magic. When it gets one wrong, it costs someone theirafternoon. What separates those two outcomes is measurement.

We have the base of an evaluation system. We're hiring a Senior Applied ML Engineer to build the feedback loop on top ofit: the thing that turns production failures into evaluation cases, cases into a quality bar the team trusts, and thatbar into shipped improvements.

This is a senior individual contributor role, and it is not research-only, prompt-only, or QA. You'll own how we measureand improve the quality of Cardboard's agent, end to end, and you'll work with product and engineering to ship theimprovements rather than hand off a report.

What you'll actually do

Define what "good" means for our agent, and build evaluation datasets we trust out of real product usage.

Build offline and online evaluations: automated checks, model graders, and human review where it's the only honestsignal.

Study real agent runs, find the failure patterns, and close them through better data, better evaluation methods, modelselection, and fine-tuning where it earns its keep.

Add regression checks and release gates, and track quality alongside latency and cost.

You've shipped and operated an LLM or agent system that real customers used, not just a demo.

You're a strong software engineer in TypeScript or Python, and you can work across both.

You've built evaluations, datasets, experiments, or AI quality systems before.

You have strong product judgment. You can take a vague complaint about the agent feeling dumb and turn it intosomething measurable, then move the number.

You don't need a PhD or foundation-model training experience. Evidence that you've built reliable AI products mattersmore to us than credentials or any specific framework.

Bonus: multimodal AI, video, media, or creative software.

Bonus: you know experiment design and statistics well.

What success looks like

Within your first six months:

We have a quality baseline for our main agent workflows that the team actually trusts.

Production failures regularly become new evaluation cases.

Important agent changes pass clear regression checks before release.

We can show measurable improvements in key editing workflows.

What you get

You'd be surrounded by people who are absurdly good at what they do. One started coding at 11 and shipped an app with 6M+ downloads in high school. One got into CS engineering at 14 and has been working on distributed systems for 8+ years. One's an ex-founder who took a company to 1.2M users and $300M+ in transactions. That's the team. We're looking for someone who'll raise the bar on how we measure and improve AI quality. Apart from that you'd get:

Competitive salary and founding-team equity.

Unlimited tokens across every AI model. Use whatever you want, as much as you want.

A healthy budget for AI tools and any peripherals you need to do your best work.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Fullstack Engineer
Fullstack Engineer

BookMyMentor • Bengaluru

On-site
INR 3,000,000 - 6,000,000
Competitive salary
Founding-team equity
Unlimited AI tokens
+1
Founding Design Engineer
Founding Design Engineer

Usecardboard • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Competitive salary and equity
Unlimited tokens across AI models
Healthy budget for AI tools and peripherals
Backend Engineer
Backend Engineer

Cardboard, Inc. • Bengaluru

On-site
INR 1,500,000 - 2,700,000
Founding-team equity
Unlimited tokens across AI models
Budget for AI tools and peripherals
Evals Engineer
Evals Engineer

PingAura AI Technologies Private Limited. • Mumbai

On-site
INR 1,000,000 - 1,500,000
Founding-tier ownership
Direct collaboration with AI experts
Competitive salary
AI Agent Architect
AI Agent Architect

Emergent Labs • Bengaluru

On-site
INR 1,500,000 - 2,500,000
Daily meals provided
Family insurance coverage
Unlimited paid time off
+1
AI Agent Engineer
AI Agent Engineer

PingAura AI Technologies Private Limited. • Mumbai

On-site
INR 1,200,000 - 2,000,000
Founding Storyteller
Founding Storyteller

Oximy • Karnataka

On-site
INR 600,000 - 900,000
Competitive salary and equity
MacBook Pro
Unlimited token budget
+3
Founding Designer
Founding Designer

Usecardboard • Bengaluru

On-site
INR 1,200,000 - 3,600,000
Competitive salary and equity
Unlimited tokens
Healthy budget for tools and equipment
AI-Native Full-Stack Engineer
AI-Native Full-Stack Engineer

Infinity Reach - The Digital Marketing Agency • Hyderabad

On-site
INR 1,200,000 - 2,400,000
Applied AI Engineer
Applied AI Engineer

Lifesight • Karnataka

On-site
INR 3,000,000 - 6,000,000