Research Engineer

Ando

San Francisco (CA)

On-site

USD 150,000 - 210,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Free Equinox membership
Equity grant (4 years vesting)
Health, dental, vision insurance

Job summary

Ando is a bold research-backed messaging platform in San Francisco seeking to push the boundaries of agent-enabled collaboration. You’ll join a small, fast-moving team working close to data and product to build robust evaluation on real workspace data.

The role emphasizes experimentation, collaboration with researchers and product teams, and publishing meaningful progress. Candidates should bring strong applied research, evaluation, and communication skills and enjoy working with imperfect data.

Qualifications

  • Strong applied research background with depth in model evaluation, benchmarking, and/or failure analysis.
  • Evidence over credentials with work samples or code demonstrating eval frameworks, benchmark suites, or tooling.
  • Strong technical communication to explain complex ideas with researchers and product teams.
  • Comfort with messy, incomplete workspace data and edge cases.
  • Bonus: familiarity with simulation, human-in-the-loop evaluation, memory/compression work, or agent observability tooling.

Responsibilities

  • Build evaluation from real data, designing label schemas and rigorous labelling pipelines.
  • Run experiments that ship, bridging offline data and production systems.
  • Decide what to build in-house versus with partners, evaluating frontier vendors and teams.
  • Publish when meaningful progress is achieved, contributing final results to stakeholders.

Skills

Applied research
Model evaluation
Benchmarking
Experiment design
Technical communication
Langsmith

Tools

Langsmith

Job description

Ando is a messaging platform where AI agents take on work alongside their human teammates. We’re rebuilding Slack from the ground up around two core ideas: durable memory and agents as first-class participants.

We have real, longitudinal, multi-party workspace data, and human / agent users whose behavior tells you whether the systems you're designing are actually meaningfully improving. Working with live business communication means permissions, redaction, and security are things we have to think about as well. If you want to have your research come into contact with reality, Ando is the place for it.

Some concrete research problems on our plate right now:

  • Proactivity: An agent embedded in a team's channels has to decide, message by message, whether to ignore, quietly track, or intervene. Evaluating that judgment means building benchmarks where the ground truth includes silence. Existing agent benchmarks are almost entirely reactive.

  • Memory: What should a workspace agent remember across weeks and months of participation, in what representation, and how do you measure whether memory is helping versus hurting agent usefulness?

  • Continual Learning: How much does basic memory, retrieval, and context impact agent performance, vs where do we genuinely need RL and continual learning?

What you’ll be doing at Ando

You'll be one of the first members of a small research team, working close to both the data and the product.

  • Build evaluation from real data - Mine production workspace data (carefully, with consent and redaction pipelines you'll help design) into benchmarks and labeled datasets. Design label schemas, run labeling with real inter-rater rigor, and build the harnesses that make expert judgment cheap to capture and hard to corrupt.

  • Run experiments that ship - Initial work happens on offline workspace data; the destination is production systems used by every Ando customer. The distance between "the benchmark improved" and "the feature shipped" should be weeks and you'll own both ends.

  • Decide what we build versus who we partner with - We won't do everything in-house. Part of the job is evaluating frontier vendors and research teams (eval infrastructure, observability, continual-learning tooling) and choosing who we build with.

  • Publish when we have something real - We expect the team to publish as we make meaningful progress. But ideas are not Ando's moat; execution, product quality, and customer experience are. Research here is in service of customers first, and the publications will be better for it: they'll describe things that actually worked on real data.

What we\'re looking for
  • Strong applied research background, with depth in model evaluation, benchmarking, and/or failure analysis. You\'ve built evals you trusted enough to make decisions with.

  • Evidence over credentials. Work samples or code that demonstrate the skills: eval frameworks, benchmark suites, failure-analysis reports or tooling, labeling infrastructure. Show us something you built to find out whether a system actually worked.

  • Strong technical communication. You can explain complex ideas simply and hold high-bandwidth, generative technical conversations with researchers and with our product team.

  • Comfort with mess. Real workspace data is incomplete, ambiguous, and full of edge cases that break clean abstractions. You treat that as signal, not noise.

  • Bonus: familiarity with simulation (Park et al.), human-in-the-loop evaluation (Scale HIL leaderboard), Cartridges and related context/memory-compression work, memory for multi-party long-running settings, or agent observability standards (setting up Langsmith or similar).

Hiring process

  1. 30 min intro call

  2. Technical conversation - walk us through an app you\'ve shipped

  3. Paid take-home or IRL work trial

Benefits

  • Free Equinox membership & other health perks

  • Generous equity grant vested over 4 years

  • Health, dental, vision insurance

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Product Engineer
Product Engineer

Ando • San Francisco (CA)

On-site
USD 180,000 - 240,000
Free Equinox membership
Health, dental, vision insurance
Generous equity grant vested over 4 4
+1
Mobile Engineer
Mobile Engineer

Ando • San Francisco (CA)

On-site
USD 120,000 - 160,000
Free Equinox membership
Equity grant
Health, dental, vision insurance
AI Researcher / Engineer / Intern
AI Researcher / Engineer / Intern

Egra • New York (NY)

On-site
USD 120,000 - 160,000
Competitive salary and meaningful equity
Platinum-tier health insurance
Uncapped compute access
Founding Research Engineer
Founding Research Engineer

Ambral (YC S25) • San Francisco (CA)

Hybrid
USD 140,000 - 210,000
Equity and ownership
Equinox membership
Free meals, coffee, snacks
+2
Founding Research Engineer
Founding Research Engineer

Ambral (YC S25) • New York (NY)

On-site
USD 215,000 - 330,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Memory-Driven AI Product Engineer
Memory-Driven AI Product Engineer

Ando • San Francisco (CA)

On-site
USD 150,000 - 210,000
Free Equinox membership
Equity grant (4 years vesting)
Health, dental, vision insurance
Founding Research Engineer
Founding Research Engineer

Jobzhr • New York (NY)

On-site
USD 140,000 - 210,000
Significant equity and ownership
Equinox membership
Free meals, coffee, and snacks
+2
Applied Research Scientist
Applied Research Scientist

Fleet AI, Inc. • Buffalo (NY)

On-site
USD 150,000 - 210,000
Founding Research Engineer
Founding Research Engineer

Ambral (YC S25) • United States

On-site
USD 215,000 - 330,000
Significant equity
Equinox membership
Free meals, coffee, and snacks
+2
Mobile Engineer
Mobile Engineer

Worky • San Francisco (CA)

On-site
USD 140,000 - 200,000
Equity grant vested over 4 years
Health, dental, vision insurance
Free Equinox membership