Applied AI Researcher, Benchmarking

Distyl

San Francisco (CA)

Hybrid

USD 150,000 - 250,000

Full time

4 days ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Equity
Medical, dental, vision
Flexible time off
401(k) and financial coaching
Wellness benefits
Complimentary in-office lunches
Access to AI models and tools

Job summary

Distyl AI is a frontier AI company seeking researchers to redefine how software is used in enterprise environments. You will design evaluation benchmarks and explore novel paradigms for intelligent systems, shaping metrics, reliability, and operational impact.

The role emphasizes hands-on prototyping, rigorous analysis, and publication-ready results. Distyl offers equity, a hybrid work model, and access to leading AI models for enterprise-scale problems.

Qualifications

  • Experience designing and running evaluations.
  • Strong statistical and analytical rigor.
  • Experience building with models, not just building models.
  • Proven track record of research results.
  • Uses AI daily with tools like ChatGPT, Cursor, Perplexity.
  • Strong programming and data analysis skills.
  • Bias toward showing results, not just ideas.

Responsibilities

  • Define evaluation frameworks and benchmarks.
  • Explore evaluation paradigms for intelligent systems.
  • Measure model behavior and establish methodologies for emergent capabilities.

Skills

Benchmarking design
Statistical rigor
AI systems
Research track record
Practical AI tools
Prototyping

Tools

ChatGPT
Cursor
Perplexity
Python

Job description

About Distyl AI

Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations. We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission‑critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys. Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board‑members of 20+ F500s.


What We Are Looking For

At Distyl we’re pushing the envelope of AI utilization in enterprise. This requires creative researchers who don’t just want to drive incremental improvements on benchmarks or optimize an existing process but instead are looking to creatively redefine how software is used. Our researchers come from many academic backgrounds but have strong research track records, operate in an AI-native way, and would be bored staying on the rails of a traditional research org.


Key Responsibilities


  • The Benchmarking team defines how progress is measured. Researchers design evaluation frameworks that capture reasoning depth, interaction quality, reliability, and operational impact. They construct benchmarks that reflect real-world complexity. Their systems become the standard by which new architectures, techniques, and releases are judged.

  • Researchers in Benchmarking explore new paradigms for evaluating intelligent systems: adversarial robustness testing, longitudinal performance tracking, and human-in-the-loop assessment. They investigate how metrics shape model behavior and establish rigorous methodologies for quantifying emergent capability. Their insights drive both Distyl’s internal research priorities and industry-wide standards.


Who You Are


  • Experience Designing and Running Evaluations: You’ve built or maintained benchmarks, test suites, or experimental frameworks to measure model or system performance

  • Statistical and Analytical Rigor: You design fair, reproducible experiments and can extract signal from noisy empirical results

  • Experience Building with Models, Not Just Building Models: We develop intelligent systems using models rather than training or fine-tuning them. Ideal candidates have expertise in compound AI systems, agentic collaboration, and associated techniques (ensembling, ReAct, graph-of-thoughts, etc.)

  • Proven Track Record of Research Results: Whether you’ve published in top journals, posted amazing work on twitter, or somewhere else we want to see what you've done

  • Uses AI Every Day: Before you can revolutionize someone else’s workflow, you need to revolutionize yours. You should be using tools like ChatGPT, Cursor, and Perplexity to accelerate your workflow

  • Strong Programming and Data Analysis Skills: While you might not consider yourself a software engineer you need to be able to build prototypes of your ideas and then perform the experiments to prove the effectiveness to a F500 Head of AI

  • Biases Towards Showing vs Telling: Our customers want to see the power of AI today vs discuss the most elegant idea that will take 5 years to realize


What We Offer


  • The base salary range for this role is $150K – $250K, depending on experience, location, and level. In addition to base compensation, this role is eligible for meaningful equity, along with a comprehensive benefits package

  • 100% coverage of medical, dental, and vision insurance for employee and dependents

  • Flexible time off

  • Retirement and financial planning benefits, including access to pre‑tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources

  • Comprehensive wellness benefits, including physical fitness, mental well‑being, and fertility and family‑building benefits through Carrot

  • Complimentary in‑office lunches and snacks provided

  • Access to state‑of‑the‑art AI models, generous usage of modern AI tools, and real‑world business problems

  • Ownership of high‑impact projects across top enterprises

  • A mission‑driven, fast‑moving culture that values curiosity, pragmatism, and excellence


Distyl has offices in San Francisco and New York. This role follows a hybrid collaboration model with 3+ days per week (Tuesday–Thursday) in‑office.


We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Researcher, System Discovery
Applied AI Researcher, System Discovery

Distyl • San Francisco (CA)

Hybrid
USD 150,000 - 250,000
Equity
Medical/dental/vision insurance
Flexible time off
+3
Applied AI Researcher, System Self-Improvement
Applied AI Researcher, System Self-Improvement

Distyl • New York (NY), San Francisco (CA)

On-site
USD 150,000 - 250,000
Equity
Health insurance
Flexible time off
+6
Applied AI Researcher, AI Systems
Applied AI Researcher, AI Systems

Distyl • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 250,000
Equity
Full benefits package
Flexible time off
+2
Applied AI Researcher, Post-Training
Applied AI Researcher, Post-Training

Distyl • New York (NY), San Francisco (CA)

Hybrid
USD 150,000 - 250,000
Equity
Medical insurance
Flexible time off
+6
Research Engineers, Data
Research Engineers, Data

Socket.dev • San Francisco (CA), New York (NY)

Hybrid
USD 150,000 - 250,000
Medical, dental, vision
401(k) with commuter benefits
Access to AI models & tools
+2
Organizational Development & Learning Lead
Organizational Development & Learning Lead

Distyl AI • New York (NY)

Hybrid
USD 170,000 - 190,000
Equity
Medical, dental, vision insurance
Flexible time off
+6
Organizational Development & Learning Lead
Organizational Development & Learning Lead

Distyl AI, Inc. • New York (NY)

Hybrid
USD 170,000 - 190,000
Equity
Benefits package
Medical coverage
+7
AI Engineer
AI Engineer

Distyl • New York (NY)

Hybrid
USD 150,000 - 250,000
100% covered medical, dental, and vision
401(k) with commuter benefits
Access to modern AI tools
Applied AI Researcher, Multi-Agent Systems
Applied AI Researcher, Multi-Agent Systems

Distyl AI • San Francisco (CA)

Hybrid
USD 150,000 - 250,000
100% covered medical, dental, and vision for employees and dependents
401(k) plan
Commuter benefits
+1
Organizational Development & Learning Lead
Organizational Development & Learning Lead

Distyl • San Francisco (CA)

Hybrid
USD 170,000 - 190,000
Equity
Medical, dental, and vision insurance
Flexible time off
+4