Software Engineer, AI Research – Prototyping

Jobtailor

Palo Alto (CA)

On-site

USD 120,000 - 180,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking a hands-on engineer to advance voice AI initiatives, focusing on LLM reasoning and agent systems. You will design controlled experiments, build rapid prototypes, and evaluate models and providers against real conversation data.

You will turn investigations into shareable artifacts and collaborate with platform engineers to push validated ideas to production, while growing our internal AI knowledge base.

Qualifications

  • Proficiency in Python for building experiments and tooling.
  • Strong ML background with hands-on experience in evaluating models and techniques.
  • Experience designing controlled experiments with baselines and honest measurement.
  • Experience with prompting and evaluating LLMs across providers and techniques.
  • Track record of teaching, workshops, technical writing, or internal tech talks.

Responsibilities

  • Track emerging techniques in voice AI, LLM reasoning, and agent systems, and identify which ones matter for us
  • Design structured experiments with real controls: know when a result is signal and when it is noise
  • Build rapid prototypes and test them against real conversation data
  • Run head-to-head evaluations of models, providers, and techniques (reasoning approaches, speech models, orchestration patterns)
  • Turn every investigation into a team-usable artifact: a benchmark, a written deep-dive, a tech talk, or a recommendation with evidence
  • Work with platform engineers to hand off validated ideas for production implementation
  • Build the internal knowledge base for how and why our AI stack works the way it does

Skills

Python Programming
Applied Machine Learning
Experimental Design
LLM Experience
Knowledge Transfer

Job description


  • Track emerging techniques in voice AI, LLM reasoning, and agent systems, and identify which ones matter for us

  • Design structured experiments with real controls: know when a result is signal and when it is noise

  • Build rapid prototypes and test them against real conversation data

  • Run head-to-head evaluations of models, providers, and techniques (reasoning approaches, speech models, orchestration patterns)

  • Turn every investigation into a team-usable artifact: a benchmark, a written deep-dive, a tech talk, or a recommendation with evidence

  • Work with platform engineers to hand off validated ideas for production implementation

  • Build the internal knowledge base for how and why our AI stack works the way it does


Requirements


  • 3+ years of software engineering or applied ML experience

  • Strong Python skills; able to build and run your own experiments end-to-end without infrastructure support

  • Hands‑on experience with LLMs: prompting, evaluation, and an intuition for how model behavior changes across techniques and providers

  • Experimental rigor: experience designing tests with controls, baselines, and honest measurement

  • A track record of teaching or knowledge transfer in some form: teaching or TA experience, workshops, technical writing, internal tech talks, well‑documented open source, or developer education

  • Intellectual honesty: comfortable reporting that a promising idea did not work.


Core Competencies

Demonstrates expertise in Python programming and applied machine learning, with a strong focus on experimental design and knowledge transfer. Capable of building prototypes and conducting rigorous evaluations of AI models and techniques.


Highest‑signal resume keywords


  • Python Programming

  • Applied Machine Learning

  • Experimental Design

  • LLM Experience

  • Knowledge Transfer


ATS Optimization Keywords

Hard Skills


  • Software Engineering

  • Experiment Design

  • Model Evaluation

  • Prototyping

  • Data Analysis


Soft Skills


  • Intellectual Honesty

  • Teaching Experience

  • Technical Writing


Industry Keywords


  • Voice AI

  • LLM Reasoning

  • Agent Systems

  • Benchmarking

  • Tech Talks

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Applied AI Scientist, Senior/Staff
Applied AI Scientist, Senior/Staff

Jobtailor • United States

On-site
USD 120,000 - 180,000
Principal Machine Learning Scientist – Agentic Experiences
Principal Machine Learning Scientist – Agentic Experiences

Jobtailor • California (MO)

On-site
USD 150,000 - 190,000
Lead Machine Learning Engineer
Lead Machine Learning Engineer

Jobtailor • California (MO)

On-site
USD 140,000 - 210,000
Research Engineer, Research Scientist, AI Systems Engineer
Research Engineer, Research Scientist, AI Systems Engineer

Jobtailor • San Francisco (CA)

On-site
USD 180,000 - 280,000
Applied Researcher II – AI Foundations, LLM Core, Agentic AI
Applied Researcher II – AI Foundations, LLM Core, Agentic AI

Jobtailor • California (MO)

On-site
USD 170,000 - 210,000
AI Implementation Engineer
AI Implementation Engineer

Jobtailor • New Jersey

On-site
USD 150,000 - 210,000
Head of AI
Head of AI

Jobtailor • New York (NY)

On-site
USD 180,000 - 280,000
Software Engineer, AI Platform
Software Engineer, AI Platform

Triwill Group • San Francisco (CA)

Hybrid
USD 140,000 - 180,000
Principal AI/ML Engineer
Principal AI/ML Engineer

Jobtailor • United States

On-site
USD 180,000 - 240,000
Senior Manager, AI Engineering – People Leader, Gen AI Platform Services
Senior Manager, AI Engineering – People Leader, Gen AI Platform Services

Jobtailor • California (MO)

On-site
USD 180,000 - 280,000