Founding Software Engineer — AI & Voice Systems

Roark

United States

Remote

USD 150,000 - 230,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Roark is seeking a founding-level engineer to build the evaluation and measurement infrastructure for an AI voice calling platform. You will score production calls in near real time, layer audio-native signals on top of LLM evaluators, and power a simulation engine that places realistic automated calls.

You will own large-scale analytics and distributed workflows, with dashboards, alerts, and post-mortems to ensure reliability and correctness across the system.

Qualifications

  • Four or more years of production experience in backend, ML, or audio systems.
  • Depth in at least one: speech/audio processing, LLM evaluation pipelines, or high-volume data infrastructure.
  • Strong commitment to correctness, idempotency, and ground-truth labels.
  • Willingness to own operations including dashboards and alerts.

Responsibilities

  • Design and operate evaluation pipelines that score large volumes of production calls with low latency.
  • Build audio-native metrics and validate them against ground truth alongside LLM evaluators.
  • Extend a simulation engine to dial real agents across phone and WebRTC with varied personas.
  • Own analytics infrastructure including columnar stores, streams, and dashboards.
  • Orchestrate long-running workflows with a durable engine and ensure idempotent behavior.
  • Operate dashboards, alerts, and perform root-cause analysis when metrics drift.

Skills

Backend development
ML systems
Audio processing
Python/TypeScript
Data infrastructure
Operational ownership
Technical writing

Tools

TypeScript
Bun
AWS (SST)
ClickHouse
DynamoDB
Temporal
Postgres
WebRTC

Job description

Role overview

A founding-level engineering position focused on building the evaluation and measurement infrastructure for an AI voice calling platform. The work centers on scoring production calls in near real time, layering audio-native signals (pronunciation, emotion, vocal stress) on top of LLM-based evaluators, and powering a simulation engine that places realistic automated calls. Getting this measurement layer right is treated as foundational to everything else in the product.

Responsibilities
  • Design and operate evaluation pipelines that score large volumes of production calls at low latency.
  • Build audio-native metrics that go beyond transcripts, validating them against human-labeled ground truth alongside LLM-based evaluators.
  • Extend a simulation engine that dials real agents across phone and WebRTC, modeling varied personas, accents, and background noise.
  • Own large-scale analytics infrastructure, including columnar data stores, key-value streams, and the query layer feeding internal dashboards.
  • Orchestrate long-running, distributed workflows (replays, backfills, load tests) using a durable workflow engine, with strong attention to idempotency and correctness.
  • Operate what you build: dashboards, alerts, and root-cause investigation when production metrics drift.
Requirements
  • Four or more years of production experience building backend, ML, or audio systems, ideally in TypeScript or Python.
  • Hands-on depth in at least one of: speech/audio processing, LLM evaluation pipelines, or high-volume data infrastructure.
  • Strong rigor around correctness, with comfort thinking in terms of idempotency, replay safety, and ground-truth labels rather than only happy paths.
  • Willingness to own operations, including dashboards, alerts, and after-hours investigation.
  • Clear written communication, with design documents and post-mortems people actually want to read.
Nice to have
  • Familiarity with the listed stack (TypeScript, Bun, AWS via SST, ClickHouse, DynamoDB, Temporal, Postgres, and WebRTC/telephony integrations).
  • Experience validating ML systems against human judgment at scale.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Engineer, Voice and Realtime
AI Engineer, Voice and Realtime

Obble • United States

Remote
USD 120,000 - 180,000
Founding AI Voice Systems Engineer: Real-Time Evaluation
Founding AI Voice Systems Engineer: Real-Time Evaluation

Roark • United States

Remote
USD 150,000 - 230,000
Senior / Staff Backend Engineer
Senior / Staff Backend Engineer

Hamming • Austin (TX)

On-site
USD 120,000 - 160,000
Flexible work hours
Career development opportunities
Senior Backend Engineer, AI Evaluations
Senior Backend Engineer, AI Evaluations

Ellipsis Health • San Francisco (CA)

On-site
USD 160,000 - 210,000
401(k) matching
Health insurance
Flexible PTO
Founding AI Engineer
Founding AI Engineer

Aimhire • San Francisco (CA)

On-site
USD 150,000 - 230,000
Senior Software Engineer
Senior Software Engineer

AIM AI • Los Angeles (CA)

On-site
USD 130,000 - 170,000
ML Engineer — Real-time Speech
ML Engineer — Real-time Speech

Sellsig • Minnesota

On-site
USD 100,000 - 130,000
ML Engineer
ML Engineer

Catalyst Labs • New York (NY)

On-site
USD 120,000 - 140,000
Competitive compensation
Bonus opportunities
Equity in the company
Software Test Engineer at Deepgram
Software Test Engineer at Deepgram

Matcha • Northern (KY)

On-site
USD 110,000 - 160,000
Software Test Engineer
Software Test Engineer

Deepgram • Northern (KY)

On-site
USD 120,000 - 180,000