Software Engineer, Agent Intelligence

Sage Care

Palo Alto (CA)

On-site

USD 120,000 - 150,000

Full time

14 days+
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flexible working hours
Health benefits

Job summary

A forward-thinking technology company is seeking a skilled backend engineer to build and improve diagnostic, observability, and root cause analysis infrastructure. This role involves creating automated systems to trace events across various technologies and to monitor real-time communication for clinical AI assistants. The ideal candidate will possess strong backend skills in Python and experience with event-driven architectures, while also being familiar with tools and frameworks that enhance observability and diagnostics.

Qualifications

  • Strong backend engineer experienced with diagnostics, observability, and event-driven tracing.
  • Expert in Python, logging systems, real-time pipelines, and distributed debugging.
  • Deep familiarity with LLM agents and telemetry or tracing frameworks.

Responsibilities

  • Build automated RCA pipelines to detect and classify failure modes.
  • Implement event tracing infrastructure capturing every agentic decision.
  • Architect a live SOP state-machine tracer with real-time transcript overlays.

Skills

Backend engineering
Event-driven tracing
Python
Logging systems
Real-time pipelines
Distributed debugging

Tools

Grafana
OpenTelemetry
Twilio
D3.js
React

Job description

Role Overview

Own and build the full diagnostic, observability, and RCA infrastructure that makes Sage Care’s AI assistant trustworthy and debuggable— in real time and post‑call. This engineer builds the visibility layer across telephony, transcription, reasoning, SOP traversal, and tool‑calling; creates dashboards for both engineers and live human supervisors; and implements automated triage + notification pipelines that surface issues to the right module owners immediately.

This role sits at the intersection of LLM orchestration, voice pipelines, transcription, SOP engines, and operations, serving as the connective tissue across the stack. Your work enables rapid root‑cause analysis, real‑time intervention, and continuous improvement of our clinical AI assistants.

Key Responsibilities
Root Cause Analysis, Tracing & Observability
  • Build automated RCA pipelines to detect and classify failure modes:
    • Hallucinations
    • Misrouted intents
    • Leaked/invalid tool calls (Transfer, SayMessage, Hangup, NOOP)
    • Unrecoverable SOP loops
    • Broken state transitions
    • Telephony dropouts / DTMF issues
  • Implement event tracing infrastructure capturing every agentic decision across LLM, telephony, and SOP execution.
  • Compare expected vs. actual SOP behavior using protocol‑driven expectations or human‑labeled ground truth.
  • Automatically compute performance, safety, reliability, and coverage metrics.
Diagnostic Dashboards & Visualization
  • Build live and post‑call dashboards that visualize:
    • Full call timeline
    • SOP/state machine traversal
    • Agent reasoning traces
    • Tool invocation history
    • Divergence from expected behavior
  • Design interactive visualizations: heatmaps, decision‑path overlays, branching comparisons, and error hotspots.
  • Build triage dashboards for engineering and operations teams to rapidly understand system health.
Integration with Core AI Modules
  • Voice + Telephony Integration
    • Trace call‑level events (dropouts, retries, audio playback issues).
    • Detect DTMF misfires and incorrect action routing.
  • Transcriber Module Integration
    • Analyze turn segmentation, word‑error‑rate drift, boosting performance, and latency.
    • Visualize errors in context (audio, transcript, aligned timecodes).
  • LLM Orchestration Integration
    • Audit intent classification accuracy and subgraph routing.
    • Trace reasoning sequences, missing tool calls, redundant tool calls, or invalid arguments.
    • Validate tool call correctness (maps, SMS, search, internal SOP tools).
Live Monitoring & Human‑in‑the‑Loop
  • Architect a live SOP state‑machine tracer with:
    • Real‑time transcript overlays
    • Current state + next expected state
    • Deviation alerts
  • Build dashboards to monitor 10–15 concurrent calls, highlighting sessions with:
    • Loops
    • Latency spikes
    • Failed tool calls
    • Repeated incorrect decisions
  • Provide human specialists with escalation alerts and context.
Command & Control Interface

Build an intervention console for on‑call specialists, enabling:

  • “Skip step”
  • “Say apology”
  • “Escalate to human”
  • “Send SMS”
  • “Repeat last message”
  • Override of SOP steps while maintaining auditability and continuity.

This system must blend seamlessly into existing agent workflows without breaking call integrity.

Failure Classification, Clustering & Pattern Detection
  • Build clustering systems (via embeddings or metadata) to group systemic failure modes:
    • Intent misroutes under noisy audio
    • Repeated missing tool calls
    • Looped state machine traversal
    • Hallucinated follow‑ups or invalid summaries
  • Generate recurring‑failure reports to guide engineering improvements.
Auto‑Triaging & Notification System (NEW)
  • Design and implement an automated triage and notification system that:
    • Detects failure category and severity in real time.
    • Routes incidents to the correct module owners:
      • Telephony
      • Transcription
      • LLM orchestration
      • SOP/decision‑tree team
      • Platform reliability
    • Sends structured payloads containing:
      • Trace graphs
      • Relevant logs
      • Transcript segments
      • SOP divergence snapshots
      • Suggested RCA labels
Notifications May Integrate With
  • PagerDuty
  • Slack (rich message blocks)
  • Jira auto‑ticket creation
  • Internal incident pipelines

This ensures rapid operational feedback loops and reduces time‑to‑resolution.

Post‑Call RCA Pipelines & Analytics
  • Extend pipelines to automatically generate human‑readable failure summaries with:
    • Call‑level trace graphs
    • Tool call sequences
    • Transcript context
    • Classified failure types
    • Suggested root causes
  • Store snapshots for operational handoff and debugging.
Required Qualifications
  • Strong backend engineer experienced with diagnostics, observability, and event‑driven tracing.
  • Expert in Python, logging systems, real‑time pipelines, and distributed debugging.
  • Deep familiarity with:
    • LLM agents
    • LangGraph or state‑machine frameworks
    • Tool‑calling architectures
    • Telemetry or tracing frameworks
  • Comfortable designing both:
    • Backend data pipelines
    • Frontend dashboards in React, D3, WebSockets, or equivalent.
Preferred Qualifications
  • Telephony/Voice: SIP, WebRTC, Twilio, audio streaming pipelines.
  • Clinical operations, call‑center workflows, or mission‑critical HITL supervision systems.
  • Observability stacks (Grafana, ELK, OpenTelemetry, Sentry).
  • Clustering/ML techniques for failure pattern discovery.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Real-Time AI Diagnostics & Observability Engineer
Real-Time AI Diagnostics & Observability Engineer

Sage Care • Palo Alto (CA)

On-site
USD 120,000 - 150,000
Software Engineer, Agent Platform
Software Engineer, Agent Platform

Sage Care • Palo Alto (CA)

On-site
USD 130,000 - 160,000
Senior Applied AI Engineer - Life Sciences & Healthcare
Senior Applied AI Engineer - Life Sciences & Healthcare

Vi • Austin (TX)

On-site
USD 120,000 - 150,000
AI Engineer – Enterprise AI Platform
AI Engineer – Enterprise AI Platform

Jobtailor • Chicago (IL)

On-site
USD 140,000 - 180,000
Sr Software Engineer, AI Agent Platform
Sr Software Engineer, AI Agent Platform

Motorola Solutions • Salt Lake City (UT)

On-site
USD 140,000 - 190,000
Senior / Staff Backend Engineer
Senior / Staff Backend Engineer

Hamming • Austin (TX)

Hybrid
USD 120,000 - 160,000
Flexible work hours
Career development opportunities
Implementation Engineer
Implementation Engineer

Storm3 • New York (NY)

On-site
USD 100,000 - 130,000
AI Prompt & Agent Developer
AI Prompt & Agent Developer

OpenCall.ai (YC W24) • New York (NY)

On-site
USD 90,000 - 130,000
Senior AI Engineer (3 roles)
Senior AI Engineer (3 roles)

STARLIMS • United States

On-site
USD 150,000 - 190,000
Senior AI Engineer
Senior AI Engineer

Apt • Dallas (TX)

On-site
USD 120,000 - 190,000