Una candidatura completa in un minuto — curriculum e lettera di presentazione personalizzati, pronti da inviare.
Clera in San Francisco, CA is seeking an AI/ML engineer to own the intelligence behind failure detection in production agents and verify fixes before deployment.
You'll focus on evaluating when a proposed patch works, turning user signals into evidence, and improving evaluation methods in real-world traces. The role emphasizes reliability, learning from failures, and building trustworthy systems.
Join an AI/ML team focused on making production agents more reliable by detecting failures, evaluating potential fixes, and helping teams build systems they can trust. This role owns the intelligence behind what gets flagged, how confidently it is identified, and whether a proposed fix works.
Detect subtle agent failures, including incorrect responses that appear compliant, omissions, and patterns that emerge across many traces.
Turn indirect user signals, such as rephrasing, abandonment, and retries, into evidence that an agent has failed.
Build evaluations that help measure detection quality and improve performance even when labeled ground truth is unavailable.
Make patch generation trustworthy by reproducing failures, verifying fixes, and avoiding low-confidence changes.
Improve the cost and quality tradeoffs of model-based evaluation, including when to use a smaller model or no model.
Review real production traces regularly to identify problems and guide improvements.
Salary range is $120,000 to $200,000 USD annually. Visa sponsorship is not available.
This is an on‑site role based in San Francisco, United States.