An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Salla is seeking a senior engineer to own the evaluation stack for its production multi-agent systems. You will design LLM-as-judge components, calibrate against human labels, and quantify agent quality per failure mode.
You will also build simulators, integrate regression detection into CI, and contribute to agent development by turning failures into improvements. This role emphasizes strong software engineering and production readiness.
We run production multi-agent systems that handle real work for a large base of users. As those systems grow, our biggest constraint is confidence: we need to know how well the agents perform, catch regressions before they ship, and keep quality steady as we release. This role owns that.
You’ll build the evaluation systems behind our agents – the judges, test harnesses, and simulators that tell us whether an agent is working and where it’s failing. The goal is to let us ship agents faster because we can trust what the evaluation tells us.
Evaluation is the focus, but it won’t be the boundary. Because you’ll understand the agents’ failure modes better than anyone there will also be opportunities to contribute to agent development itself, building and improving the agents alongside the systems that evaluate them.