An application made for this job — a tailored resume and cover letter that speak straight to the posting.
Brilliant Systems is seeking an engineer to own the evaluation of ML/LLM systems, turning client requirements into measurable evaluation sets and robust tooling. You will build harnesses, monitor production drift, and ensure evaluation is integral to the product lifecycle.
You must demonstrate statistical literacy, strong Python skills, and a knack for developer-friendly tooling, while candidly reporting results that may be inconvenient.
Decide what done means. If we cannot score it, we have not specified it, and we certainly cannot sell it.
We sell fixed-price phases, so we need a defensible answer to whether a thing is finished. For deterministic software that is a test suite. For a system that is probabilistic by design it is an evaluation harness, and building good ones is a speciality. This role owns that speciality across client engagements and our own products.