Verschicke keinen generischen Lebenslauf — erstelle einen Lebenslauf und ein Anschreiben, die genau auf diese Rolle zugeschnitten sind.
Obsidian is seeking experts to advance frontier engineering-reasoning evaluations in collaboration with a leading AI research lab. You will design simulations and specifications, guiding models to produce artifacts that meet precise criteria.
The role requires strong domain knowledge in control systems or related fields, Python simulations expertise, and a knack for rigorous, measurable thresholds within a browser-based workflow.
A frontier engineering-reasoning evaluation run in collaboration with a leading AI research lab. The work measures whether state-of-the‑art models can reason from first principles in your engineering domain rather than retrieve facts from training data — and you get visibility into the model's internal reasoning on your own tasks.
Domain experts write simulations and specification sheets; the model attempts to design an artifact — controller gains, circuit parameters, geometry — that satisfies every spec.
You define the simulation, a set of specs with pass/fail thresholds, and the prompt. The model probes your simulation with a limited number of calls, then submits a final design. An agentic grader runs the simulation and scores spec satisfaction.
Onboarding calls run daily, alongside internal tooling built to help you work faster. Prior model-evaluation experience is not required.