Erhalte mehr Antworten von Arbeitgebern
Versende in nur wenigen Minuten einen passgenauen Lebenslauf.
Xcede Recruitment Solutions is seeking an Applied Scientist with expertise in LLM benchmarking, based in Berlin or Munich, with some remote work possible within Germany. The candidate will enhance prompt optimisation, evaluate models against performance metrics, and develop an evaluation framework.
Ideal applicants should have strong Python skills, hands-on experience with LLMs, and a solid understanding of benchmarking methodologies. This role involves collaboration with engineering teams to enhance production outcomes.
Confidential search for a fast-growing, Series C agentic AI company building conversational AI for global enterprise brands.
This isn't a research seat and it isn't a data analytics role. You'd own prompt optimisation and LLM benchmarking end to end: building the evaluation framework that decides which models the company adopts, comparing quality against latency across real production use cases, and generalising that framework so it can eventually judge any model, not just the ones already wired into the product. Six months from now, the ambition is for this to be a benchmark other companies reference.
Research background is a plus, so is experience comparing models at scale. Not a fit if you’re a pure analyst or a theoretical researcher who hasn’t shipped to production.
Berlin or Munich preferred, remote within Germany considered.