Principal Applied Scientist

Xcede Recruitment Solutions

Berlin

Vor Ort

EUR 65.000 - 90.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Xcede Recruitment Solutions is seeking an Applied Scientist with expertise in LLM benchmarking, based in Berlin or Munich, with some remote work possible within Germany. The candidate will enhance prompt optimisation, evaluate models against performance metrics, and develop an evaluation framework.

Ideal applicants should have strong Python skills, hands-on experience with LLMs, and a solid understanding of benchmarking methodologies. This role involves collaboration with engineering teams to enhance production outcomes.

Qualifikationen

  • Strong Python programming skills are essential.
  • Experience with LLMs and prompt engineering is required.
  • Understanding of benchmarking and evaluation methodologies is necessary.
  • Systems thinking to build tooling rather than running one-off notebooks is crucial.

Aufgaben

  • Improve and optimise prompts for production use cases.
  • Build and maintain benchmarking scenarios for model evaluation.
  • Assess new model releases and provide data-driven recommendations.
  • Design a generalised evaluation framework for diverse models.
  • Collaborate with engineering teams for production changes.
  • Track quality and latency trends of models over time.

Kenntnisse

Strong Python
Hands-on experience with LLMs
Benchmarking and evaluation methodology
Systems thinking

Jobbeschreibung

Applied Scientist | LLM Benchmarking | Berlin / Munich / Remote in Germany

Confidential search for a fast-growing, Series C agentic AI company building conversational AI for global enterprise brands.

This isn't a research seat and it isn't a data analytics role. You'd own prompt optimisation and LLM benchmarking end to end: building the evaluation framework that decides which models the company adopts, comparing quality against latency across real production use cases, and generalising that framework so it can eventually judge any model, not just the ones already wired into the product. Six months from now, the ambition is for this to be a benchmark other companies reference.

What you'll do
  • Improve and optimise prompts for real production use cases and agent workflows, systematically evaluating performance across different prompting strategies and models.
  • Build and maintain benchmarking scenarios that evaluate models (Gemini, GPT-class systems, and whatever comes next) against task performance, end-to-end system integration, latency, and quality.
  • Assess new model releases as they land, validate them against our performance and latency requirements, and give data‑driven recommendations on whether we adopt them.
  • Design and evolve a generalised evaluation framework, starting from our current benchmarking tooling and gradually decoupling it from our core agent system so it can assess any model, including ones we haven't integrated yet.
  • Work closely with our agent and platform engineering teams to turn findings into production changes.
  • Track quality and latency trends across model versions over time.
What you'll need
  • Strong Python
  • Hands‑on experience with LLMs and prompt engineering
  • A real understanding of benchmarking and evaluation methodology
  • Systems thinking – you'll be building tooling, not running one‑off notebooks

Research background is a plus, so is experience comparing models at scale. Not a fit if you’re a pure analyst or a theoretical researcher who hasn’t shipped to production.

Berlin or Munich preferred, remote within Germany considered.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

German Senior Prompt Engineer: LLM Migration & Optimization
German Senior Prompt Engineer: LLM Migration & Optimization

Welo Data • Deutschland

Remote
EUR 60.000 - 80.000
Senior Prompt Engineer: LLM Migration & Optimization
Senior Prompt Engineer: LLM Migration & Optimization

Welo Data • Deutschland

Remote
EUR 42.000 - 69.000
French Senior Prompt Engineer: LLM Migration & Optimization
French Senior Prompt Engineer: LLM Migration & Optimization

Welo Data • Deutschland

Remote
EUR 83.000 - 165.000
Applied AI Engineer
Applied AI Engineer

TechShack • München

Vor Ort
EUR 70.000 - 110.000
Visa sponsorship
Relocation support
Learning budget
+1
Senior AI Engineer
Senior AI Engineer

Bluefish • Berlin

Hybrid
EUR 70.000 - 90.000
Applied Scientist - LLM, Alexa Conversational Modelling Intelligence
Applied Scientist - LLM, Alexa Conversational Modelling Intelligence

Amazon • Berlin

Vor Ort
EUR 110.000 - 170.000
Italian Senior Prompt Engineer: LLM Migration & Optimization
Italian Senior Prompt Engineer: LLM Migration & Optimization

Welo Data • Deutschland

Remote
EUR 60.000 - 100.000
Flexible hours
Remote work (100% )
Freelance/Independent contractor
Japanese Senior Prompt Engineer: LLM Migration & Optimization
Japanese Senior Prompt Engineer: LLM Migration & Optimization

Welo Data • Deutschland

Remote
EUR 48.000 - 73.000
Remote work
Flexible schedule
Competitive freelance rates
Prompt Engineer: LLM Migration & Optimization
Prompt Engineer: LLM Migration & Optimization

Welo Data • Deutschland

Remote
EUR 68.000 - 103.000
Portuguese (Portugal) Senior Prompt Engineer: LLM Migration & Optimization
Portuguese (Portugal) Senior Prompt Engineer: LLM Migration & Optimization

Welo Data • Deutschland

Remote
EUR 83.000 - 131.000