AI Engineer

muffintech

Berlin

Hybrid

EUR 90.000 - 140.000

Vollzeit

Vor 9 Tagen

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

muffintech is hiring an AI Engineer to build end-to-end AI capabilities at the heart of our product. You will own features from domain context to production, including retrieval, agent orchestration, and backend services.

Expect to define evaluation strategies and tackle failure modes of language models, while prioritizing latency and cost as product constraints. You will work with Python across AI and backend layers, employing retrieval-augmented generation and agent workflows over

Qualifikationen

  • Experience delivering LLM-based features in production environments.

Aufgaben

  • Own AI features end to end from requirement to production measurements.
  • Build and evolve retrieval and agent systems, ingestion, chunking, and orchestration.
  • Own the backend APIs, data ingestion, and pipelines for features.
  • Define how quality is measured with eval sets and production metrics.
  • Design for language model failure modes and robust system boundaries.
  • Collaborate with product to define what is

Kenntnisse

Python
Backend engineering
LLM production
Retrieval systems
Evaluations
Language model failure modes
Technical judgment
Communication with product
AI tooling

Tools

REST
WebSocket
Streaming

Jobbeschreibung

Mid to Senior · Full-time · Hybrid/Remote

About the role

muffintech is at the forefront of AI-driven solutions for the insurance industry, empowering companies to seamlessly integrate cutting-edge technologies into their operations. We focus on vertical scalability – automating a broad range of insurance processes through a unified, AI-first platform rather than just tackling a single task.

We are hiring an AI Engineer to build the AI capabilities at the heart of our product. This is a high-ownership, product-minded role: you will join the engineers already working on our AI layer and take your own features end to end — from retrieval and agent logic through to the backend services and pipelines behind them. Product brings the problem and the domain context; you decide how it gets solved, with tech leadership available on major architectural calls.

AI features differ from the rest of the product in one important way: a passing test suite does not tell you whether they work. Much of this role is deciding how a capability should be measured, building the evaluation that answers it, and telling a real improvement from a plausible-looking one. Fluent use of modern AI-assisted tooling in your own work is expected.

What you’ll do

  • Own AI features end to end — from a product requirement and its domain context through to a measured capability running in production that you keep healthy over time.
  • Build and evolve our retrieval and agent systems — ingestion and chunking, retrieval quality, tool design, and the orchestration that holds them together.
  • Own the backend your features depend on, including the APIs, data ingestion, and pipelines that feed and serve them.
  • Define how quality is measured: build the eval sets and production measurements that show whether a change improved the product, and use them before you ship.
  • Design for the failure modes language models bring — unsupported answers, prompt injection through retrieved content, partial tool failures, and destructive actions against live systems.
  • Work closely with product — establishing what “correct” means for a feature, surfacing edge cases early, and being straight about what the system can be trusted to do.
  • Treat latency and cost as product constraints alongside quality.
  • Python across the AI and backend layer.
  • Retrieval-augmented generation and agentic workflows over insurance domain documents.
  • Language models from several providers — called directly, with orchestration frameworks where they earn their place.
  • REST and WebSocket for network communication, including streaming responses.
  • Evaluation and observability tooling for model behaviour in production.

What we’re looking for

  • Demonstrable experience taking LLM-based features into production — not prototypes or notebooks.
  • Solid Python and backend engineering: you can design, build, and run the services, APIs, and data pipelines your features depend on.
  • Real depth in retrieval: you understand how retrieval systems fail and can diagnose one from evidence rather than by trial and error.
  • Experience designing and running evaluations for non-deterministic systems, and the judgement to recognise when a measured improvement is real.
  • A clear-eyed view of language model failure modes, and the instinct to design around them at the system boundary rather than prompt around them.
  • Sound technical judgment: you can weigh quality against latency and cost and set conventions.
  • Proactive, clear communication with product about what the system can reliably deliver.
  • Fluency with modern AI-assisted development tooling as part of your everyday workflow.

Nice to have

  • Experience in insurance, or another regulated domain where being wrong carries real cost.
  • Experience with tool-calling agents against production systems, and the guardrails that makes necessary.

Who this role suits

This role suits someone comfortable owning a problem whose right answer is not knowable up front and has to be established by measurement. You will join engineers already working on the AI layer, but will be trusted with your own area and expected to set its direction.

How we work

Location requirement

This position must be worked from within Germany. Because we handle sensitive data, the data-protection and regulatory commitments we make to our clients place requirements on where that data is accessed and processed. We are therefore unable to accept applicants who would be working from outside Germany. Remote work is fully supported — your place of work simply needs to be in Germany.

  • Arrangement: Hybrid, or fully remote both possible
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

AI Engineer
AI Engineer

Remotely • Berlin

Hybrid
EUR 90.000 - 135.000
AI-Native Senior Software Engineer (Customer-Facing)
AI-Native Senior Software Engineer (Customer-Facing)

GoHiring GmbH • Berlin

Remote
EUR 90.000 - 130.000
AI Engineer (m/f/d)
AI Engineer (m/f/d)

Vectoryx USA Inc. • Berlin

Vor Ort
EUR 90.000 - 130.000
Applied AI Engineer
Applied AI Engineer

TechShack • München

Vor Ort
EUR 70.000 - 110.000
Visa sponsorship
Relocation support
Learning budget
+1
Founding AI Engineer Onsite (Munich, Germany)
Founding AI Engineer Onsite (Munich, Germany)

S27a • München

Vor Ort
EUR 95.000 - 150.000
AI Go-to-Market Insurance DACH Lead (w/m/x)
AI Go-to-Market Insurance DACH Lead (w/m/x)

NTT DATA Europe & Latam • Rosenheim

Hybrid
EUR 180.000 - 240.000
Flexitime with working-time rules
30 days annual leave + special days
Sabbatical option up to 1 year
+6
AI Go-to-Market Insurance DACH Lead (w/m/x)
AI Go-to-Market Insurance DACH Lead (w/m/x)

NTT DATA Europe & Latam • Frankfurt

Hybrid
EUR 150.000 - 210.000
Flexible working hours
30 days annual leave
Sabbatical up to one year
+3
AI Go-to-Market Insurance DACH Lead (w/m/x)
AI Go-to-Market Insurance DACH Lead (w/m/x)

NTT DATA Europe & Latam • München

Hybrid
EUR 180.000 - 240.000
Flexible hours
30 days annual leave
Remote work from EU
+5
AI Go-to-Market Insurance DACH Lead (w/m/x)
AI Go-to-Market Insurance DACH Lead (w/m/x)

NTT DATA Europe & Latam • Berlin

Hybrid
EUR 120.000 - 180.000
Flexible working hours
30 days annual leave
Sabbatical option
+4
AI Go-to-Market Insurance DACH Lead (w/m/x)
AI Go-to-Market Insurance DACH Lead (w/m/x)

NTT DATA Europe & Latam • Köln

Vor Ort
EUR 150.000 - 210.000
Flexible working hours
30 days annual leave + holidays
Remote work from EU countries
+1