AI Engineer (LLM/RAG) (m/w/d)

Meyandy LLC

Köln

Vor Ort

EUR 85.000 - 120.000

Vollzeit

Vor 2 Tagen
Sei unter den ersten Bewerbenden

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

Nejo in Cologne seeks an experienced AI Engineer to own live LLM pipelines and RAG systems. You will improve reliability and extend use cases, working primarily with TypeScript and Next.js. The role does not involve training or fine-tuning foundation models.

The stack includes TypeScript, Next.js, Node.js, Postgres with pgvector, Docker, Azure, Anthropic and OpenAI APIs, and Vercel AI SDK. You will ensure stability, observability, and cost-efficiency across production systems.

Qualifikationen

  • Professional experience with at least one LLM-based system in production use, including responsibility for its operation and incident handling.
  • Advanced TypeScript/Node.js skills: strict typing of non-deterministic model output, async/concurrency, streaming, structured error handling.
  • Production Next.js experience: App Router, route handlers, server actions, client streaming.

Aufgaben

  • Take over and maintain the existing LLM pipelines, assess architecture, identify failure modes, fix and refactor without disrupting production.
  • Own the RAG systems end to end: document ingestion, chunking, indexing, hybrid retrieval, query rewriting, reranking, grounded generation with citations.
  • Implement and maintain chunk-level access control, index freshness, tenant isolation across retrieval systems.
  • Develop content generation pipelines delivering consistent quality at volume, with human review steps.
  • Build and operate automated workflows against internal/third-party business systems with durable, idempotent execution, retries, dead-letter handling, and approval steps.
  • Establish an evaluation framework with golden datasets from production failures and retrieval metrics.
  • Implement observability across the full request path.
  • Optimise cost and latency via prompt caching, batching, model routing, and smaller models where appropriate.
  • Assess where deterministic logic is preferred and implement accordingly.
  • Collaborate with non-technical colleagues to specify and validate automated processes.

Kenntnisse

TypeScript
Node.js
Next.js
Postgres
pgvector
Docker
Azure
Anthropic/OpenAI APIs
Vercel AI SDK
LLM/RAG
Observability
Cost/Latency Optimization

Tools

JSON Schema
Zod
LangChain/Agent SDKs
OpenTelemetry

Jobbeschreibung

For a company in the fast-growing AI implementation market we are looking for an experienced AI Engineer, starting immediately. The company operates LLM-based systems in production: content generation pipelines, retrieval-augmented generation (RAG) over internal documents, and automated workflows deeply integrated with their business systems. The stack is TypeScript and Next.js end to end. These systems are already live. As AI Engineer you take ownership of them, improve their reliability and quality, and extend them to new use cases. The role combines applied LLM engineering with solid backend engineering in TypeScript. It does not involve training or fine-tuning foundation models.

Stack: TypeScript, Next.js, Node.js, Postgres with pgvector, Docker, Azure, Anthropic and OpenAI APIs, Vercel AI SDK.

Tasks
  • Take over and maintain the existing LLM pipelines: assess the current architecture, identify failure modes, prioritise fixes, and refactor and extend without disrupting production
  • Own the RAG systems end to end: document ingestion and parsing, chunking, indexing, hybrid retrieval (BM25 and vector), query rewriting, reranking, grounded generation with citations
  • Implement and maintain chunk-level access control, index freshness and tenant isolation across retrieval systems
  • Develop content generation pipelines that deliver consistent quality at volume, including human review steps
  • Build and operate automated workflows against internal and third-party business systems (ERP, CRM, email, internal APIs), with durable and idempotent execution, retry and dead-letter handling, and approval steps for irreversible actions
  • Establish an evaluation framework for systems currently running without one: golden datasets derived from observed production failures, retrieval metrics and more
  • Implement observability across the full request path
  • Optimise cost and latency through prompt caching, batching, model routing and use of smaller models where appropriate
  • Assess where deterministic logic is the better solution and implement it accordingly
  • Work directly with non-technical colleagues to specify and validate automated processes
Requirements
  • Professional experience with at least one LLM-based system in production use, including responsibility for its operation and incident handling
  • TypeScript and Node.js at an advanced level: strict typing of non-deterministic model output, async and concurrency patterns, streaming responses, structured error handling
  • Next.js in production: App Router, route handlers, server actions, streaming to the client
  • Practical retrieval expertise: hybrid search, embedding model selection, cross-encoder reranking, metadata filtering, permission-aware retrieval, and structured diagnosis of poor retrieval quality
  • Experience processing real-world documents: PDFs with tables, scanned material, DOCX, HTML, including layout-aware parsing, OCR and evidence-based chunking
  • Structured outputs and tool calling as part of your everyday work: JSON Schema, Zod or comparable runtime validation, function calling, handling of malformed or partial output, context window management
  • Designed and run LLM evaluations
  • Experience with LLM tracing and evaluation tooling in a TypeScript codebase (e.g. Braintrust, Langfuse, Promptfoo, OpenTelemetry or Arize Phoenix)
  • Familiar with Postgres including vector search (pgvector or a comparable vector store), Docker, Git, CI/CD and one major cloud platform
  • Working experience with the Anthropic and/or OpenAI TypeScript SDKs
  • Confident communication in English, German is a plus
Plus
  • Durable workflow execution for long-running, unattended processes (Temporal, Inngest or comparable)
  • Agent orchestration in production, tool calling, recovery, multi-step workflows (Vercel AI SDK, LangGraph, Mastra, Claude Agent SDK, MCP TypeScript SDK)
  • Integration experience with enterprise systems such as ERP or CRM platforms like SAP
  • Security and data protection in LLM systems, prompt injection and data exfiltration defences, PII handling, GDPR-compliant design, EU-hosted or self-hosted inference structured or graph-base…

AI Engineer (LLM/RAG) (m/w/d) — Nejo, Cologne.

Mots-clés : Software Development.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Tech Lead AI Engineering (m/w/d)
Tech Lead AI Engineering (m/w/d)

Meyandy LLC • Köln

Vor Ort
EUR 90.000 - 150.000
AI Engineer (LLM/RAG) (m/w/d)
AI Engineer (LLM/RAG) (m/w/d)

Nejo • Köln

Hybrid
Vertraulich
Hybrid setup
Office in Cologne
Overtime compensation
+5
Generative AI/Agentic AI Engineer (d/f/m) [26084]
Generative AI/Agentic AI Engineer (d/f/m) [26084]

Soprasteria1 • Ulm

Vor Ort
EUR 70.000 - 120.000
Agentic AI Software Engineer
Agentic AI Software Engineer

Meyandy LLC • München

Hybrid
EUR 90.000 - 130.000
Deutschlandticket for public transit
Junior AI Application Engineer – AI Products (LLM & RAG) (m/f/d)
Junior AI Application Engineer – AI Products (LLM & RAG) (m/f/d)

Reply S.p.A. • Berlin

Hybrid
EUR 55.000 - 80.000
Mobility package
Gym subsidy
Corporate Savings Plan
+2
Senior AI Engineer
Senior AI Engineer

Meyandy LLC • Würzburg

Hybrid
EUR 90.000 - 140.000
Real ownership
Small team
Flexible hours
+2
AI Application Engineer – AI Products (LLM & RAG) (m/f/d)
AI Application Engineer – AI Products (LLM & RAG) (m/f/d)

Machine Learning Reply GmbH • Berlin

Hybrid
EUR 65.000 - 85.000
Mobility package
Gym subsidy
WellPass
+2
Senior Forward Deployed Engineer (m/f/d)
Senior Forward Deployed Engineer (m/f/d)

EPAM Systems • München

Hybrid
EUR 120.000 - 180.000
30 days holiday per annum
Company Pension Scheme
Regular performance assessments
+7
AI Application Engineer – AI Products (LLM & RAG) (m/f/d)
AI Application Engineer – AI Products (LLM & RAG) (m/f/d)

Reply S.p.A. • Berlin

Vor Ort
EUR 60.000 - 80.000
Mobility package
Gym subsidy & WellPass
Insurance & Pension Scheme
+2
Senior AI-Native Fullstack Engineer (m/w/d)
Senior AI-Native Fullstack Engineer (m/w/d)

Meyandy LLC • München

Hybrid
EUR 90.000 - 130.000
Work across AI-powered SaaS
Modern AI tooling daily
Room to contribute ideas