Una candidatura completa en un minuto: currículum y carta de presentación adaptados, listos para enviar.
Value Crew seeks a Lead Data & AI Engineer to define how an AI system should operate in production. You’ll be the senior technical reference for a small team, delivering 60–70% hands-on engineering while steering architecture with Product, Security and Data.
You’ll own end-to-end AI system design, tooling, evaluation and observability, shaping deterministic software and cost-aware decisions for scalable production workflows.
We aren't looking for someone to add an LLM API to an existing product.
We're looking for the engineer who can decide how the AI system should actually work in production.
Our client is a fast-growing European B2B software company helping distributed organisations automate high-volume operational and compliance-heavy workflows.
Its platform combines enterprise data, documents, external systems and AI agents to automate work that historically required people to move information between emails, PDFs, legacy applications and back-office systems.
The company has passed the experimental stage. Customers already use AI-powered workflows in production.
The challenge now is reliability and scale.
How do you make agents predictable? How do you measure them? What belongs in the prompt, retrieval layer or deterministic software? When should a model be allowed to take an action? How do you manage permissions, cost, latency and fallbacks?
Those are the problems you'll own.
As Lead Data & AI Engineer, you'll be the senior technical reference for a small group of Data and AI Engineers.
This is a hands-on technical leadership role, not a people-management job disguised as engineering.
Expect approximately 60-70% hands-on engineering, with the rest spent on architecture, technical direction, mentoring and working with Product, Security and Data.
Define the architecture of production AI systems built around LLMs, retrieval and agentic workflows.
Design multi-step agents, tool use, function calling, memory and state-management patterns.
Decide where deterministic software should replace probabilistic model behaviour.
Build reliable RAG systems covering ingestion, chunking, embeddings, retrieval, re-ranking and grounding.
Design model abstraction and routing so the product isn't unnecessarily locked into one provider.
Introduce systematic AI evaluation: offline datasets, regression suites, production metrics and experimentation.
Build observability covering traces, prompts, retrieval results, model responses, latency, cost and failures.
Implement guardrails, structured output validation, retries, fallbacks and human review for higher-risk actions.
Design AI-ready data architectures and interfaces between enterprise data and AI services.
Own LLMOps/MLOps patterns for versioning, testing, rollout and rollback.
Partner with Security and Legal on access control, privacy, responsible AI and EU regulatory requirements.
Review architecture and code while still shipping meaningful production code yourself.
Mentor engineers and help define technical standards as the team grows.
Translate business problems into engineering decisions rather than treating “AI” as the requirement.
The exact stack evolves, but today you'll encounter:
Core: Python, FastAPI, SQL, Git
AI: Azure OpenAI, Azure AI Foundry, LangGraph / Semantic Kernel
Retrieval: embeddings, vector/hybrid search, re-ranking
Data & ML: Databricks, MLflow
Platform: containers, Kubernetes, Terraform, CI/CD
Observability: distributed tracing, AI eval tooling, production monitoring
Experience with equivalent technologies is absolutely valid.
We care more about whether you understand the underlying engineering problem than whether you've used one particular framework.
7+ years in software, data or ML engineering, with meaningful production ownership.
Previous technical leadership or architecture responsibility.
Strong Python/software-engineering fundamentals.
Hands‑on experience taking an LLM, RAG or agent‑based system beyond prototype stage and into production.
Deep understanding of the failure modes of LLM applications.
Experience with cloud architecture, APIs and distributed systems.
A strong evaluation mindset: “it seemed to work in testing” isn't enough.
Ability to reason across data, ML and software architecture rather than treating them as separate worlds.
Experience making technical decisions involving security, privacy, reliability and cost.
Ability to explain complex systems clearly to engineers, Product and business stakeholders.
Enterprise search, MCP, model fine-tuning, responsible AI, Azure AI, MLflow, Databricks, event-driven architectures or experience building software in regulated domains.
After your first months, the team should be able to answer questions it currently can't answer reliably:
Why did the agent make that decision?
Which retrieval change improved quality?
Did the new model actually outperform the old one?
What does one successful workflow cost?
Which AI failures reach customers?
Can we roll this version back safely?
And, crucially: should this particular problem use an LLM at all?
€75k-90k base salary depending on scope and experience.
Annual performance bonus and/or equity package.
Hybrid setup: around two office days per week in Barcelona or Madrid.
High-end equipment and home-office support.
Direct influence over the AI architecture rather than inheriting decisions made several layers above you.
1. Talent / role conversation: 45 min
2. AI systems deep dive: 75 minYou'll discuss a system you've built and design a production AI architecture with one of our senior engineers.
3. Leadership & Product conversation: 60 minTrade-offs, influence, technical leadership and business thinking.
4. References and offer.
No five-round interview marathon and no take-home project that consumes your weekend.
We hire this profile on an ongoing basis. Depending on current team needs, the process may lead to an immediate opportunity or to our priority talent network.