Описание:
Adyen provides payments, data, and financial products in a single solution for customers such as Meta, Uber, H&M, and Microsoft. Its Payment Solutions team processes billions of transactions and helps businesses create seamless payment experiences through a global financial technology platform. *Instagram и Facebook принадлежат компании Meta Platforms Inc., деятельность которой признана экстремистской и запрещена на территории РФ
Задачи:
- Discover critical challenges across product and engineering teams and rapidly build AI prototypes to demonstrate value
- Own the end-to-end development of bespoke AI tools for merchant experience, pricing models, and internal workflows
- Define and lead evaluation strategies for agentic systems and LLMs
- Design internal benchmarks covering domain complexity, edge cases, capabilities, and failure modes
- Build reusable evaluation infrastructure embedded in the development process
- Provide technical expertise on agentic frameworks, retrieval and search strategies, and agent tool-use approaches across partner teams
- Identify connections across AI initiatives and help teams avoid duplicated work or incorrect approaches
- Set engineering standards for the team and company
- Mentor through problem decomposition, research methodology, and code review
- Promote reproducibility, documentation, and rigorous evaluation practices across the AI organization
Требования:
- 7+ Years of hands-on experience in applied AI/ML research or engineering
- A proven track record of shipping AI systems, including agentic or LLM-powered systems, in production
- Deep expertise in language models and Generative AI
- Hands-on experience with architecture, post-training, inference optimization, context engineering, and failure modes at scale
- Experience designing and operating agentic systems at scale, including multi-agent orchestration, tool use, memory and context management, state handling for long-running workflows, and human-in-the-loop design
- Experience designing evaluation frameworks or internal benchmarks beyond standard metrics
- Understanding of LLM-as-judge failure modes and meaningful system evaluation
- Strong foundation in supervised learning, ensemble methods, optimization, probabilistic modeling, and statistics
- Ability to write clean, well-structured, production-ready Python code
- Hands-on experience with at least one production-grade agentic framework
- Nice to have: Familiarity with financial data, payments, fraud detection, or risk systems, publications, conference presentations, open-source contributions, observability and evaluation tooling, MLOps, model deployment pipelines in large-scale environments.
Условия:
The role is based out of the Amsterdam office; The company is office-first and values in-person collaboration; Remote-only work is not offered.