No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.
XenonStack is seeking a Machine Learning Engineer, Model Evaluation to ensure large language models meet enterprise-grade accuracy, safety, and trustworthiness in real-world workflows. You will design evaluation pipelines, develop benchmarking tools, and run stress tests; you will also partner with ML engineers, product managers, and Responsible AI teams to align metrics with business goals.
This role offers opportunities to advance RLHF/RLAIF and maintain central repositories of test cases,
About Xenonstack XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.
About Xenonstack XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.
We Build Enterprise-grade Platforms Across The Agentic Stack
Our mission is to accelerate the world’s transition to AI + Human Intelligence by making agentic systems reliable, responsible, and enterprise-ready.
We are seeking an Machine Learning Engineer, Model Evaluation to ensure that large language models (LLMs) and agentic AI systems meet enterprise-grade standards of accuracy, safety, and trustworthiness.
This role focuses on evaluating, benchmarking, and stress-testing LLMs in real-world workflows, building frameworks for reliability, robustness, and continuous improvement. If you thrive at the intersection of AI research, applied testing, and responsible deployment, this is the role for you.
Ensure reliability in cutting-edge AI platforms that are redefining enterprise adoption.
Be part of one of the fastest-growing AI Foundries, powering Fortune 500 enterprises with trustworthy AI.
Grow into roles such as AI Systems Architect, Responsible AI Engineer, or Reliability Engineering Lead.
Work on enterprise-scale evaluation challenges across BFSI, Healthcare, Telecom, and GRC.
Your evaluations will directly shape production-grade AI agents used in mission-critical systems.
Our values — Agency, Taste, Ownership, Mastery, Impatience, and Customer Obsession — empower you to innovate fearlessly.
Join a company that prioritizes trustworthy, explainable, and compliant AI.
At XenonStack, we believe in shaping the future of intelligent systems. We foster a culture of cultivation built on bold, human-centric leadership principles, where deep work, simplicity, and adoption define everything we do.
Be part of our mission to accelerate the world’s transition to AI + Human Intelligence — by making AI agents not just powerful, but trustworthy and reliable.