Machine Learning Engineer, Model Evaluation

XenonStack Moments

San Martin Cp3 Hualtaco I

Presencial

PEN 308.000 - 444.000

Jornada completa

hace 27 horas
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Descripción de la vacante

XenonStack is seeking a Machine Learning Engineer, Model Evaluation to ensure large language models meet enterprise-grade accuracy, safety, and trustworthiness in real-world workflows. You will design evaluation pipelines, develop benchmarking tools, and run stress tests; you will also partner with ML engineers, product managers, and Responsible AI teams to align metrics with business goals.

This role offers opportunities to advance RLHF/RLAIF and maintain central repositories of test cases,

Formación

  • 3–6 years in AI/ML, NLP, or applied model evaluation.
  • Strong understanding of LLM architectures, prompt engineering, and failure modes.
  • Hands-on with evaluation frameworks (Eval harnesses, Ragas, OpenAI Evals, DeepEval).
  • Proficiency in Python and libraries like LangChain, LangGraph, LlamaIndex, Hugging Face.
  • Experience with vector databases, RAG pipelines, and knowledge graph integration.
  • Familiarity with bias/fairness testing and Responsible AI frameworks.

Responsabilidades

  • Evaluation Frameworks: Design and implement LLM evaluation pipelines covering accuracy, robustness, safety, and bias.
  • Evaluation Frameworks: Develop automated systems for benchmarking models on enterprise-relevant tasks.
  • Reliability Engineering: Conduct stress tests, adversarial testing, and edge-case evaluations.
  • Reliability Engineering: Build tools to measure latency, consistency, and error recovery in multi-turn interactions.
  • Metrics & Monitoring: Define KPIs such as factual accuracy, hallucination rate, toxicity, and compliance alignment.
  • Metrics & Monitoring: Establish real-time monitoring for drift, anomalies, and performance regressions.
  • Collaboration & Alignment: Partner with ML engineers, product managers, and domain experts to align evaluation with business objectives.
  • Collaboration & Alignment: Work with Responsible AI teams to implement ethical, explainable, and compliant evaluation practices.
  • Continuous Improvement: Feed insights from evaluation into fine-tuning, RLHF/RLAIF pipelines, and model selection.
  • Continuous Improvement: Maintain a central repository of test cases, benchmarks, and evaluation results.
  • Research & Innovation: Stay current with state-of-the-art LLM evaluation techniques, from academic benchmarks to applied enterprise metrics.
  • Research & Innovation: Explore automated evaluation using agentic test harnesses and synthetic data generation.

Conocimientos

AI/ML 3-6 yrs
LLM architectures
Prompt engineering
Evaluation frameworks
Python + LangChain
Vector databases & RAG
Bias/fairness testing

Herramientas

Eval harnesses
Ragas
OpenAI Evals
DeepEval
LangChain
LlamaIndex
Hugging Face

Descripción del empleo

About Xenonstack XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.

About Xenonstack XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.

We Build Enterprise-grade Platforms Across The Agentic Stack

  • Akira AI — Reasoning and agent orchestration. Turn models into collaborative, policy-governed agents that learn and act together.
  • ElixirData — Agentic analytics intelligence. Explainable, decision-centric analytics for measurable business outcomes.
  • NexaStack — Agentic infrastructure automation. Secure, compliant AI deployment across cloud, edge, and on-prem.
  • MetaSecure — Trust, compliance and defense. Continuous assurance with AI-BOMs, risk scoring, and agentic security.

Our mission is to accelerate the world’s transition to AI + Human Intelligence by making agentic systems reliable, responsible, and enterprise-ready.

THE OPPORTUNITY

We are seeking an Machine Learning Engineer, Model Evaluation to ensure that large language models (LLMs) and agentic AI systems meet enterprise-grade standards of accuracy, safety, and trustworthiness.

This role focuses on evaluating, benchmarking, and stress-testing LLMs in real-world workflows, building frameworks for reliability, robustness, and continuous improvement. If you thrive at the intersection of AI research, applied testing, and responsible deployment, this is the role for you.

Key Responsibilities
  • Evaluation Frameworks
    • Design and implement LLM evaluation pipelines covering accuracy, robustness, safety, and bias.
    • Develop automated systems for benchmarking models on enterprise-relevant tasks.
  • Reliability Engineering
    • Conduct stress tests, adversarial testing, and edge-case evaluations.
    • Build tools to measure latency, consistency, and error recovery in multi-turn interactions.
  • Metrics & Monitoring
    • Define KPIs such as factual accuracy, hallucination rate, toxicity, and compliance alignment.
    • Establish real-time monitoring for drift, anomalies, and performance regressions.
  • Collaboration & Alignment
    • Partner with ML engineers, product managers, and domain experts to align evaluation with business objectives.
    • Work with Responsible AI teams to implement ethical, explainable, and compliant evaluation practices.
  • Continuous Improvement
    • Feed insights from evaluation into fine-tuning, RLHF/RLAIF pipelines, and model selection.
    • Maintain a central repository of test cases, benchmarks, and evaluation results.
  • Research & Innovation
    • Stay current with state-of-the-art LLM evaluation techniques, from academic benchmarks to applied enterprise metrics.
    • Explore automated evaluation using agentic test harnesses and synthetic data generation.
Skills & Qualifications
Must-Have
  • 3––6 years in AI/ML, NLP, or applied model evaluation.
  • Strong understanding of LLM architectures, prompt engineering, and failure modes.
  • Hands-on with evaluation frameworks (Eval harnesses, Ragas, OpenAI Evals, DeepEval).
  • Proficiency in Python and libraries like LangChain, LangGraph, LlamaIndex, Hugging Face.
  • Experience with vector databases, RAG pipelines, and knowledge graph integration.
  • Familiarity with bias/fairness testing and Responsible AI frameworks.
Good-to-Have
  • Experience with reinforcement learning (RLHF, RLAIF) and reward modeling.
  • Exposure to agentic evaluation frameworks (multi-agent stress testing, synthetic user simulators).
  • Knowledge of compliance and safety requirements for BFSI, GRC, or SOC use cases.
  • Contributions to open-source evaluation libraries or research papers.
WHY SHOULD YOU JOIN US?
  • Agentic AI Product Company

Ensure reliability in cutting-edge AI platforms that are redefining enterprise adoption.

  • A Fast‑Growing Category Leader

Be part of one of the fastest-growing AI Foundries, powering Fortune 500 enterprises with trustworthy AI.

  • Career Mobility & Growth

Grow into roles such as AI Systems Architect, Responsible AI Engineer, or Reliability Engineering Lead.

  • Global Exposure

Work on enterprise-scale evaluation challenges across BFSI, Healthcare, Telecom, and GRC.

  • Create Real Impact

Your evaluations will directly shape production-grade AI agents used in mission-critical systems.

  • Culture of Excellence

Our values — Agency, Taste, Ownership, Mastery, Impatience, and Customer Obsession — empower you to innovate fearlessly.

  • Responsible AI First

Join a company that prioritizes trustworthy, explainable, and compliant AI.

XENONSTACK CULTURE – JOIN US & MAKE AN IMPACT!

At XenonStack, we believe in shaping the future of intelligent systems. We foster a culture of cultivation built on bold, human-centric leadership principles, where deep work, simplicity, and adoption define everything we do.

Our Cultural Values
  • Agency – Be self-directed and proactive.
  • Taste – Sweat the details and build with precision.
  • Ownership – Take responsibility for outcomes.
  • Mastery – Commit to continuous learning and growth.
  • Impatience – Move fast and embrace progress.
  • Customer Obsession – Always put the customer first.
Our Product Philosophy
  • Obsessed with Adoption – Making AI accessible, reliable, and enterprise-ready.
  • Obsessed with Simplicity – Turning complex evaluation challenges into seamless, automated frameworks.

Be part of our mission to accelerate the world’s transition to AI + Human Intelligence — by making AI agents not just powerful, but trustworthy and reliable.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Machine Learning Engineer, Model Evaluation
Machine Learning Engineer, Model Evaluation

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 308.000 - 444.000
Machine Learning Engineer, Reinforcement Learning
Machine Learning Engineer, Reinforcement Learning

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 305.000 - 407.000
MLOps Engineer
MLOps Engineer

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 90.000 - 130.000
Applied Scientist
Applied Scientist

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 305.000 - 441.000
MLOps Engineer
MLOps Engineer

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 60.000 - 100.000
Machine Learning Engineer, Agentic Systems
Machine Learning Engineer, Agentic Systems

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 60.000 - 100.000
Continuous Learning
Certifications & Workshops
Cutting-edge Projects
+4
Site Reliability Engineer, AI Platform
Site Reliability Engineer, AI Platform

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 308.000 - 444.000
Applied Scientist
Applied Scientist

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 120.000 - 210.000
Site Reliability Engineer, AI Platform
Site Reliability Engineer, AI Platform

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 376.000 - 478.000
Machine Learning Engineer, Reinforcement Learning
Machine Learning Engineer, Reinforcement Learning

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 100.000 - 180.000