Site Reliability Engineer, AI Platform

XenonStack Moments

Vallecito 15-8 Hualtaco I

Presencial

PEN 308.000 - 444.000

Jornada completa

Hace 4 días
Sé de los primeros/as/es en solicitar esta vacante
Generador de candidaturas

Una candidatura hecha para este puesto de trabajo — un currículum y una carta de presentación adaptados que responden directamente a la oferta.

Supera los filtros ATS

Descripción de la vacante

XenonStack is hiring an Site Reliability Engineer, AI Platform to design end-to-end observability for AI-native and multi-agent systems. You will monitor agents, pipelines and infrastructure, building dashboards and alerting for real-time reliability.

You will define SLOs/SLAs, investigate root causes, and integrate observability into CI/CD and AgentOps, collaborating with cross-functional teams to improve reliability at scale.

Formación

  • 3–6 years of experience in SRE/DevOps for AI platforms.
  • Strong observability with Prometheus, Grafana, ELK and OpenTelemetry.
  • Cloud and Kubernetes monitoring on AWS/GCP/Azure.
  • Scripting in Python, Go or Bash for automation.
  • Understanding of AI/LLM pipelines and data pipelines.
  • Experience with CI/CD and monitoring-as-code.

Responsabilidades

  • Design end-to-end observability for agentic AI systems.
  • Build dashboards, logs, traces, and cost telemetry pipelines.
  • Monitor LLM usage, token allocation, and multi-agent interactions.
  • Define SLOs/SLIs/SLAs for agent workflows and infra.
  • Integrate observability into CI/CD and AgentOps pipelines.
  • Provide executive reporting on reliability and adoption metrics.

Conocimientos

Observability
Prometheus
Grafana
OpenTelemetry
Jaeger
Kubernetes monitoring
Python
Go
CI/CD
SRE / DevOps

Educación

Bachelor’s degree in CS/Eng or equivalent

Herramientas

LangChain
LangGraph
MCP
RAG pipelines
OpenTelemetry
Prometheus
Grafana
ELK
Jaeger
CI/CD tooling
Weigths & Biases
Arize AI

Descripción del empleo

About Xenonstack

XenonStack is a Data and AI Foundry for Agentic Systems, enabling enterprises to design, deploy, operate, and scale intelligent agents across digital and physical environments.

We Build Enterprise-grade Platforms Across The Agentic Stack
  • Akira AI — Reasoning and agent orchestration. Turn models into collaborative, policy-governed agents that learn and act together.
  • ElixirData — Agentic analytics intelligence. Explainable, decision-centric analytics for measurable business outcomes.
  • NexaStack — Agentic infrastructure automation. Secure, compliant AI deployment across cloud, edge, and on-prem.
  • MetaSecure — Trust, compliance and defense. Continuous assurance with AI-BOMs, risk scoring, and agentic security.

Our mission is to accelerate the world’s transition to AI + Human Intelligence by making agentic systems reliable, responsible, and enterprise-ready.

THE OPPORTUNITY

We are seeking an Site Reliability Engineer, AI Platform to design and implement end-to-end observability frameworks for AI-native and multi-agent systems.

This role sits at the heart of AgentOps and Reliability Engineering — ensuring that agents, pipelines, and infrastructure are monitored, measurable, and continuously optimized.

If you thrive on metrics, monitoring, and making complex systems transparent and reliable, this role offers a chance to define observability for the next generation of enterprise AI.

Key Responsibilities
  • Observability Frameworks
    • Design and implement observability pipelines covering metrics, logs, traces, and cost telemetry for agentic systems.
    • Build dashboards and alerting systems to monitor reliability, performance, and drift in real-time.
  • Agentic AI Monitoring
    • Track LLM usage, context windows, token allocation, and multi-agent interactions.
    • Build monitoring hooks into LangChain, LangGraph, MCP, and RAG pipelines.
  • Reliability & Performance
    • Define and monitor SLOs, SLIs, and SLAs for agentic workflows and inference infrastructure.
    • Conduct root cause analysis of agent failures, latency issues, and cost spikes.
  • Automation & Tooling
    • Integrate observability into CI/CD and AgentOps pipelines.
    • Develop custom plugins/scripts to extend observability for LLMs, agents, and data pipelines.
  • Collaboration & Reporting
    • Work with AgentOps, DevOps, and Data Engineering teams to ensure system-wide observability.
    • Provide executive-level reporting on reliability, efficiency, and adoption metrics.
  • Continuous Improvement
    • Implement feedback loops to improve agent performance and reduce downtime.
    • Stay updated with state-of-the-art observability and AI monitoring frameworks.
Skills & Qualifications

Must-Have

  • 3–6 years of experience in SRE, DevOps, or Site Reliability Engineer, AI Platforming.
  • Strong knowledge of observability tools (Prometheus, Grafana, ELK, OpenTelemetry, Jaeger).
  • Experience with cloud-native infrastructure (AWS, GCP, Azure) and Kubernetes monitoring.
  • Proficiency in Python, Go, or Bash for scripting and automation.
  • Understanding of AI/LLM pipelines, RAG systems, and vector databases.
  • Hands-on with CI/CD pipelines and monitoring-as-code.

Good-to-Have

  • Experience with AgentOps tools (LangSmith, PromptLayer, Arize AI, Weights & Biases).
  • Exposure to AI-specific observability (token usage, model latency, hallucination tracking).
  • Knowledge of Responsible AI monitoring frameworks.
  • Background in BFSI, GRC, SOC, or other regulated industries.
WHY SHOULD YOU JOIN US?
  • Agentic AI Product Company

Build observability frameworks for next-gen enterprise AI systems.

  • A Fast-Growing Category Leader

Be part of one of the fastest-growing AI Foundries, powering mission-critical agent deployments.

  • Career Mobility & Growth

Advance into roles like Reliability Architect, AgentOps Lead, or Head of Observability.

  • Global Exposure

Work on observability challenges across Fortune 500 enterprises and global innovators.

  • Create Real Impact

Ensure transparency, trust, and resilience in production-grade AI systems.

  • Culture of Excellence

Our values — Agency, Taste, Ownership, Mastery, Impatience, and Customer Obsession — give you autonomy to innovate and accountability to deliver.

  • Responsible AI First

Help enterprises adopt AI that is not just powerful, but explainable and auditable.

XENONSTACK CULTURE – JOIN US & MAKE AN IMPACT!

At XenonStack, we believe in shaping the future of intelligent systems. We foster a culture of cultivation built on bold, human-centric leadership principles, where deep work, simplicity, and adoption define everything we do.

Our Cultural Values
  • Agency – Be self-directed and proactive.
  • Taste – Sweat the details and build with precision.
  • Ownership – Take responsibility for outcomes.
  • Mastery – Commit to continuous learning and growth.
  • Impatience – Move fast and embrace progress.
  • Customer Obsession – Always put the customer first.
Our Product Philosophy
  • Obsessed with Adoption – Making observability and trust an integral part of enterprise AI.
  • Obsessed with Simplicity – Turning complex monitoring into seamless, actionable insights.

Be part of our mission to accelerate the world’s transition to AI + Human Intelligence — by making agentic AI systems transparent, observable, and reliable at scale.

Consigue la evaluación confidencial y gratuita de tu currículum.

o arrastra y suelta tu archivo aquí

Similar jobs

Puestos de trabajo similares que vale la pena comparar

Site Reliability Engineer, AI Platform
Site Reliability Engineer, AI Platform

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 376.000 - 478.000
MLOps Engineer
MLOps Engineer

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 90.000 - 130.000
MLOps Engineer
MLOps Engineer

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 60.000 - 100.000
Forward Deployed Solutions Engineer, Agentic Systems
Forward Deployed Solutions Engineer, Agentic Systems

XenonStack Moments • San Martin Cp3 Hualtaco I

Híbrido
PEN 305.000 - 475.000
Hybrid work model
Global exposure
Career growth opportunities
Forward Deployed Solutions Engineer, Agentic Systems
Forward Deployed Solutions Engineer, Agentic Systems

XenonStack Moments • Vallecito 15-8 Hualtaco I

Híbrido
PEN 308.000 - 444.000
Machine Learning Engineer, Reinforcement Learning
Machine Learning Engineer, Reinforcement Learning

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 100.000 - 180.000
Solution Architect, DevOps
Solution Architect, DevOps

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 140.000 - 210.000
Strategic leadership opportunities
Performance-based rewards
Freedom to innovate with emerging tech
+2
Solution Architect, DevOps
Solution Architect, DevOps

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 373.000 - 475.000
Comprehensive medical insurance
Leadership opportunities
Professional development budget
Applied Scientist
Applied Scientist

XenonStack Moments • San Martin Cp3 Hualtaco I

Presencial
PEN 305.000 - 441.000
Applied Scientist
Applied Scientist

XenonStack Moments • Vallecito 15-8 Hualtaco I

Presencial
PEN 120.000 - 210.000