Get more replies from employers
Send a job-specific resume in minutes.
XenonStack is seeking an Agentic Infrastructure Observability Engineer to design and implement end-to-end observability for AI-native and multi-agent systems. You will monitor agents, pipelines, and infrastructure to ensure reliability, transparency, and cost efficiency.
You will work with AgentOps, DevOps, and Data teams to embed observability into CI/CD and build plugins for LLMs, agents, and data pipelines, shaping enterprise AI reliability at scale.
XenonStack is the fastest-growingData and AI Foundry for Agentic Systems, enabling enterprises to gainreal-time and intelligent business insights.
Inference AI Infrastructure for Agentic Systems→nexastack.ai
Our mission is to accelerate the world’s transition toAI + Human Intelligenceby building platforms that arescalable, reliable, and observable by design.
We are seeking anAgentic Infrastructure Observability Engineerto design and implementend-to-end observability frameworksfor AI-native and multi-agent systems.
This role sits at the heart ofAgentOps and Reliability Engineering— ensuring thatagents, pipelines, and infrastructureare monitored, measurable, and continuously optimized.
If you thrive onmetrics, monitoring, and making complex systems transparent and reliable, this role offers a chance to define observability for the next generation of enterprise AI.
Design and implementobservability pipelinescovering metrics, logs, traces, and cost telemetry for agentic systems.
TrackLLM usage, context windows, token allocation, and multi-agent interactions.
Build monitoring hooks intoLangChain, LangGraph, MCP, and RAG pipelines.
Define and monitorSLOs, SLIs, and SLAsfor agentic workflows and inference infrastructure.
Conduct root cause analysis ofagent failures, latency issues, and cost spikes.
Integrate observability intoCI/CD and AgentOps pipelines.
Develop custom plugins/scripts to extend observability for LLMs, agents, and data pipelines.
Work withAgentOps, DevOps, and Data Engineering teamsto ensure system-wide observability.
Provideexecutive-level reportingon reliability, efficiency, and adoption metrics.
Implementfeedback loopsto improve agent performance and reduce downtime.
Stay updated withstate-of-the-art observability and AI monitoring frameworks.
Build observability frameworks fornext-gen enterprise AI systems.
Be part of one of the fastest-growingAI Foundries, powering mission-critical agent deployments.
Advance into roles likeReliability Architect, AgentOps Lead, or Head of Observability.
Work on observability challenges acrossFortune 500 enterprises and global innovators.
Ensuretransparency, trust, and resiliencein production-grade AI systems.
Our values —Agency, Taste, Ownership, Mastery, Impatience, and Customer Obsession— give you autonomy to innovate and accountability to deliver.
Help enterprises adopt AI that isnot just powerful, but explainable and auditable.
At XenonStack, we believe inshaping the future of intelligent systems. We foster aculture of cultivationbuilt on bold, human-centric leadership principles, wheredeep work, simplicity, and adoptiondefine everything we do.
Be self-directed and proactive.
Sweat the details and build with precision.
Take responsibility for outcomes.
Commit to continuous learning and growth.
Move fast and embrace progress.
Always put the customer first.
Obsessed with Adoption – Making observability and trust an integral part of enterprise AI.
Obsessed with Simplicity – Turning complex monitoring into seamless, actionable insights.
Be part of our mission toaccelerate the world’s transition to AI + Human Intelligence— by making agentic AI systemstransparent, observable, and reliable at scale.