Описание
ClickStack is ClickHouse’s open-source observability platform, unifying logs, metrics, traces, and session replays so engineers can find root causes quickly. It is developing an agent layer that can investigate incidents, propose root causes, and provide concise summaries.
Задачи
- Build agents that investigate incidents, surface anomalies, and answer why production is broken using ClickStack
- Build a library of reusable skills that captures debugging, root-cause analysis, ClickHouse query writing, and incident response workflows
- Own the agent stack end to end, including context engineering, tool design, evaluations, tracing, and cost
- Build MCP servers, SDKs, and integrations that let customers’ agents read telemetry, take action, and remain observable
- Collaborate with open-source contributors and customers, troubleshoot their problems, and feed learnings back into the product
- Address latency, cost, context window limits, evaluation coverage, and hallucinations on real telemetry
Требования
- 5+ Years of software engineering experience, including 1–2 years working on LLM-powered systems or agents in production
- Strong backend skills in TypeScript/Node.js and/or Python; comfortable with both, even if one is the primary language
- Hands-on experience building and shipping agents with multi-step tool use, planning, memory, and error recovery
- Experience designing skills using Markdown-based workflow encodings, such as Anthropic-style skills, and understanding when to use a skill, a tool, or both
- Experience with MCP, including building servers, designing tools, and considering authentication, scoping, and observability for agentic systems
- Strong evaluation practices, including golden sets, LLM-as-judge, and regression detection
- Proficiency in SQL and ability to write ClickHouse queries directly
- Comfortable with Docker and Kubernetes
- Active in open source and the developer community
- Будет плюсом: production agents in observability, incident response, or SRE; agent observability expertise, including tracing, cost attribution, evaluation pipelines, or OpenTelemetry for agents; prompt caching, context compaction, or other techniques for running agents on production telemetry volumes; columnar databases and event ingestion pipelines; contributions to or maintenance of an open-source AI/agent project; Go, Rust, or other systems languages for integrations and high-throughput infrastructure
Условия
- Flexible work environment
- Employer contributions toward healthcare
- Stock options for every new team member
- Flexible time off in the US and generous entitlement in other countries
- USD$500 home office setup for remote employees
- Opportunities to engage with colleagues at company-wide offsites
- The role is at a rapidly scaling start-up, with an opportunity to help shape its culture