About kaiko
Kaiko is building a next‑generation agentic clinical AI assistant that helps clinicians reason across patient data, guidelines, and diagnostics.
About the role
This is a frontline reliability role focused on keeping the platform healthy through on‑call, incident response, and continuous improvement of the SRE programme. You’ll have room to shape how reliability works at kaiko rather than inherit a rigid process.
Responsibilities
- Own the reactive frontline: carry on‑call, triage alerts quickly and methodically, and drive incidents to resolution with clear communication and clean handoffs.
- Make alerts trustworthy: treat noisy or low‑value alerts as defects to fix, not background noise; move toward structured, queryable signals and precise, actionable alerting.
- Turn incidents into durable fixes: run and contribute to blameless post‑mortems, then follow through across teams so action items land in our services.
- Strengthen the SRE programme: write automation and tooling that removes toil, improves runbooks and dashboards, and leaves the on‑call rotation better documented.
Qualifications
- Deep Kubernetes experience: debug real failure modes under pressure—scheduling, networking, resource pressure, control‑plane versus workload issues.
- Strong Linux fundamentals: comfortable debugging at the OS level—processes, networking, filesystems, resource limits.
- Solid programming ability: eliminate toil, develop maintainable automation and tooling; experience building products shows good problem‑solving.
- Observability and incident‑response fluency: comfortable in metrics, logs, traces; write queries to isolate problems and stay composed during incidents.
- Internalized SRE principles: SLO/SLI, error‑budget thinking, bias toward prevention, instinct for reducing toil.
- Experience of roughly 3‑6 years in production SRE, platform, or infrastructure roles—skill demonstrated over experience.
Soft Skills
- Solution‑oriented and pragmatic; willing to build end‑to‑end solutions when needed.
- Collaborative communicator who coaches teams and writes clear, actionable guidance.
- Bias to automate and remove toil.
Nice to Have
- Experience with structured logging and modern alerting practices.
- Infrastructure‑as‑code and CI/CD fluency (Terraform, Helm, GitOps, etc.).
- Familiarity with incident‑management tooling (Rootly, PagerDuty, incident.io) and related post‑mortem discipline.
- Exposure to regulated or high‑stakes domains (health, fintech, critical infrastructure).
Benefits
- Competitive salary, good pension plan and 25 vacation days per year.
- Great offsites and team events.
- EUR 1000 learning and development budget.
- Autonomy to work flexible style.
- Annual commuting subsidy.
Equal Employment Opportunity
Kaiko is committed to building a diverse team and to an inclusive, equitable hiring process. We welcome applicants of all backgrounds.