Kulu builds an AI agent that joins live meetings — it listens, speaks, sees the user's shared screen, and takes actions in real time. Everything runs on realtime infrastructure: WebRTC media, streaming LLM sessions, a Python backend, and the cloud underneath.
We're hiring an infrastructure engineer to run it with us. You'll join our Foundation team, reporting directly to the CTO, and your work will centre on three things in the first half year: migrating us from AWS to GCP, building our monitoring from the ground up, and keeping production healthy when things break.
We're a small, high-bandwidth team, deliberately building toward a more in-person culture in Bali.
Tasks
- Design, build, and maintain the realtime platform Kulu's AI assistant runs on — WebRTC media infrastructure, streaming multimodal LLM sessions, and the Python backend behind them: the foundation every user conversation rides on.
- Own reliability, performance, and observability of Kulu's core services: the realtime pipeline end to end — rooms, agents, audio in/out, and the recording pipeline — and the telemetry that turns “the agent felt slow” into a number with a cause.
- Own the LLM streaming session layer as an engineering system — connection lifecycle, streaming, reconnection and resumption, tool-call plumbing, timeouts and watchdogs — so that behaviour changes designed by the AI product lead land on a reliable substrate.
- Own deployment and infrastructure across our AWS environment and CI/CD pipelines, and drive our infrastructure-as-code programme (OpenTofu).
- Maintain the platform in production and respond to incidents: triage, root-cause, and fix live issues across the stack, then feed every incident back into telemetry, runbooks, and infrastructure as code so it can't happen silently twice.
- Build services that handle sensitive meeting data with the privacy, correctness, and auditability our customers expect.
- Contribute to architectural decisions as Kulu scales, with a particular focus on system correctness, data integrity, and security.
- Bring strong engineering fundamentals to every problem: well-tested, maintainable code, clear data models, and systems that degrade gracefully under pressure.
- Maintain strong documentation and champion operational excellence across the engineering culture.
- Ship end to end: you write the migration, deploy the service, and watch the dashboards after.
Requirements
- 3+ years of production experience in a DevOps, SRE, Cloud Engineer, Platform Engineer, or backend infrastructure role.
- Production experience writing and owning infrastructure as code (Terraform/OpenTofu).
- Knowledge of monitoring and logging tooling (Prometheus/Grafana, Datadog, ELK, or equivalent).
- Strong programming skills in Python (asyncio, FastAPI or similar) and solid shell scripting (Bash).
- Solid operational experience with AWS or GCP — ideally some of both, since your first project is migrating us from one to the other — including IAM, secrets management, VPC and networking.
- Hands-on experience building CI/CD pipelines (GitHub Actions, GitLab CI, or similar); proficiency with Git and branching strategies.
- Strong grasp of distributed systems, event-driven architecture, and fault-tolerant service design.
- Solid relational database skills (PostgreSQL, schema migrations); Redis in production.
- Strong understanding of infrastructure and application security best practices.
- Strong AI-tool skills: AI is part of how you read, write, and debug code.
- Clear communicator; fluent spoken and written English — our company language is exclusively English.
Nice to Have
- AWS or GCP certifications (e.g., Solutions Architect, DevOps Engineer).
- Security compliance experience — supporting a certification programme (SOC 2 / ISO 42001), access management, or audit preparation.
- Experience supporting eval or model-quality infrastructure.
Our Stack
Python (FastAPI, asyncio) · WebRTC media infrastructure · streaming multimodal LLM APIs · PostgreSQL · Redis · React + TypeScript · AWS · GitHub Actions · OpenTofu · Datadog
Benefits
IDR 25–40m/month gross
What Success Looks Like
- First month: you can deploy, configure, and debug our infrastructure on your own, and the GCP migration plan is ready — target architecture chosen, quotas requested, steps written down.
- Three months: the migration is finished — our environments run on GCP, fully managed as code — the first dashboards and alerts are live, and when something breaks in production, you're the person who handles it.
- Six months: AWS is fully shut down; "is production healthy" is one look at a dashboard you built; incidents are diagnosed in minutes instead of hours; the team ships every week without worrying about infrastructure.
Our Operating Principles
- Always Be Hustlin' – Move fast, outwork the competition, stay scrappy.
- Relentlessly Curious – Reason from first principles, question assumptions and explore better ways.
- Super Pumped – Show relentless enthusiasm and drive.
- Make Magic – Create experiences that delight customers.
- Obsess Over the Details – Perfect the product experience, no matter how small.
- Big Bold Bets – Think big, take risks and go after huge, transformative opportunities.
- Ownership is Key – Own the outcome, not just the task.
As Kulu grows, this role naturally expands with the Foundation team — a full-stack engineer will join on Simu, and there is a clear path to owning our infrastructure and reliability function outright.