Infra Engineer (AI Agents) | IDR 30-45m p/m

Antler

Denpasar

On-site

IDR 279,000,000 - 446,400,000

Full time

8 days ago
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Kulu in Bali is hiring an infrastructure engineer to run a realtime AI agent platform. You will join the Foundation team, report to the CTO, and focus on migrating from AWS to GCP, building monitoring from the ground up, and keeping production healthy when things break.

You will own reliability and observability of core services, manage the LLM streaming layer, and drive IaC across deployments, with a strong emphasis on security and data integrity.

Qualifications

  • 3+ years of production experience in DevOps/SRE/Cloud/Platform or backend infra.
  • Production experience owning infrastructure as code (Terraform/OpenTofu).
  • Monitoring and logging tooling knowledge (Prometheus/Grafana, Datadog, ELK).
  • Strong Python (asyncio) and Bash scripting skills.
  • Operational AWS or GCP experience and IAM/secrets/VPC knowledge.
  • CI/CD pipelines experience (GitHub Actions, GitLab CI) and Git.
  • Distributed systems knowledge and fault tolerance design.
  • PostgreSQL and Redis in production environments.
  • Security best practices for infrastructure and apps.
  • AI tooling experience and fluent English communication.

Responsibilities

  • Design, build, and maintain the realtime platform infrastructure (WebRTC, LLM sessions, Python backend).
  • Own reliability, performance, and observability end-to-end of core services.
  • Own the LLM streaming session layer: connection lifecycle, streaming, tool-call plumbing.
  • Manage deployment across AWS and CI/CD; push IaC with OpenTofu.
  • Maintain production platform and triage incidents; feed learnings into runbooks.
  • Handle sensitive meeting data with privacy and auditability.
  • Contribute to architecture, focusing on correctness, data integrity, and security.
  • Write well-tested, maintainable code with clear data models.
  • Maintain documentation and champion operational excellence.
  • Ship end-to-end: migrate, deploy, and monitor dashboards.

Skills

Python asyncio
FastAPI
Bash
Cloud: AWS/GCP
CI/CD
Git workflows
Distributed systems
PostgreSQL
Redis
Security best practices
AI tooling
English communication

Tools

Terraform/OpenTofu
Prometheus/Grafana
Datadog
ELK

Job description

Kulu builds an AI agent that joins live meetings — it listens, speaks, sees the user's shared screen, and takes actions in real time. Everything runs on realtime infrastructure: WebRTC media, streaming LLM sessions, a Python backend, and the cloud underneath.

We're hiring an infrastructure engineer to run it with us. You'll join our Foundation team, reporting directly to the CTO, and your work will centre on three things in the first half year: migrating us from AWS to GCP, building our monitoring from the ground up, and keeping production healthy when things break.

We're a small, high-bandwidth team, deliberately building toward a more in-person culture in Bali.

Tasks
  • Design, build, and maintain the realtime platform Kulu's AI assistant runs on — WebRTC media infrastructure, streaming multimodal LLM sessions, and the Python backend behind them: the foundation every user conversation rides on.
  • Own reliability, performance, and observability of Kulu's core services: the realtime pipeline end to end — rooms, agents, audio in/out, and the recording pipeline — and the telemetry that turns “the agent felt slow” into a number with a cause.
  • Own the LLM streaming session layer as an engineering system — connection lifecycle, streaming, reconnection and resumption, tool-call plumbing, timeouts and watchdogs — so that behaviour changes designed by the AI product lead land on a reliable substrate.
  • Own deployment and infrastructure across our AWS environment and CI/CD pipelines, and drive our infrastructure-as-code programme (OpenTofu).
  • Maintain the platform in production and respond to incidents: triage, root‑cause, and fix live issues across the stack, then feed every incident back into telemetry, runbooks, and infrastructure as code so it can't happen silently twice.
  • Build services that handle sensitive meeting data with the privacy, correctness, and auditability our customers expect.
  • Contribute to architectural decisions as Kulu scales, with a particular focus on system correctness, data integrity, and security.
  • Bring strong engineering fundamentals to every problem: well‑tested, maintainable code, clear data models, and systems that degrade gracefully under pressure.
  • Maintain strong documentation and champion operational excellence across the engineering culture.
  • Ship end to end: you write the migration, deploy the service, and watch the dashboards after.
Requirements
  • 3+ years of production experience in a DevOps, SRE, Cloud Engineer, Platform Engineer, or backend infrastructure role.
  • Production experience writing and owning infrastructure as code (Terraform/OpenTofu).
  • Knowledge of monitoring and logging tooling (Prometheus/Grafana, Datadog, ELK, or equivalent).
  • Strong programming skills in Python (asyncio, FastAPI or similar) and solid shell scripting (Bash).
  • Solid operational experience with AWS or GCP — ideally some of both, since your first project is migrating us from one to the other — including IAM, secrets management, VPC and networking.
  • Hands‑on experience building CI/CD pipelines (GitHub Actions, GitLab CI, or similar); proficiency with Git and branching strategies.
  • Strong grasp of distributed systems, event‑driven architecture, and fault‑tolerant service design.
  • Solid relational database skills (PostgreSQL, schema migrations); Redis in production.
  • Strong understanding of infrastructure and application security best practices.
  • Strong AI-tool skills: AI is part of how you read, write, and debug code.
  • Clear communicator; fluent spoken and written English — our company language is exclusively English.
Nice to Have
  • AWS or GCP certifications (e.g., Solutions Architect, DevOps Engineer).
  • Security compliance experience — supporting a certification programme ("SOC 2 / ISO 42001"), access management, or audit preparation.
  • Experience supporting eval or model‑quality infrastructure.
Our Stack

Python (FastAPI, asyncio) · WebRTC media infrastructure · streaming multimodal LLM APIs · PostgreSQL · Redis · React + TypeScript · AWS · GitHub Actions · OpenTofu · Datadog

Benefits

IDR 25–40m/month gross

What Success Looks Like
  • First month: you can deploy, configure, and debug our infrastructure on your own, and the GCP migration plan is ready — target architecture chosen, quotas requested, steps written down.
  • Three months: the migration is finished — our environments run on GCP, fully managed as code — the first dashboards and alerts are live, and when something breaks in production, you're the person who handles it.
  • Six months: AWS is fully shut down; "is production healthy" is one look at a dashboard you built; incidents are diagnosed in minutes instead of hours; the team ships every week without worrying about infrastructure.
Our Operating Principles:
  • Always Be Hustlin' – Move fast, outwork the competition, stay scrappy.
  • Relentlessly Curious – Reason from first principles, question assumptions and explore better ways.
  • Super Pumped – Show relentless enthusiasm and drive.
  • Make Magic – Create experiences that delight customers.
  • Obsess Over the Details – Perfect the product experience, no matter how small.
  • Big Bold Bets – Think big, take risks and go after huge, transformative opportunities.
  • Ownership is Key – Own the outcome, not just the task.

As Kulu grows, this role naturally expands with the Foundation team — a full-stack engineer will join on Simu, and there is a clear path to owning our infrastructure and reliability function outright.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Infra Engineer (AI Agents) | IDR 30-45m p/m
Infra Engineer (AI Agents) | IDR 30-45m p/m

frontierinteractionscom • Denpasar

On-site
IDR 279,000,000 - 446,400,000
Infra Engineer (On-site, Bali) | IDR 30-45m p/m
Infra Engineer (On-site, Bali) | IDR 30-45m p/m

frontierinteractionscom • Jakarta Pusat

On-site
IDR 334,800,000 - 502,200,000
IDR 30–45m/month gross
Infra Engineer (AI Agents) | IDR 25-45m p/m
Infra Engineer (AI Agents) | IDR 25-45m p/m

frontierinteractionscom • Denpasar

On-site
IDR 279,000,000 - 446,400,000
Infra Engineer (On-site, Bali) | IDR 30-45m p/m
Infra Engineer (On-site, Bali) | IDR 30-45m p/m

Kulu • Jakarta Pusat

On-site
Confidential
Infra Engineer (AI Agents) | IDR 25-40m p/m
Infra Engineer (AI Agents) | IDR 25-40m p/m

Kulu • Denpasar

On-site
Confidential
Infra Engineer (AI Agents) | IDR 30-45m p/m
Infra Engineer (AI Agents) | IDR 30-45m p/m

Kulu • Denpasar

On-site
IDR 279,000,000 - 446,400,000
Infra Engineer (On-site, Bali) | IDR 30-45m p/m
Infra Engineer (On-site, Bali) | IDR 30-45m p/m

amIT Global Solutions Sdn Bhd • Indonesia

On-site
IDR 334,800,000 - 502,200,000
Lead Product Engineer (Platform) | IDR 40-70m p/m + equity
Lead Product Engineer (Platform) | IDR 40-70m p/m + equity

Kulu • Denpasar

On-site
Confidential
IDR 40-70m/month gross + 0.2-0.5%equiv
Lead AI Product Manager @ Top AI Startup - Rp IDR 60-100m plus 0.5-1% equity
Lead AI Product Manager @ Top AI Startup - Rp IDR 60-100m plus 0.5-1% equity

Kulu • Denpasar

On-site
Confidential
Equity option
AI Solutions Consultant - AI Startup (Bali) - Rp IDR 22-32m/month
AI Solutions Consultant - AI Startup (Bali) - Rp IDR 22-32m/month

Antler • Denpasar

On-site
IDR 245,520,000 - 357,120,000
Rp IDR 22–32m/month gross, dependings
Amazing Bali Office