Build the ops layer for the next generation of software companies.
We're a small, senior team solving hard problems at the intersection of infrastructure engineering and AI. Our customers are some of the fastest-growing startups in the world, and we're the team that keeps their systems running while they scale.
This isn't a monitoring dashboard company. We embed directly into customer infrastructure - real production systems, real incidents, real architecture decisions. If you want to work across dozens of different stacks and see every failure mode that exists, this is the job.
What we value
You'll own outcomes for real customer infrastructure. No tickets, no sprints, no standups about standups. You see a problem, you fix it.
Clear communication over constant meetings
We're a distributed team. Written communication matters. You should be able to explain a complex system issue clearly to both engineers and founders.
Building things that compound
We're building AI agent skills alongside our customer work. Every incident you resolve, every system you learn - it feeds back into tooling that makes the next one faster.
Open Roles
You'll be embedded directly with Pagerfree customers, owning their ops layer end-to-end. That means on-call coverage, incident response, performance optimization, and proactive infrastructure improvements across multiple customer environments.
What you'll do
- Own on-call and incident response for customer production systems. When things break at 2am, you're the one who fixes them - and then you fix the root cause so it doesn't happen again
- Performance engineering on real systems under real load. Slow queries, resource bottlenecks, scaling walls - diagnose and resolve
- Architecture review and system design guidance for customer engineering teams shipping new features
- Database optimization, particularly Postgres. Query tuning, indexing strategy, migration planning, capacity forecasting
- Build and improve our internal AI agent tooling that helps you and the rest of the team move faster over time
What we're looking for
- You've operated production infrastructure at meaningful scale. Kubernetes, AWS or GCP, container orchestration, CI/CD pipelines
- Strong Postgres experience. You've tuned queries on large datasets, planned migrations, and thought about replication and failover
- You've been on-call and you're good at it. Fast triage, clear communication during incidents, thorough postmortems after
- You can context-switch across multiple customer environments without losing depth
- You communicate well in writing. Incident summaries, architecture docs, customer-facing updates - all part of the job
Nice to have
- Experience with observability tooling (Datadog, Grafana, Prometheus)
- Familiarity with Python, Go, Rust, or TypeScript backend systems
- Background in healthcare, fintech, or other regulated industries
- Experience building developer tools or internal platforms
You'll build the AI agent infrastructure that powers Pagerfree's operations. Our agents analyze customer logs, pull context from past incidents, and help our engineers diagnose issues faster. You'll make those agents smarter, faster, and more reliable.
What you'll do
- Design and build AI agent workflows that interact with real customer infrastructure safely. Sandboxed execution, credential isolation, audit logging
- Improve how our agents pull context from diverse infrastructure signals - logs, metrics, alerting systems, codebases, Slack history
- Build reusable skills that encode how to diagnose and resolve specific classes of infrastructure problems. These compound over time and are a core part of our value proposition
- Work closely with our embedded engineers to understand what context they need during incidents and build tooling that delivers it automatically
- Ship fast in a small team. You'll own large pieces of the system with minimal overhead
What we're looking for
- Strong backend engineering fundamentals. You've built and operated production systems. Python, TypeScript, or Go preferred
- Familiarity with LLM APIs and agent architectures. You've built something real with language models, not just prototyped a chatbot
- Systems thinking. You understand how infrastructure components connect and fail. Experience with AWS, Kubernetes, Postgres, or similar is valuable because our agents need to understand these systems too
- Security-minded. Our agents touch production infrastructure. You think about sandboxing, credential management, and audit trails by default
Nice to have
- Experience with sandboxed execution environments
- Background in infrastructure engineering or SRE
- Experience building internal developer tools
- Familiarity with observability and monitoring systems
We're always interested in hearing from exceptional infrastructure engineers, backend engineers, and people who are deeply technical and want to work on hard problems.