Senior Infrastructure Engineer

Pagerfree, Inc.

Northern (KY)

Hybrid

USD 120,000 - 180,000

Full time

10 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Job summary

Pagerfree, Inc. is seeking a senior infrastructure engineer to own the ops layer end-to-end for our production customers. You’ll be embedded with clients, handle on-call incidents, optimize performance, and drive proactive infrastructure improvements across multiple environments.

We value clear communication, rapid root-cause analysis, and building AI-enabled tooling that compounds over time. You'll work with Kubernetes, AWS, Postgres, and observability tooling to keep systems reliable as our

Qualifications

  • Experience operating production infrastructure at scale.
  • Strong backend and systems thinking.
  • Experience with AWS, Kubernetes, and Postgres is valuable.

Responsibilities

  • Own on-call and incident response for customer production systems.
  • Performance engineering to diagnose slow queries and bottlenecks.
  • Architecture review and system design guidance for customer teams.
  • Database optimization for Postgres including tuning and migrations.
  • Build and improve AI agent tooling to speed up incidents.
  • Design and build AI agent workflows for real customer infrastructure with sandboxing and audit logging.
  • Improve context gathering from logs, metrics, alerts, code, and Slack history.
  • Create reusable diagnostic skills that scale with use.
  • Collaborate with embedded engineers to deliver contextual tooling.
  • Ship fast in a small team with ownership of major system components.

Skills

Production-scale infra
Kubernetes
AWS or GCP
CI/CD pipelines
Postgres tuning
On-call experience
Backend development

Tools

Datadog
Grafana
Prometheus
Python
Go
TypeScript

Job description

Build the ops layer for the next generation of software companies.

We're a small, senior team solving hard problems at the intersection of infrastructure engineering and AI. Our customers are some of the fastest-growing startups in the world, and we're the team that keeps their systems running while they scale.

This isn't a monitoring dashboard company. We embed directly into customer infrastructure - real production systems, real incidents, real architecture decisions. If you want to work across dozens of different stacks and see every failure mode that exists, this is the job.

What we value

You'll own outcomes for real customer infrastructure. No tickets, no sprints, no standups about standups. You see a problem, you fix it.

Clear communication over constant meetings

We're a distributed team. Written communication matters. You should be able to explain a complex system issue clearly to both engineers and founders.

Building things that compound

We're building AI agent skills alongside our customer work. Every incident you resolve, every system you learn - it feeds back into tooling that makes the next one faster.

Open Roles

You'll be embedded directly with Pagerfree customers, owning their ops layer end-to-end. That means on-call coverage, incident response, performance optimization, and proactive infrastructure improvements across multiple customer environments.

What you'll do
  • Own on-call and incident response for customer production systems. When things break at 2am, you're the one who fixes them - and then you fix the root cause so it doesn't happen again
  • Performance engineering on real systems under real load. Slow queries, resource bottlenecks, scaling walls - diagnose and resolve
  • Architecture review and system design guidance for customer engineering teams shipping new features
  • Database optimization, particularly Postgres. Query tuning, indexing strategy, migration planning, capacity forecasting
  • Build and improve our internal AI agent tooling that helps you and the rest of the team move faster over time
What we're looking for
  • You've operated production infrastructure at meaningful scale. Kubernetes, AWS or GCP, container orchestration, CI/CD pipelines
  • Strong Postgres experience. You've tuned queries on large datasets, planned migrations, and thought about replication and failover
  • You've been on-call and you're good at it. Fast triage, clear communication during incidents, thorough postmortems after
  • You can context-switch across multiple customer environments without losing depth
  • You communicate well in writing. Incident summaries, architecture docs, customer-facing updates - all part of the job
Nice to have
  • Experience with observability tooling (Datadog, Grafana, Prometheus)
  • Familiarity with Python, Go, Rust, or TypeScript backend systems
  • Background in healthcare, fintech, or other regulated industries
  • Experience building developer tools or internal platforms

You'll build the AI agent infrastructure that powers Pagerfree's operations. Our agents analyze customer logs, pull context from past incidents, and help our engineers diagnose issues faster. You'll make those agents smarter, faster, and more reliable.

What you'll do
  • Design and build AI agent workflows that interact with real customer infrastructure safely. Sandboxed execution, credential isolation, audit logging
  • Improve how our agents pull context from diverse infrastructure signals - logs, metrics, alerting systems, codebases, Slack history
  • Build reusable skills that encode how to diagnose and resolve specific classes of infrastructure problems. These compound over time and are a core part of our value proposition
  • Work closely with our embedded engineers to understand what context they need during incidents and build tooling that delivers it automatically
  • Ship fast in a small team. You'll own large pieces of the system with minimal overhead
What we're looking for
  • Strong backend engineering fundamentals. You've built and operated production systems. Python, TypeScript, or Go preferred
  • Familiarity with LLM APIs and agent architectures. You've built something real with language models, not just prototyped a chatbot
  • Systems thinking. You understand how infrastructure components connect and fail. Experience with AWS, Kubernetes, Postgres, or similar is valuable because our agents need to understand these systems too
  • Security-minded. Our agents touch production infrastructure. You think about sandboxing, credential management, and audit trails by default
Nice to have
  • Experience with sandboxed execution environments
  • Background in infrastructure engineering or SRE
  • Experience building internal developer tools
  • Familiarity with observability and monitoring systems

We're always interested in hearing from exceptional infrastructure engineers, backend engineers, and people who are deeply technical and want to work on hard problems.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Software Engineer - Infrastructure
Software Engineer - Infrastructure

Emergent • San Francisco (CA)

On-site
USD 180,000 - 240,000
401(k)
Health, dental, and vision insurance
Unlimited Paid Time Off
+1
Platform Engineer
Platform Engineer

Docflow Labs Inc • Los Angeles (CA)

On-site
USD 150,000 - 185,000
Health insurance
Platform Engineer
Platform Engineer

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Meals provided (breakfast, lunch, and/
Snacks, drinks and coffee daily
+2
Staff Engineer
Staff Engineer

360 Privacy • Brentwood (TN)

On-site
USD 180,000 - 260,000
Digital privacy protection
AI Infrastructure Engineer
AI Infrastructure Engineer

Percepta • New York (NY)

On-site
USD 140,000 - 190,000
Senior Site Reliability Engineer, AI-DNA, $100k/year USD
Senior Site Reliability Engineer, AI-DNA, $100k/year USD

IgniteTech • United States

On-site
USD 120,000 - 160,000
Senior Infra Engineer: AI-Driven Ops & On-Call Mastery
Senior Infra Engineer: AI-Driven Ops & On-Call Mastery

Pagerfree, Inc. • Northern (KY)

Hybrid
USD 120,000 - 180,000
Senior Platform Engineer
Senior Platform Engineer

Rifa AI • United States

Remote
USD 140,000 - 190,000
Staff Engineer - Distributed Systems
Staff Engineer - Distributed Systems

United States Digital Space LLC • United States

Remote
USD 180,000 - 240,000
Software Engineer, Infrastructure & Reliability
Software Engineer, Infrastructure & Reliability

CrewAI, Inc. • Northern (KY)

Hybrid
USD 150,000 - 190,000