Site Reliability Engineer

Aisle

New York (NY)

On-site

USD 140,000 - 190,000

Full time

15 hours ago
Be an early applicant

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aisle is building real-time retail tech that nudges actions and optimizes outcomes as they happen. We’re seeking a Senior SRE/DevOps engineer to own reliability, scalability, and observability of our GCP-based infrastructure in the US market.

You will design and implement monitoring with Datadog, manage IAM and security, participate in on-call rotations, incident response, and post-mortems, and drive durable improvements across distributed systems and production workflows.

Qualifications

  • 4+ years in SRE, DevOps, or infrastructure roles.
  • Strong hands-on experience with GCP.
  • Experience with event-driven/serverless architectures (Pub/Sub, Cloud Functions).
  • Familiarity with IAM, security best practices, and observability (Datadog).

Responsibilities

  • Own reliability, scalability, and observability of our GCP-based infra.
  • Design monitoring and alerting with Datadog; create monitors-as-code.
  • Manage IAM, service accounts, and security practices across cloud.
  • Participate in on-call rotations and post-mortems; drive systemic improvements.
  • Stabilize core infrastructure under heavy concurrent load.

Skills

SRE/DevOps
GCP
Event-driven systems
Observability (Datadog)
Security & IAM
Node.js
Backend engineering
AI in engineering
On-call & incident response

Tools

Cloud Functions
Pub/Sub
Workflows
Kubernetes
Vercel
PostgreSQL
Prisma
Redis
BullMQ
OpenClaw
LLM gateway
pgbouncer
Node.js
TypeScript

Job description

Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened.

Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact.

This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time.

Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment.

We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis.

You’ll win here if…

What is Aisle?

Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened.

Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact.

This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time.

Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment.

We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis.

  • You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actio
  • Can solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the system
  • You’re comfortable moving quickly, shipping improvements, and iterating in production
  • AI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems faster
  • You treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering)
  • You have experience building or orchestrating AI/agent workflows - in production or through serious side projects
About your role
Reliability & Infrastructure
  • Own the reliability, scalability, and observability of our infrastructure across GCP and Vercel
  • Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards
  • Manage IAM, service accounts, and security best practices across our cloud environment
  • Participate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements
Core Infrastructure & State
  • Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflows
  • Build and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments
  • Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution
  • Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)
  • Automate infrastructure provisioning, deployments, and operational workflows
AI & Next-Gen Tooling: Agent Ops
  • Build agent operations infrastructure that enables AI agents to run safely and reliably in production
  • Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry
  • Help define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loop
  • Own visibility into AI usage, reliability, and spend as our agent footprint scales
Cross-functional Impact
  • Partner closely with engineering and product teams to maintain reliability without slowing development velocity
  • Act as a force multiplier across the team — helping engineers ship faster and more safely
About your skillsMust haves
  • 4+ years in SRE, DevOps, or infrastructure/platform engineering
  • Strong, hands-on experience with a major cloud platform (preferable GCP)
  • Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)
  • Solid understanding of IAM, security, and cloud best practices
  • Experience with observability tools like Datadog
  • Familiarity with Node.js environments
  • AI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.
  • Experience building or orchestrating AI/agent workflows (work or serious personal projects)
  • High ownership, strong curiosity, and a bias toward action
Nice to haves
  • Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls
  • Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plus
  • Hands-on experience with GCP Workflows for orchestration
  • Familiarity with Prisma, pgbouncer, and PostgreSQL connection pooling
  • Experience with Vercel deployment and edge computing
  • Familiarity with the k8s ecosystem
  • Familiarity with Redis and BullMQ
  • Understanding of SOC 2 compliance requirements and implementation
  • Previous experience in a high-growth startup environment
  • Previous backend engineering experience to bridge the gap between infrastructure and code
About the stack
  • Cloud: Google Cloud Platform (GCP)
  • Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes
  • Database: PostgreSQL with pgbouncer, Prisma
  • Observability: Datadog
  • Runtime: Node.js, TypeScript
  • Deployment: Vercel, GCP
  • Frontend: React, Next.js, TypeScript
  • Backend: TypeScript, PostgreSQL, Next.js
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Software Engineer, Infrastructure & Reliability
Software Engineer, Infrastructure & Reliability

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Software Engineer - Infrastructure
Software Engineer - Infrastructure

Emergentlabsinc • San Francisco (CA)

On-site
USD 110,000 - 150,000
401(k)
Health, dental, and vision insurance
Unlimited Paid Time Off
+1
Agent Builder
Agent Builder

Aisle • New York (NY)

On-site
USD 140,000 - 180,000
Staff AI Engineer
Staff AI Engineer

HubSync • United States

On-site
USD 130,000 - 160,000
Principal AI Software Engineer
Principal AI Software Engineer

NICE • Seattle (WA)

On-site
USD 140,000 - 180,000
Forward Deployment Engineer
Forward Deployment Engineer

TechDigital Group • Bloomfield (CT)

On-site
USD 120,000 - 190,000
AI Engineer, Agent Builder (Remote).
AI Engineer, Agent Builder (Remote).

Catalyst Wayfare • United States

Remote
USD 120,000 - 170,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

hardrockdigital • United States

Hybrid
USD 120,000 - 160,000
Competitive pay and benefits
Flexible vacation allowance
Startup culture with global brand support
+1
Software Engineer - AI Agent Platforms
Software Engineer - AI Agent Platforms

Workato • San Francisco (CA)

On-site
USD 140,000 - 200,000
Sr. Full Stack Engineer - AI Forward
Sr. Full Stack Engineer - AI Forward

CarParts.com • Long Beach (CA)

On-site
USD 140,000 - 210,000