Site Reliability Engineer

BAM Ventures

New York (NY)

On-site

USD 180,000 - 240,000

Full time

8 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Aisle is building a real-time control layer for physical retail, enabling brands to trigger actions and optimize outcomes as they happen. We are scaling infrastructure to support AI-driven production paths and high-volume event processing in a fast-moving startup environment.

We are looking for a Platform/Infrastructure engineer with 4+ years in SRE or DevOps, strong GCP experience, and hands-on work with serverless systems, IAM, and observability.

Qualifications

  • 4+ years in SRE, DevOps, or infrastructure/platform engineering.
  • Strong hands-on experience with a major cloud platform (GCP preferred).
  • Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.).
  • Solid understanding of IAM, security, and cloud best practices.
  • Familiarity with observability tools like Datadog and Node.js environments.
  • AI is part of your daily engineering workflow and you use AI tools to improve velocity and reliability.
  • Experience building or orchestrating AI/agent workflows in production or serious side projects.
  • High ownership, curiosity, and bias toward action.

Responsibilities

  • Own reliability, scalability, and observability of infrastructure across GCP and Vercel.
  • Design and implement monitoring and alerting with Datadog, including dashboards.
  • Manage IAM, service accounts, and security across cloud environments.
  • Participate in on-call rotation and post-mortems; turn learnings into systemic improvements.
  • Stabilize core infrastructure under heavy concurrent load and keep stateful systems robust.

Skills

SRE/DevOps
GCP
Serverless
IAM/Security
Datadog
Node.js
AI in DevOps
AI Agent Workflows
Ownership

Tools

GCP Workflows
PostgreSQL
Prisma
pgbouncer
Vercel
Kubernetes
Redis
BullMQ
OpenClaw

Job description

What is Aisle?

Physical retail has always been a black box - brands ship product, run static promotions, and wait weeks (or months) to understand what actually happened.

Aisle builds a real-time control layer on top of physical retail. We connect shopper actions, incentives, and outcomes into a continuous feedback loop, so brands can trigger, measure, and optimize retail performance while it’s happening - not after the fact.

This is the foundation for autonomous retail execution: systems that don’t just report on what happened, but actively decide what should happen next - which offers to show, when to show them, and how to maximize outcomes in real time.

Today, 1600+ brands and 5M+ consumers use Aisle to drive measurable results in brick-and-mortar. Under the hood, that means high-volume event ingestion, real-time decisioning, and infrastructure that has to be both reliable and low-latency to influence behavior in the moment.

We’ve gotten here with a small, fast-moving team, and now we’re scaling the system to handle significantly more volume, complexity, and automation - including AI systems that are part of the production path, not just offline analysis.

You’ll win here if…
  1. You enjoy communicating across technical and non-technical teams looking to build shared understanding that can turn into collaborative actio
  2. Can solve problems in stages with appropriate level of urgency when needed: immediate triage, temporary patch, and long term fix that increases durability of the system
  3. You’re comfortable moving quickly, shipping improvements, and iterating in production
  4. AI is a core part of how you work - you’ve gone well beyond basic copilots and actively use AI tools to design, debug, and ship systems faster
  5. You treat AI agents as chaotic, untrusted components - your instinct is isolation, rate limiting, and permissions scoping, etc. (gist of harness engineering)
  6. You have experience building or orchestrating AI/agent workflows - in production or through serious side projects
About your role
Reliability & Infrastructure
  • Own the reliability, scalability, and observability of our infrastructure across GCP and Vercel
  • Design and implement monitoring and alerting in Datadog, including monitors-as-code and product-level dashboards
  • Manage IAM, service accounts, and security best practices across our cloud environment
  • Participate in on-call rotation, incident response, and post-mortems — and turn learnings into systemic improvements
Core Infrastructure & State
  • Stabilize core infrastructure and stateful systems under heavy concurrent load from both consumer traffic and agent-driven workflows
  • Build and maintain our event-driven GCP infrastructure, including Cloud Functions, Pub/Sub, Workflows, and GKE/Kubernetes deployments
  • Investigate performance inconsistencies (e.g., identical replicas behaving differently) and drive durable root-cause resolution
  • Stabilize and simplify parts of the stack that create friction today (e.g., Prisma, connection management, cron workflows)
  • Automate infrastructure provisioning, deployments, and operational workflows
AI & Next-Gen Tooling: Agent Ops
  • Build agent operations infrastructure that enables AI agents to run safely and reliably in production
  • Develop secure isolation layers, LLM gateways, workflow monitoring, anomaly detection, and production telemetry
  • Help define CI/CD for an agent-driven engineering environment - including how code gets shipped, reviewed, validated, and deployed when agents are in the loop
  • Own visibility into AI usage, reliability, and spend as our agent footprint scales
Cross-functional Impact
  • Partner closely with engineering and product teams to maintain reliability without slowing development velocity
  • Act as a force multiplier across the team — helping engineers ship faster and more safely
About your skillsMust haves
  • 4+ years in SRE, DevOps, or infrastructure/platform engineering
  • Strong, hands-on experience with a major cloud platform (preferable GCP)
  • Experience with event-driven or serverless systems (Pub/Sub, Cloud Functions, etc.)
  • Solid understanding of IAM, security, and cloud best practices
  • Experience with observability tools like Datadog
  • Familiarity with Node.js environments
  • AI is part of your daily engineering workflow you use AI tools thoughtfully to improve engineering velocity, reliability, and operational excellence beyond basic code generation.
  • Experience building or orchestrating AI/agent workflows (work or serious personal projects)
  • High ownership, strong curiosity, and a bias toward action
Nice to haves
  • Experience with harness engineering patterns for untrusted agents, including isolation, rate limiting, permission scoping, auditability, and safety controls
  • Familiarity with OpenClaw, agent orchestration frameworks, or LLM gateway patterns is a plus
  • Hands‑on experience with GCP Workflows for orchestration
  • Familiarity with Prisma, pgbouncer, and PostgreSQL connection pooling
  • Experience with Vercel deployment and edge computing
  • Familiarity with the k8s ecosystem
  • Familiarity with Redis and BullMQ
  • Understanding of SOC 2 compliance requirements and implementation
  • Previous experience in a high-growth startup environment
  • Previous backend engineering experience to bridge the gap between infrastructure and code
About the stack
  • Cloud: Google Cloud Platform (GCP)
  • Infrastructure: Cloud Functions, Pub/Sub, Workflows, Kubernetes
  • Database: PostgreSQL with pgbouncer, Prisma
  • Observability: Datadog
  • Runtime: Node.js, TypeScript
  • Deployment: Vercel, GCP
  • Frontend: React, Next.js, TypeScript
  • Backend: TypeScript, PostgreSQL, Next.js
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Aisle • New York (NY)

On-site
USD 140,000 - 190,000
Agent Builder
Agent Builder

BAM Ventures • New York (NY)

On-site
USD 140,000 - 190,000
Software Engineer, Infrastructure & Reliability
Software Engineer, Infrastructure & Reliability

crewAI, Inc. • San Francisco (CA)

On-site
USD 140,000 - 210,000
Forward Deployment Engineer
Forward Deployment Engineer

TechDigital Group • Bloomfield (CT)

On-site
USD 120,000 - 190,000
Agent Builder
Agent Builder

Aisle • New York (NY)

On-site
USD 140,000 - 180,000
Senior Software Engineer, Infrastructure
Senior Software Engineer, Infrastructure

Talanto • Northern (KY)

Hybrid
USD 1,786,000 - 2,120,000
AI Engineer, Agent Builder (Remote).
AI Engineer, Agent Builder (Remote).

Catalyst Wayfare • United States

Remote
USD 120,000 - 170,000
Software Engineer, Agent
Software Engineer, Agent

Vercel • New York (NY)

Hybrid
USD 232,000 - 348,000
Equity
Mentorship & events
Flexible time off
+1
Infrastructure Engineer
Infrastructure Engineer

Pursuit Talent Advisory • Austin (TX)

On-site
USD 150,000 - 210,000
Staff AI Engineer
Staff AI Engineer

HubSync • United States

On-site
USD 130,000 - 160,000