A complete application in a minute — tailored resume and cover letter, ready to send.
Endear, a modern CRM for retail teams, is hiring a Senior Site Reliability Engineer to become Endear’s first dedicated reliability hire. This hands-on builder role focuses on reducing on-call burden by shaping observability, alerting policy, capacity planning, and runbook standards across services.
You will own incident response, drive root-cause fixes, and partner with engineers on database, queue, and event-processing reliability.
At Endear, we’re building a modern CRM for retail teams-starting with the frontline. Our software helps sales associates have more personal, effective customer conversations through AI-powered tools that drive measurable revenue.
Despite retail being significantly larger than eCommerce, most software overlooks the in-store experience. Endear helps brands turn real customer relationships into growth through intuitive software and thoughtful design.
As Endear grows, we are investing in the reliability and platform systems that keep our product fast, resilient, and easy for engineering teams to operate.
We’re hiring a Senior Site Reliability Engineer to become Endear’s first dedicated reliability hire. This is a hands-on builder role for someone who wants to solve the underlying systems problems that create on-call burden-not simply respond to pages.
You will own observability standards, alerting policy and routing, capacity planning method, incident response process and runbook standards, reliability requirements in project scoping, and tooling selection. You will investigate recurring incidents, drive root cause fixes, and partner with engineers on database, queue, and event processing reliability. Engineers keep instrumenting their own services, responding to their own alerts, and writing their own postmortems
You will also help establish a nearshore triage layer for routine, well-understood issues, so product engineers can spend more time building product and less time on operational interruptions.
Build a clear view of Endear’s highest-impact reliability risks, recurring incidents, and on-call pain points.
Establish a prioritized reliability backlog and drive root-cause fixes for the most important issues.
Make on call cost measurable, so every later claim is arguable from data.
Propose the observability standard and the alerting policy, agreed across engineering rather than published, and land them in one team's real workflow.
Improve alert quality, severity definitions, escalation paths, and runbooks for common incidents.
Strengthen observability, queue health, database capacity planning, and operational readiness ahead of peak retail periods.
Create the foundation for a nearshore triage process for low-priority, repeatable issues.
Make on-call materially quieter and less disruptive for product engineers.
Build scalable systems for observability, alerting, incident response, database reliability, and queue/event-processing health.
Make reliability requirements part of how work gets scoped, not something discovered after launch.
Own and improve the nearshore triage relationship, playbooks, and escalation process.
Help establish the technical roadmap and future resourcing plan for Endear’s broader platform and reliability function.
Have deep hands-on experience with Kubernetes, production databases, and event-driven systems.
Have operated and improved high-volume production systems with meaningful reliability, performance, and data-scale requirements.
Like teaching. You have run on call training, written the observability guide, or paired with product engineers to instrument their services.
Enjoy finding root causes, fixing repeat incidents, and building tooling that makes engineers’ lives easier.
Have experience with observability, alerting, incident response, capacity planning, and operational runbooks.
Can work effectively as a senior IC: owning complex technical work directly while coordinating across teams.
Are comfortable in a lean environment where priorities move quickly and you will need to make practical trade-offs.
Bring experience from a scaling, mid-size company rather than only an early-stage startup or hyperscaler environment.
Have GCP or ClickHouse experience, which are strong pluses.
You’ll partner with:
Ini Becker, Engineering Manager, and JP Grace, CTO: Align on reliability priorities, technical risks, and the roadmap for improving on call health.
Engineering team: Partner on root-cause fixes, platform improvements, and architecture decisions that improve reliability.
Nearshore triage partner: Build and maintain playbooks, escalation paths, and expectations for routine incident handling.
Product and Support: Help ensure issues are surfaced, prioritized, and resolved with the right level of urgency.
Endear is a lean, remote team where individuals have broad ownership. This role will directly shape how the company handles production reliability as it grows.
Recruiter screen - 30 minutes
Behavioral interview with Hiring Manager - 60 minutes
Technical panel with Engineering - 60 minutes
Final conversation with CTO + Co-Founders
Offer
Base salary: $140,000-180,000
Fully remote, U.S.-based role
Comprehensive healthcare, including medical, dental, and vision, plus a 401(k) plan
Monthly stipend for co-working and home-office setup
Flexible PTO and unlimited vacation
Opportunity to build Endear’s first dedicated reliability function from the ground up