Get more replies from employers
Send a job-specific resume in minutes.
Alcor is hiring a Site Reliability Engineer to own production reliability for a real-time platform where uptime and latency are the product. You will manage SLOs, incident response, on-call rotations, and production scaling with a focus on blameless postmortems and fast recovery.
Expect a startup tempo with weekly deploys, 1-week sprints, and a culture that values fault-tolerance and AI-assisted tooling to boost velocity while keeping reliability at the center of every decision.
The role. Own production reliability for a real-time platform where uptime and latency ARE the product — voice, desktop, intelligence, and AI combined; an agent mid-call can't wait for a retry. First SRE hired immediately (Day 0–14) for production scaling and SLO ownership; a second joins at the start of Phase 3 for 24/7 coverage. Pairs with C1 Platform Foundation on observability and tenancy isolation. Startup environment: weekly deploys, 1-week sprints, fail fast, move forward — reliability engineering at that speed, not against it.
What you'll own. SLOs and error budgets per tenant / service · incident response and blameless postmortems · production scaling and capacity · observability depth (p50/p95/p99 per event hop) · uptime as a personal mission · on-call rotation with DevOps · your committed timelines.
Who you are. Self-starter, grit, show-me mentality — you prove reliability with dashboards and drills, not assertions. A ways-to-YES engineer: weekly deploys are the heartbeat and your job is making them safe, never slowing them. You love new technology, adapt fast when the stack changes under you, use AI tools daily to multiply velocity, and consider yourself exceptional. Calm in an incident, relentless after it. Team player who likes winning.
Requirements