Site Reliaibility engineer

Alcor

Kraków

On-site

PLN 250,000 - 380,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Alcor is hiring a Site Reliability Engineer to own production reliability for a real-time platform where uptime and latency are the product. You will manage SLOs, incident response, on-call rotations, and production scaling with a focus on blameless postmortems and fast recovery.

Expect a startup tempo with weekly deploys, 1-week sprints, and a culture that values fault-tolerance and AI-assisted tooling to boost velocity while keeping reliability at the center of every decision.

Qualifications

  • 8+ years operating production systems at scale.
  • Strong Go or Python – you automate reliability.
  • Deep on event-driven and real-time systems reliability.
  • Strong monitoring and uptime mindset with proactive alerting.
  • Good networking understanding (TCP/UDP, TLS, WebSocket, DNS, load balancing).
  • GCP at scale; multi-cloud literacy a plus and multi-tenancy experience.

Responsibilities

  • Own SLOs and error budgets per tenant / service.
  • Incident response and blameless postmortems.
  • Production scaling and capacity planning.
  • Observability depth (p50/p95/p99 per event hop).
  • On-call rotation with DevOps; communicate outages clearly.
  • Deploy-safety collaboration with DevOps and automated rollbacks.

Skills

SRE & incident response
Go
Python
Cloud native / GCP
Observability & monitoring
Automation / runbooks
Systems at scale
Networking fundamentals
Chaos engineering
On-call ownership

Tools

NATS
WebSocket
Streaming pipelines
CI/CD pipelines

Job description

Site Reliability Engineer (SRE)

The role. Own production reliability for a real-time platform where uptime and latency ARE the product — voice, desktop, intelligence, and AI combined; an agent mid-call can't wait for a retry. First SRE hired immediately (Day 0–14) for production scaling and SLO ownership; a second joins at the start of Phase 3 for 24/7 coverage. Pairs with C1 Platform Foundation on observability and tenancy isolation. Startup environment: weekly deploys, 1-week sprints, fail fast, move forward — reliability engineering at that speed, not against it.


What you'll own. SLOs and error budgets per tenant / service · incident response and blameless postmortems · production scaling and capacity · observability depth (p50/p95/p99 per event hop) · uptime as a personal mission · on-call rotation with DevOps · your committed timelines.


Who you are. Self-starter, grit, show-me mentality — you prove reliability with dashboards and drills, not assertions. A ways-to-YES engineer: weekly deploys are the heartbeat and your job is making them safe, never slowing them. You love new technology, adapt fast when the stack changes under you, use AI tools daily to multiply velocity, and consider yourself exceptional. Calm in an incident, relentless after it. Team player who likes winning.


Requirements


  • 8+ years operating production systems at scale; owns SLOs, error budgets, incident command.

  • Strong Go or Python — you automate reliability, you don't toil at it. Everything you build is code: runbooks execute, remediation is automatic, toil trends to zero.

  • Deep on event-driven and real-time systems reliability — NATS-class buses, WebSocket fleets, streaming pipelines — and the failure physics underneath: state, race conditions, locking, ordering, back-pressure, cascading load. You've debugged these in production.

  • Strong monitoring and uptime mindset — metrics, logs, traces wired to alerting that catches it before the customer does; you know the difference between a noisy alert and a real signal.

  • Good networking understanding — protocols and how they work (TCP/UDP, TLS, WebSocket, DNS, load balancing); RTP/SIP a strong plus for our media paths.

  • GCP at scale; multi-cloud literacy a plus. Multi-tenancy isolation experience a strong plus.

  • Capacity modeling and load testing partnership with QA — find the knee of the curve before customers do.

  • Chaos engineering — failure injection as routine practice; prove graceful degradation, don't assume it.

  • Deploy-safety partnership with DevOps — canary analysis, automatic rollback triggers, error-budget-driven release gates.

  • AI-aware reliability — monitoring model latency, drift, and cost as production signals, not just CPU and memory.

  • Incident communication craft — clear, fast, blameless; execs and customers get truth at the right altitude.

  • A master debugger of production — reads the trace, the metric, the flame graph, and sees it; narrows an incident to the service, the deploy, the event.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineer
Site Reliability Engineer

Balyasny Asset Management L.P. • Warszawa

On-site
PLN 180,000 - 300,000
Senior Devops engineer
Senior Devops engineer

Alcor • Kraków

On-site
PLN 180,000 - 320,000
Senior Systems Site Reliability Engineer, B2B
Senior Systems Site Reliability Engineer, B2B

Jobtailor • Poland

On-site
PLN 180,000 - 320,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Grid Dynamics • Województwo pomorskie

On-site
PLN 80,000 - 120,000
Medical insurance
Sports benefits
Professional development opportunities
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Grid Dynamics • Kraków

On-site
PLN 254,000 - 340,000
Medical insurance
Sports benefits
Professional development opportunities
+2
Site Reliability Engineer
Site Reliability Engineer

Caspian One • Warszawa

On-site
PLN 180,000 - 280,000
Site Reliability Engineer
Site Reliability Engineer

Fáilte Ireland • Poland

On-site
PLN 180,000 - 280,000
Equity program
Cloud SRE — Reliability, Observability & Equity in InsurTech
Cloud SRE — Reliability, Observability & Equity in InsurTech

Fáilte Ireland • Poland

On-site
PLN 180,000 - 280,000
Real-Time Site Reliability Engineer — Production Uptime & Scale
Real-Time Site Reliability Engineer — Production Uptime & Scale

Alcor • Kraków

On-site
PLN 250,000 - 380,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Województwo małopolskie

On-site