Senior Site Reliability Engineer (Performance and Scalability)

Digital Zone

Poland

On-site

PLN 320,000 - 540,000

Full time

2 days ago
Be an early applicant
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Top-of-the-market compensation
Work with regional talent

Job summary

DigitalZone is seeking an experienced SRE/Platform Engineer to scale our platform and prepare it for high campaign traffic. You’ll own capacity planning, observability, and reliability across Go, TypeScript, and PHP/Laravel services, with a focus on AWS and Postgres performance.

You will establish load testing standards, drive incident response improvements, and partner with engineering teams to make scaling self-sufficient. A calm, methodical approach is valued in this high-growth environment.

Qualifications

  • 5+ years in SRE, platform, or backend engineering with ownership of large-scale systems.
  • Proven track record scaling systems through real traffic spikes.
  • Deep AWS experience and Postgres performance expertise.
  • Fluency with observability tooling and infrastructure-as-code.

Responsibilities

  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation for campaign spikes.
  • Establish load and failure testing as a standard engineering practice with frameworks and runbooks.
  • Own SLOs, error budgets, and the observability stack across TypeScript, Go, and PHP/Laravel services.
  • Harden Postgres and AWS infrastructure for performance and availability, reducing toil via automation and IaC.
  • Lead incident response and blameless postmortems, driving systemic fixes upstream.
  • Partner with engineering teams to help them scale their own systems.

Skills

SRE
Platform engineering
Backend engineering
Observability tooling
IaC
Go/TypeScript

Tools

Postgres
AWS
Infrastructure as Code

Job description

Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.

What you'll do
  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load
  • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results
  • Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale
  • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC
  • Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on
  • Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems
What you'll bring
  • 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute
  • A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted
  • Deep AWS experience and a solid grasp of Postgres performance and scaling
  • Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar
  • A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep
Benefits
  • Immediate, large-scale impact on a high-growth business
  • Top-of-the-market compensation packages
  • Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform SRE: Scale, Observability & Enablement
Platform SRE: Scale, Observability & Enablement

Digital Zone • Poland

On-site
PLN 320,000 - 540,000
Top-of-the-market compensation
Work with regional talent
Staff Security Engineer
Staff Security Engineer

Digital Zone • Poland

On-site
PLN 260,000 - 360,000
Top compensation
High-growth environment
Collaborative team
Site Reliability Engineer (SRE)
Site Reliability Engineer (SRE)

XM • Poland

On-site
PLN 180,000 - 240,000
Competitive remuneration package
International training opportunities
Confidential recruitment
Engineering Manager
Engineering Manager

Digital Zone • Poland

On-site
PLN 480,000 - 900,000
Impactful role
Competitive compensation
Prestigious team
Site Reliability Engineer
Site Reliability Engineer

Balyasny Asset Management L.P. • Warszawa

On-site
PLN 180,000 - 300,000
DevOps Engineer
DevOps Engineer

Joy Studios • Warszawa

On-site
PLN 260,000 - 360,000
Site Reliability Engineer
Site Reliability Engineer

SIX • Warszawa

On-site
PLN 260,000 - 380,000
Service Manager & Site Reliability Consultant
Service Manager & Site Reliability Consultant

GFT Technologies Poland • Wrocław

On-site
PLN 130,000 - 210,000
Service Manager & Site Reliability Consultant
Service Manager & Site Reliability Consultant

GFT Technologies Poland • Łódź

On-site
PLN 180,000 - 320,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind • Kraków

Remote
PLN 180,000 - 240,000
Private healthcare and insurance
Multisport card
Language classes
+3