Senior Site Reliability Engineer (Performance and Scalability)

Digital Zone

Turkey

On-site

TRY 180,000 - 300,000

Full time

8 days ago
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Top salary package
Regional talent
Impactful work

Job summary

Digital Zone in Turkey is expanding its platform and seeks a seasoned SRE/Platform engineer to scale services across Go, TypeScript, and PHP/Laravel environments. This role focuses on enabling teams through scalable infrastructure and reliable practices.

You will own capacity planning, autoscaling, caching, and the incident lifecycle, driving improvements across the observability stack and automation.

Qualifications

  • 5+ years in SRE, platform, or backend engineering with ownership of large-scale systems.
  • Proven experience scaling systems during traffic spikes and running load/failure tests.
  • Fluent with AWS, Postgres performance, and observability tooling.

Responsibilities

  • Build scalability foundations: capacity planning, autoscaling, caching, queues, and graceful degradation.
  • Establish load and failure testing as standard practice across teams, with frameworks and runbooks.
  • Own SLOs, error budgets, and the observability stack across TypeScript, Go, and PHP/Laravel services; instrument for scale.
  • Harden Postgres and AWS infrastructure for performance and availability; automate toil via IaC.
  • Lead incident response and blameless postmortems; drive systemic fixes upstream into design and campaign planning.
  • Collaborate with engineering teams early on capacity and resilience to enable self-sufficiency in scaling their systems.

Skills

SRE / platform ownership
Scale systems under traffic spikes
AWS experience
Postgres performance tuning
Go / TS scripting

Tools

Postgres
AWS
IaC tooling

Job description

Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.

What you'll do
  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load
  • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results
  • Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale
  • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC
  • Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on
  • Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems
What you'll bring
  • 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute
  • A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted
  • Deep AWS experience and a solid grasp of Postgres performance and scaling
  • Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar
  • A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep
Benefits
  • Immediate, large-scale impact on a high-growth business
  • Top-of-the-market compensation packages
  • Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Staff Software Engineer
Staff Software Engineer

Digital Zone • Turkey

On-site
TRY 600,000 - 1,200,000
Top-of-the-market compensation
Impact on high-growth business
Work with regional talent
Senior Software Engineer - Backend
Senior Software Engineer - Backend

Digital Zone • Turkey

On-site
TRY 600,000 - 900,000
Immediate impact
Top compensation
Regional talent peers
Senior SRE: Scale & Enable Platform Reliability
Senior SRE: Scale & Enable Platform Reliability

Digital Zone • Turkey

On-site
TRY 180,000 - 300,000
Top salary package
Regional talent
Impactful work
Staff Site Reliability Engineer
Staff Site Reliability Engineer

Jobgether • Turkey

Remote
TRY 4,382,000 - 7,303,000
Fully remote work
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • Turkey

On-site
TRY 4,859,000 - 6,803,000
Private health insurance
Professional development
English courses
+2
Site Reliability Engineer
Site Reliability Engineer

MetLife México • Fatih

On-site
TRY 320,000 - 540,000
Private health insurance
Pension plan
Work from home allowance
+1
Site Reliability Engineer
Site Reliability Engineer

OBSS • Fatih

Hybrid
TRY 1,818,000 - 2,728,000
Flexible working arrangements
Training programs
Certifications
+1
Customer Success Engineer
Customer Success Engineer

Medianova • Şişli

On-site
TRY 240,000 - 360,000
Head of Platform & DevOps
Head of Platform & DevOps

RedCloud • Fatih

On-site
TRY 2,272,000 - 3,182,000
25 Days Annual leave
Enhanced Company Pension (Matched up to 5%)
Healthcare Cashplan
+3
Senior DevOps Engineer
Senior DevOps Engineer

Dgpays • Fatih

Hybrid
TRY 350,000 - 550,000
Hybrid working
Private health insurance
Educational materials access
+2