Senior Site Reliability Engineer (Performance and Scalability)

JobCubby

Turkey

On-site

TRY 600,000 - 1,200,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Immediate impact
Top compensation
Regional talent

Job summary

JobCubby is seeking a senior SRE to own reliability across our large-scale platform, handling tens of thousands of requests per minute. You will shape capacity planning, implement autoscaling, caching, queues, and graceful degradation to survive campaign spikes.

You will lead load testing, define SLOs and error budgets, manage the observability stack, and Harden Postgres and AWS. Collaborate with engineers early to multiply their scaling and resilience in a fast-growing environment.

Qualifications

  • 5+ years in SRE, platform, or backend engineering with ownership of large-scale systems.
  • Experience scaling systems against real traffic spikes.
  • Deep AWS experience and Postgres performance tuning.
  • Fluency with observability tools and IaC; scripting in Go or TypeScript.
  • Calm, methodical incident response and cross-team communication.

Responsibilities

  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.
  • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.
  • Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.
  • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
  • Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.
  • Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.

Skills

SRE/Backend
Capacity planning
AWS expertise
Postgres tuning
Observability
IaC
Go
TypeScript
Incident comms

Tools

Postgres
AWS
Observability tooling
IaC tooling

Job description

What you'll do


  • Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.

  • Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.

  • Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.

  • Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.

  • Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.

  • Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.


What you'll bring


  • 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.

  • A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted.

  • Deep AWS experience and a solid grasp of Postgres performance and scaling.

  • Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar.

  • A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep.



  • Immediate, large-scale impact on a high-growth business

  • Top-of-the-market compensation packages

  • Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Site Reliability Engineering Manager
Site Reliability Engineering Manager

n11 • Fatih

On-site
TRY 800,000 - 1,000,000
Site Reliability Engineer
Site Reliability Engineer

MetLife México • Fatih

On-site
TRY 320,000 - 540,000
Private health insurance
Pension plan
Work from home allowance
+1
Site Reliability Engineer
Site Reliability Engineer

OBSS • Fatih

Hybrid
TRY 1,818,000 - 2,728,000
Flexible working arrangements
Training programs
Certifications
+1
Load Test Engineer
Load Test Engineer

Magic Media • Turkey

Hybrid
TRY 450,000 - 750,000
Ownership of platform
Direct contact with founders
Competitive compensation
+2
Software Engineer, Site Reliability
Software Engineer, Site Reliability

The Consensus • Turkey

On-site
TRY 600,000 - 1,000,000
Interesting and challenging work
Learning and growth opportunities
Regular team events and offsites
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • Turkey

On-site
TRY 600,000 - 900,000
Private health insurance
Continuous upskilling & development
English courses
+1
Head of Platform & DevOps
Head of Platform & DevOps

RedCloud • Fatih

On-site
TRY 2,272,000 - 3,182,000
25 Days Annual leave
Enhanced Company Pension (Matched up to 5%)
Healthcare Cashplan
+3
SRE Leadership Manager - Scale & Reliability
SRE Leadership Manager - Scale & Reliability

n11 • Fatih

On-site
TRY 800,000 - 1,000,000
Senior SRE: Data Reliability, AI-Driven Infra (Remote)
Senior SRE: Data Reliability, AI-Driven Infra (Remote)

Embedded Shishya • Turkey

Remote
TRY 2,684,000 - 5,100,000
Platform Reliability & Scale Lead (SRE)
Platform Reliability & Scale Lead (SRE)

JobCubby • Turkey

On-site
TRY 600,000 - 1,200,000
Immediate impact
Top compensation
Regional talent