Staff Site Reliability Engineer

Jobgether

Turkey

Remote

TRY 4,382,000 - 7,303,000

Full time

5 days ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Fully remote work

Job summary

Jobgether is seeking a Staff Site Reliability Engineer based in Turkey for a fully remote, globally distributed engineering org. You will be the first dedicated SRE, embedding reliability practices across teams and shaping AI-enabled incident response and observability.

You will own SLIs/SLOs, incident leadership, and reliability culture while partnering with architects and leadership to design resilient systems with strong engineering standards.

Qualifications

  • 10+ years in engineering with 3+ years in SRE or reliability-focused leadership across multiple teams.
  • Own reliability at platform or organizational level, not just individual services.
  • Experience designing and deploying SLIs, SLOs, and error budgets.
  • Strong incident leadership for high-severity incidents and postmortems.
  • Deep understanding of distributed-system failure modes and resilience patterns.
  • Hands-on with Kubernetes, AWS, and modern observability platforms (Datadog).
  • Ability to read and write production code in Go, TypeScript, or similar.
  • Influence teams and change engineering practices without formal authority.
  • Coaching skills to grow engineers into reliability owners.
  • Excellent written and verbal communication for engineers and executives.
  • Preference for asynchronous decision-making and clear technical docs.
  • Experience using AI tools for incident investigation, telemetry, and tooling.
  • Know-how to structure operational data and runbooks for safe AI-assisted ops.
  • Pragmatic reliability balance of risk, investment, speed, and business priorities.
  • Experience with fraud detection, identity, payments, or real-time/adversarial environments is a plus.
  • Multi-region or cell-based architectures experience is a plus.
  • Experience with Elasticsearch, Redis, DynamoDB, or Kafka at scale and their failure modes.
  • Familiarity with FinOps and cloud cost-reliability trade-offs is a plus.
  • Authorized to work from the hiring location; visa sponsorship not provided.

Responsibilities

  • Define and implement SLIs and SLOs for critical production paths and tie them to engineering decisions.
  • Introduce error budgets as a framework balancing reliability with delivery.
  • Maintain reliability metrics used by leadership to guide investments.
  • Own the incident management lifecycle from detection to postmortems.
  • Improve alerting, anomaly detection, escalation, and tooling with infra teams.
  • Lead reliability assessments for high-risk changes and new services.
  • Introduce chaos engineering practices to validate safety margins.
  • Collaborate with engineering teams to tackle complex reliability challenges.
  • Coach staff and lead engineers to become reliability advocates.
  • Develop lightweight operational standards for runbooks, on-call, and change safety.
  • Embed reliability and failure tolerance into system design with architects.
  • Be hands-on during incidents and build tooling, dashboards, and reference implementations.
  • Promote AI use for incident investigation, telemetry analysis, and tooling.
  • Structure operational data so AI agents can safely diagnose and operate.
  • Contribute production fixes and improvements directly through code and infra changes.

Skills

Incident leadership
Distributed systems
Communication
Coaching/mentoring
Influencing without authority
Asynchronous decision making
AI tooling for reliability

Tools

Kubernetes
AWS
Datadog
Terraform
Elasticsearch
Redis
DynamoDB
Kafka

Job description

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Turkey.

This is a high-impact reliability leadership role within a fully remote engineering organization operating globally.
You will be the first dedicated SRE, helping establish reliability practices across multiple engineering teams and critical production systems.
The role combines hands-on engineering with organization-wide influence, covering observability, incident response, operational readiness, and resilience.
You will work closely with engineering leadership, infrastructure specialists, architects, and product teams to make reliability measurable and actionable.
A major focus will be embedding SRE principles into engineering culture rather than simply owning individual services.
You will also help shape how AI is used for incident investigation, operational tooling, observability, and safe system operations.
The position offers substantial autonomy to define standards, coach engineers, and build practices that scale with the organization.

Accountabilities
  • Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.

  • Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.

  • Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.

  • Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.

  • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.

  • Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.

  • Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.

  • Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.

  • Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.

  • Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.

  • Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.

  • Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.

  • Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.

  • Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.

  • Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.

Requirements
  • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.

  • Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.

  • Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.

  • Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.

  • Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.

  • Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.

  • Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.

  • Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.

  • Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.

  • Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.

  • Strong preference for asynchronous, documented decision-making and clear technical communication.

  • Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.

  • Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.

  • Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.

  • Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.

  • Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.

  • Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.

  • Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.

  • Must be authorized to work from the hiring location; visa sponsorship is not provided.

Benefits
  • Fully remote working environment.

  • Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.

  • High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.

  • Opportunity to work across multiple engineering teams and critical production systems.

  • Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.

  • Opportunity to shape AI-assisted reliability practices and the future of production operations.

  • Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.

  • Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.

  • For US-based employees, the stated cash compensation range is$177,000-$240,000 USD, with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.

  • Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Remote Staff Site Reliability Engineer (First Dedicated SRE)
Remote Staff Site Reliability Engineer (First Dedicated SRE)

Jobgether • Turkey

Remote
TRY 4,382,000 - 7,303,000
Fully remote work
Site Reliability Engineer
Site Reliability Engineer

MetLife México • Fatih

On-site
TRY 320,000 - 540,000
Private health insurance
Pension plan
Work from home allowance
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • Turkey

On-site
TRY 4,859,000 - 6,803,000
Private health insurance
Professional development
English courses
+2
Site Reliability Engineer
Site Reliability Engineer

OBSS • Fatih

Hybrid
TRY 1,818,000 - 2,728,000
Flexible working arrangements
Training programs
Certifications
+1
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • Turkey

On-site
TRY 600,000 - 900,000
Private health insurance
Continuous upskilling & development
English courses
+1
Senior SRE: Data Reliability, AI-Driven Infra (Remote)
Senior SRE: Data Reliability, AI-Driven Infra (Remote)

Embedded Shishya • Turkey

Remote
TRY 2,684,000 - 5,100,000
Senior Site Reliability Engineer (Performance and Scalability)
Senior Site Reliability Engineer (Performance and Scalability)

Digital Zone • Turkey

On-site
TRY 180,000 - 300,000
Top salary package
Regional talent
Impactful work
Site Reliability Engineer
Site Reliability Engineer

Akakçe • Çankaya

On-site
TRY 300,000 - 540,000
Senior SRE: AI-Driven Cloud Reliability & Kubernetes
Senior SRE: AI-Driven Cloud Reliability & Kubernetes

Sezzle • Turkey

On-site
TRY 2,684,000 - 5,100,000
Remote Site Reliability Engineer – AI-Driven Incident Response
Remote Site Reliability Engineer – AI-Driven Incident Response

Storyteller • Turkey

Remote
TRY 1,051,000 - 1,752,000