Site Reliability Engineering Manager

n11

Fatih

On-site

TRY 800,000 - 1,000,000

Full time

11 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

n11 is seeking a Site Reliability Manager to lead the SRE/DevOps team in the Technology/Infrastructure Department. You will guide system options, drive SRE best practices, and mentor engineers while coordinating with software, security and product teams. This role emphasizes automation, incident response, and architectural roadmap execution.

The ideal candidate has 7+ years in infra/SRE/DevOps, strong leadership, and deep experience with Linux, Java app servers, GCP, and container platforms.

Qualifications

  • Bachelor's degree in a related field.
  • 7+ years in infrastructure/SRE/DevOps with at least 2 years in people management.
  • Hands-on with Linux systems administration and Java app servers (Tomcat/WebLogic).
  • Strong cloud experience, especially Google Cloud Platform.
  • Container management with Kubernetes and OpenShift.
  • Experience with CI/CD systems (Jenkins, GitLab, Bitbucket).
  • Automation/Configuration management (Ansible, Terraform).
  • Monitoring/alerting tools (Zabbix, Prometheus, Grafana).
  • Designing and troubleshooting large-scale distributed systems.
  • Excellent leadership and coaching abilities.
  • Knowledge of Kafka/RabbitMQ and caching technologies.
  • Familiarity with load balancers (Netscaler/F5) and APM/log aggregation tools.

Responsibilities

  • Provide guidance on system options, risks, impact, costs, and benefits.
  • Define and drive SRE best practices including SLIs, SLOs, and automation.
  • Lead, mentor and grow the Site Reliability Engineering team.
  • Define the team's technical roadmap and priorities.
  • Collaborate with Software Engineering, Security and Product teams.
  • Install, configure solutions; develop interfaces, stubs, and simulators; maintain scripts.
  • Monitor metrics and use systems to ensure performance, scalability, and stability.
  • Lead infrastructure lifecycle management including provisioning and deployments.
  • Oversee production incident response and root cause analysis.
  • Recommend improvements for performance and optimal solutions.

Skills

Linux systems
Java app servers
DevOps
GCP
Kubernetes/OpenShift
CI/CD
Automation/Config Mgmt
Monitoring/Alerting
Distributed systems
Leadership

Education

Bachelor's degree in a related field

Tools

Jenkins
GitLab
Bitbucket
Ansible
Terraform
Zabbix
Prometheus
Grafana
Elasticsearch
Splunk
Kubernetes
OpenShift
Netscaler
F5
New Relic
AppDynamics
Datadog
Kafka
RabbitMQ

Job description

Get ready to take your place on n11, an open market platform has made valuable contributions to the e-commerce sector since its establishment by bringing more than 330 thousand registered business partners to customers.

We are looking for "Site Reliability Manager" to join our team in Technology/Infrastructure Department.

What you’ll do:
  • Provide guidance and expertise on system options, risk, impact, costs, benefits etc.
  • Define and drive SRE best practices including SLIs, SLOs, operational excellence and automation initiatives.
  • Lead, mentor and grow the Site Reliability Engineering team.
  • Define the team’s technical roadmap and operational priorities.
  • Collaborate closely with Software Engineering, Security and Product teams.
  • Install and configure solutions, implement reusable components, translate technical requirements, assist with all stages of test data, develop interface stubs and simulators and perform script maintenance and updates
  • Keep up with metrics and use monitoring systems to provide the best performance, scalability and stability
  • Lead infrastructure lifecycle management, including provisioning, software deployments, monitoring and alerting.
  • Lead production incident response and root cause analysis.
  • Give recommendations for enhancing performance, identifying the most practical alternative solutions, and assisting with modifications
Who you are:
  • Bachelor's degree in a related field
  • 7+ years of experience in infrastructure, site reliability engineering, DevOps including at least 2 years in a people management role.
  • Extensive hands‑on experience with Linux systems administration, Java application servers (Tomcat, WebLogic), and deep knowledge of DevOps principles.
  • Strong experience with cloud platforms, especially Google Cloud Platform (GCP)
  • Strong experience with container management platforms such as Kubernetes and OpenShift
  • Experience with CI/CD systems such as Jenkins, GitLab or Bitbucket
  • Experience with Automation/Configuration management tools such as Ansible, Terraform
  • Experience with monitoring and alerting services such as Zabbix & Prometheus/Grafana
  • Experience in designing, analyzing and troubleshooting large‑scale distributed systems
  • Excellent leadership skills
  • Ability to lead, coach and develop the team
  • Knowledge of distributed messaging systems such as Kafka, RabbitMQ
  • Knowledge of various caching technologies such as Varnish, Couchbase, Redis
  • Knowledge of load balancers such as Netscaler, F5
  • Knowledge of Application Performance Management (APM) tools such as New Relic, AppDynamics or Datadog
  • Knowledge of log aggregation solutions such as Elasticsearch, Splunk or equivalents
  • Knowledge of network protocols including IP, TCP, HTTP, DNS, SSL

As n11.com, we care about your Personal Data Security. Please find the Personal Data Protection Information Notice from the link below.

https://n11scdn.akamaized.net/custom/upload/51/79/2889579912657586679.pdf

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

SRE Leadership Manager - Scale & Reliability
SRE Leadership Manager - Scale & Reliability

n11 • Fatih

On-site
TRY 800,000 - 1,000,000
Site Reliability Engineer
Site Reliability Engineer

MetLife México • Fatih

On-site
TRY 320,000 - 540,000
Private health insurance
Pension plan
Work from home allowance
+1
Senior Site Reliability Engineer (Performance and Scalability)
Senior Site Reliability Engineer (Performance and Scalability)

JobCubby • Turkey

On-site
TRY 600,000 - 1,200,000
Immediate impact
Top compensation
Regional talent
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • Turkey

On-site
TRY 600,000 - 900,000
Private health insurance
Continuous upskilling & development
English courses
+1
Site Reliability Engineer
Site Reliability Engineer

OBSS • Fatih

Hybrid
TRY 1,818,000 - 2,728,000
Flexible working arrangements
Training programs
Certifications
+1
Senior SRE: GenAI-Driven Reliability & Cloud Ops
Senior SRE: GenAI-Driven Reliability & Cloud Ops

EPAM Systems, Inc. • Turkey

On-site
TRY 600,000 - 900,000
Private health insurance
Continuous upskilling & development
English courses
+1
Senior DevOps Engineer
Senior DevOps Engineer

WAGNIFY Bilgi Teknolojileri • Turkey

On-site
TRY 350,000 - 600,000
Mid - Senior Network Planning Engineer
Mid - Senior Network Planning Engineer

n11 • Sarıyer

On-site
TRY 300,000 - 540,000
Sitecore Team Leader
Sitecore Team Leader

PeopleCert • Çankaya

On-site
TRY 80,000 - 100,000
Competitive remuneration
Learning opportunities
Diversity and inclusion initiatives
+1
Software Engineering Manager
Software Engineering Manager

NGSS • Fatih

On-site
TRY 650,000 - 950,000