Senior Site Reliability Engineer

EPAM Systems

Turkey

On-site

TRY 4,859,000 - 6,803,000

Full time

17 hours ago
Be an early applicant
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Benefits offered by this job

Private health insurance
Professional development
English courses
Certification programs
Learning platform access

Job summary

EPAM Systems is seeking a highly skilled Senior Site Reliability Engineer to join our team. You will collaborate with software developers and operations to ensure high reliability, scalability, and efficiency of our systems, with a focus on meeting customer expectations.

You will lead the adoption of AI-enabled platform capabilities (GenAI, AI agents, AIOps), define SLOs, manage error budgets, and reduce toil through automation.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field.
  • 5+ years in Site Reliability Engineering or related roles.
  • Experience with cloud providers AWS/GCP/Azure.
  • Experience implementing SRE practices (SLO/SLI, error budgets, incident management).
  • Proficiency in Python or similar scripting language.
  • Strong background in monitoring and observability tools.
  • CI/CD, IaC, and configuration management experience.
  • Container orchestration with Kubernetes and Docker.
  • English proficiency at Upper-Intermediate level (B2) or higher.

Responsibilities

  • Collaborate with development, security, QA, and operations teams to implement SRE practices and ensure reliability.
  • Define and support required level of reliability, availability, and performance for services and applications.
  • Design and deliver Cloud-based solutions tailored to client needs.
  • Troubleshoot, mitigate, and support fixing of infrastructure and application issues in a timely manner.
  • Implement a monitoring system for the infrastructure and application reliability.
  • Guide adoption of AI technologies on the platform (GenAI, AI agents, AIOps) to improve operations.
  • Communicate technical concepts clearly to engineering teams and management stakeholders.

Skills

SRE practices
Python scripting
Monitoring tools
CI/CD tools
Infrastructure as code
Kubernetes
Docker
English (B2+)

Education

Bachelor's degree in CS/Engineering

Tools

AWS
GCP
Azure

Job description

We are seeking a highly skilled and motivated Senior Site Reliability Engineer (SRE) to join our team.

In this critical role, you will collaborate closely with software developers and operations teams to ensure high reliability, scalability, and efficiency of our systems, with a strong focus on meeting and exceeding customer expectations. Your expertise will be crucial in deploying, maintaining, and automating our infrastructure and application environments to ensure seamless user experiences.

Your proactive involvement will be key to enhancing system reliability, optimizing resource utilization, and ensuring continuous improvement in our operational practices.

You will have the opportunity to lead the adoption of AI-enabled platform capabilities, such as generative AI (GenAI), autonomous agents, and AIOps, driving innovation and operational excellence.

Your responsibilities will include defining and tracking Service Level Objectives (SLOs), managing error budgets, and reducing toil through automation. You will play a pivotal role in driving the success of technology initiatives, maximizing their impact across the organization, and ensuring that solutions consistently meet the high standards our customers expect.

Responsibilities
  • Collaborate with development, security, quality, and operation teams to implement SRE practices and ensure system reliability
  • Define and support required level of reliability, availability, and performance for services and applications
  • Design and deliver Cloud-based solutions tailored to client needs
  • Troubleshoot, mitigate, and support fixing of the infrastructure and application issues in a timely manner
  • Implement a monitoring system for the infrastructure and application reliability
  • Guide adoption of AI technologies on the platform (GenAI, AI agents, AIOps) to improve operations
  • Communicate technical concepts clearly to both engineering teams and management stakeholders
Requirements
  • Bachelor’s degree in Computer Science, Engineering, or a related field
  • 5+ years of hands‑on experience in Site Reliability Engineering or related roles
  • Proven experience in any cloud (AWS/GCP/Azure)
  • Experience with implementing SRE practices such as SLO/SLI, Error budgets, Postmortems, Reducing Toil, capacity planning, and Incident Management
  • Python or other scripting/programming language
  • Strong background in monitoring tools
  • Proficiency in CI/CD tools, infrastructure as code, and configuration management
  • Solid knowledge of container orchestration technologies (Kubernetes, Docker)
  • English language proficiency at an Upper-Intermediate level (B2) or higher
Nice to have
  • Certification in Kubernetes, AWS/GCP/Azure, or similar technologies
  • Proven experience in DevOps
  • Expertise in deployment and management of LLMs, including technologies like RAG
  • Knowledge of managing and optimizing AI/ML models in production environments, including basic deployment, monitoring, and maintenance
  • Native AI cloud services: AWS Bedrock, Google Vertex AI, Azure AI
  • Experience in designing, building, and operating AI agents and agentic frameworks
  • Coding Agents: Claude Code, OpenCode, Cursor, Trae, Antigravity
We offer
  • CONTINUOUS UPSKILLING, LEARNING & DEVELOPMENT
    • Diversity of tasks and projects
    • Assessment center for objective review of competency level
    • Personal development plan
    • Mentoring programs and leadership development
    • Certification and professional development support
    • Access to learning platforms including more than 2,500 internal courses
    • English courses taught by certified teachers
  • CORPORATE BENEFITS
    • Extra leave days
    • Referral bonuses
  • COMPENSATION PACKAGE
    • Competitive compensation paid in USD
    • Regular salary and performance reviews
  • MEDICAL & HEALTHCARE
    • Private health insurance
    • Well‑being events
  • WORKING ENVIRONMENT
    • Recreation areas and kitchens
    • Tea, coffee and snacks
    • Sports equipment and game consoles
    • IT Equipment
    • Microsoft’s Software Assurance Home Use Program (HUP)

Please note that our Talent Attraction Team reviews applications and CVs submitted in English.

EPAM is global leader in AI transformation engineering and integrated consulting, serving Forbes Global 2000 companies and ambitious startups. With over thirty years of expertise in custom software, product and platform engineering, we empower our clients to become AI‑Native enterprises, driving measurable value from innovation and digital investments.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems, Inc. • Turkey

On-site
TRY 600,000 - 900,000
Private health insurance
Continuous upskilling & development
English courses
+1
Senior AWS DevOps Engineer
Senior AWS DevOps Engineer

EPAM Systems • Turkey

Hybrid
TRY 300,000 - 540,000
Continuous upskilling
Private health insurance
English courses
+3
Senior Generative AI Operations (GenAI Ops) Engineer
Senior Generative AI Operations (GenAI Ops) Engineer

EPAM Systems • Turkey

On-site
TRY 2,917,000 - 4,375,000
Private health insurance
English courses
Continuous upskilling
Lead Azure Cloud Engineer
Lead Azure Cloud Engineer

EPAM Systems • Turkey

On-site
TRY 400,000 - 800,000
Upskilling program
Private health insurance
Mentoring programs
+1
Senior AI Engineer with AWS Bedrock
Senior AI Engineer with AWS Bedrock

EPAM Systems • Turkey

Hybrid
TRY 3,885,000 - 6,799,000
Continuous upskilling
Diversity of tasks
Mentoring programs
+5
Forward Deployed Engineer/Chief Role
Forward Deployed Engineer/Chief Role

EPAM Systems • Turkey

On-site
TRY 4,375,000 - 5,834,000
Private health insurance
English courses and learning stipend
Mentoring programs
Senior Data Engineer, AI
Senior Data Engineer, AI

EPAM Systems • Turkey

On-site
TRY 4,375,000 - 6,806,000
Continuous upskilling
Diversity of tasks and projects
Mentoring programs
+5
Generative AI Operations Engineer (GenAI Ops)
Generative AI Operations Engineer (GenAI Ops)

EPAM Systems • Turkey

On-site
TRY 300,000 - 460,000
Private health insurance
English courses
Professional development support
+2
Senior Python Developer
Senior Python Developer

EPAM Systems • Turkey

On-site
TRY 5,834,000 - 8,751,000
Private health insurance
English courses
Referral bonuses
+1
Head of Enterprise Support and Service Operations
Head of Enterprise Support and Service Operations

EPAM Systems • Turkey

On-site
TRY 8,751,000 - 11,667,000
Private health insurance
English courses
Access to learning platforms (2,500+ C
+1