Senior Systems Site Reliability Engineer, B2B

Jobtailor

Poland

On-site

PLN 180,000 - 320,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Jobtailor is seeking an experienced Site Reliability Engineer to partner with engineering, define SLOs, and investigate complex issues end-to-end using AI. You will produce documentation, runbooks, and postmortems while leading toil reduction through automation and AI-enabled tooling.

You will work on safe AI integration, participate in ceremonies, and mentor teams on SRE practices, including fault injection and disaster recovery exercises.

Qualifications

  • Minimum 5 years in software/SRE/production ops roles.
  • Strong troubleshooting across the stack with first principles.
  • Experience in Agile development processes.
  • Hands-on AWS capabilities (EC2, S3, EKS, RDS/Aurora, CloudFront).
  • Observability tooling experience (Grafana, Prometheus, LogicMonitor).
  • Clear technical docs for technical and non-technical audiences.
  • Infrastructure as code experience.
  • Automation development in a general-purpose language (Python/Go/Java).
  • Judgement in applying AI safely in production access and data handling.
  • Hands-on with agentic tools (Claude; Copilot; Cursor).
  • Improve team AI adoption: reusable skills, repository context, prompt patterns.
  • SQL optimization and DB tuning (preferred).
  • CI/CD tooling experience (GitHub Actions, Jenkins).
  • Exposure to chaos engineering and disaster recovery (preferred).
  • FinOps familiarity (preferred).
  • Bachelor's degree or equivalent considered.

Responsibilities

  • Partner with engineering to define SLOs and error budgets and improve reliability investments.
  • Investigate production issues end-to-end across stack using AI-assisted analysis.
  • Create runbooks, architecture notes, postmortems, and reusable proofs of concept.
  • Identify toil sources and lead automation and tooling improvements.
  • Set conditions for AI agents to work safely with tests and guardrails.
  • Participate in team ceremonies and drive collaboration opportunities.
  • Drive cross-team collaboration to influence roadmaps and mentor on SRE/AI usage.
  • Advise senior leadership during customer escalations, translating tech impact to business.
  • Contribute to scaling the SRE practice with better standards and tooling.

Skills

5+ years experience in software/SRE/PU
Production troubleshooting
Agile development experience
AI-guided reliability
Cross-team collaboration
AI agent tooling usage
Team tooling and automation
Reliable production access governance
Reliability practices with AI
Toil reduction through automation
SQL query optimization
CI/CD tooling
Chaos engineering / DR exercises
FinOps awareness

Education

Bachelor's degree or equivalent

Tools

AWS
Grafana
Prometheus
LogicMonitor
GitHub Actions
Jenkins
Terraform / IaC
Claude Code / Cursor / Copilot

Job description

Responsibilities
  • Partner with engineering teams to define service-level objectives, error budgets, and supporting indicators for their services, and help them use those measures to inform prioritization and reliability investment.
  • Investigate complex production issues end-to-end across application, data, infrastructure, and network layers, using AI to correlate logs, metrics, and code and to pressure-test hypotheses before acting.
  • Produce clear technical documentation, runbooks, architecture notes, postmortems and proofs of concept for both technical and non-technical audiences, in a form that engineers and AI tools can re-use.
  • Identify systemic sources of toil and lead the work to eliminate them through automation, AI agents, tooling, and process change.
  • Set the conditions for AI agents to do reliable work in our environment, including repository context, well-specified tasks, integrations such as MCP servers that give AI safe access to the systems it needs, and the tests and guardrails needed for AI-authored change to be trusted.
  • Participate in team ceremonies to identify and refine work, communicate findings, and drive opportunities to collaborate.
  • Drive cross-team and cross-department collaboration on reliability initiatives, including reviewing designs, influencing roadmaps, and mentoring engineers on SRE practices, including effective AI use in their reliability work.
  • Advise senior leadership and stakeholders during critical customer escalations, translating between technical reality and business impact.
  • Contribute to scaling the SRE practice itself: improving our standards, our tooling, and how we partner with product engineering teams.
Requirements
  • Minimum of 5 years experience in software engineering, SRE or production operations roles. (Required)
  • Strong production troubleshooting skills across the stack. Ability to diagnose issues from first principles using the tools available (profilers, heap and thread dumps, query plans, traces, logs, metrics). (Required)
  • Experience working within a form of the Agile development framework process. (Required)
  • Hands‑on experience operating production services on AWS (e.g. EC2, S3, EKS, RDS/Aurora, CloudFront). (Required)
  • Experience utilizing observability tools (i.e. Grafana, Prometheus, LogicMonitor). (Required)
  • Experience creating clear and concise technical documentation that is targeted at both technical and non-technical audiences. (Required)
  • Experience writing infrastructure as a code. (Required)
  • Experience writing automation in a general-purpose language (e.g. Python, Go, Java, or similar) to a production standard. (Required)
  • Strong judgement about how to apply AI effectively across the full range of SRE work, including high-stakes areas such as production access and sensitive data, knowing how to scope and verify work to make it safe. (Required)
  • Hands‑on experience using agentic development tools (e.g. Claude Code, Cursor, Copilot) to deliver engineering and operational work, scoping and delegating bounded tasks, verifying the output, and shipping with confidence. (Required)
  • Experience improving how a team works with AI, for example authoring reusable skills, repository context files, or prompt patterns that others adopt. (Required)
  • Experience optimizing SQL queries and database engine tuning. (Preferred)
  • Experience with CI/CD Tooling (e.g. Github Actions, Jenkins). (Preferred)
  • Exposure to chaos engineering, fault injection and disaster recovery exercises. (Preferred)
  • Familiar with FinOps practices. (Preferred)
  • Bachelor's degree or a combination of relevant experience and education may be considered.
Core Competencies

Demonstrates expertise in Site Reliability Engineering (SRE) practices, including production troubleshooting, automation, and effective use of AI in operational contexts. Proficient in creating technical documentation and collaborating across teams to enhance reliability and operational efficiency.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Site Reliability Engineer
Senior Site Reliability Engineer

Grid Dynamics • Kraków

On-site
PLN 254,000 - 340,000
Medical insurance
Sports benefits
Professional development opportunities
+2
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Grid Dynamics • Województwo pomorskie

On-site
PLN 80,000 - 120,000
Medical insurance
Sports benefits
Professional development opportunities
+2
Site Reliability Engineer
Site Reliability Engineer

Balyasny Asset Management L.P. • Warszawa

On-site
PLN 180,000 - 300,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Kraków

On-site
PLN 180,000 - 270,000
Site Reliaibility engineer
Site Reliaibility engineer

Alcor • Kraków

On-site
PLN 250,000 - 380,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Grid Dynamics • Wrocław

On-site
Competitive salary
Flexible schedule
Medical insurance
+4
Site Reliability Engineer
Site Reliability Engineer

Caspian One • Warszawa

On-site
PLN 180,000 - 280,000
Site Reliability Engineer
Site Reliability Engineer

Fáilte Ireland • Poland

On-site
PLN 180,000 - 280,000
Equity program
Senior SRE: AI-Driven Reliability & Automation
Senior SRE: AI-Driven Reliability & Automation

Jobtailor • Poland

On-site
PLN 180,000 - 320,000
Site Reliability Engineer
Site Reliability Engineer

Qualient Technology Solutions UK Limited • Województwo małopolskie

Hybrid
PLN 180,000 - 260,000
Hybrid work model
Learning culture