Senior Site Reliability Engineer (Python/Kubernetes)

Luxoft Poland

Poland

On-site

PLN 180,000 - 240,000

Full time

31 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Job summary

Luxoft Poland is seeking an experienced automation and AI-enabled platform engineer to resolve L3 operational issues and build AI agents for L1/L2 automation. The role focuses on monitoring, incident response and developing scalable automation across a cloud observability platform.

You will work across AI/ML-driven automation, container orchestration with Kubernetes, and CI/CD pipelines while supporting 24x7 shifts.

Qualifications

  • Strong Python development skills for automation and AI agent development.
  • Kubernetes administration and troubleshooting.
  • Experience with observability and monitoring platforms (Grafana, Prometheus, Jaeger, Splunk).
  • Incident management experience with severity-based response processes (P1/P2/P3).
  • CI/CD pipeline experience (GitHub Actions, ArgoCD).
  • Experience with AI/ML concepts for building intelligent automation.
  • Strong root cause analysis skills.
  • Willingness to work in 24x7 shift rotation.

Responsibilities

  • Independently resolve L3 operational issues using engineering expertise, AI tooling, and purpose-built automations
  • Build and maintain AI agents for L1 & L2 automated operations
  • Develop automation solutions that reduce manual operational effort
  • Participate in development tasks for run-operation
  • Monitor platform availability and manage incident response/resolution by severity (P1/P2/P3)
  • Perform root cause analysis and determine if issues require code-level fixes
  • Participate in 24x7 shift rotation during weekdays and on-call on weekends

Skills

Python
Security Monitoring & Observability
Telemetry

Tools

Kubernetes

Job description

We're building a Cloud Observability platform for market leading large Germany based company with 440,000 customers worldwide.

Originally known for leadership in enterprise resource planning (ERP) software, the company has evolved to become a market leader in end-to-end enterprise application software, database, analytics, intelligent technologies, and experience management. A top cloud company with 200 million users worldwide, the company helps businesses of all sizes and in all industries to operate profitably, adapt continuously, and achieve their purpose.

Responsibilities
  • Independently resolve L3 operational issues using engineering expertise, AI tooling, and purpose-built automations
  • Build and maintain AI agents for L1 & L2 automated operations
  • Develop automation solutions that reduce manual operational effort
  • Participate in development tasks for run-operation
  • Monitor platform availability and manage incident response/resolution by severity (P1/P2/P3)
  • Perform root cause analysis and determine if issues require code-level fixes
  • Participate in 24x7 shift rotation during weekdays and on-call on weekends
Mandatory Skills
  • Python
  • Security Monitoring & Observability
  • Telemetry
Mandatory Skills Description
  • Strong Python development skills for automation and AI agent development
  • Kubernetes administration and troubleshooting
  • Experience with observability and monitoring platforms (Grafana, Prometheus, Jaeger, Splunk)
  • Incident management experience with severity-based response processes (P1/P2/P3)
  • CI/CD pipeline experience (GitHub Actions, ArgoCD)
  • Experience with AI/ML concepts for building intelligent automation
  • Strong root cause analysis skills
  • Willingness to work in 24x7 shift rotation
Nice-to-Have Skills Description
  • Experience with Generative AI / LLM for operational automation
  • Knowledge of OpenTelemetry and Kafka
  • Experience with OpenSearch / Elasticsearch
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Service Manager & Site Reliability Consultant
Service Manager & Site Reliability Consultant

GFT Technologies Poland • Łódź

On-site
PLN 180,000 - 320,000
Senior Site Reliability Engineer
Senior Site Reliability Engineer

EPAM Systems • Warszawa

Hybrid
PLN 240,000 - 360,000
Hybrid work model
Relocation opportunities
Work abroad opportunities
+2
Senior DevOps Engineer
Senior DevOps Engineer

iFindTech Ltd • Poland

On-site
PLN 260,000 - 420,000
AI-Driven Site Reliability Engineer for Cloud Observability
AI-Driven Site Reliability Engineer for Cloud Observability

Luxoft Poland • Poland

On-site
PLN 180,000 - 240,000
Principal Platform Engineer (Python)
Principal Platform Engineer (Python)

Intellias • Poland

On-site
PLN 260,000 - 380,000
Senior Cloud Engineer
Senior Cloud Engineer

Link Group • Warszawa

Hybrid
PLN 200,000 - 350,000
Senior Site Reliability Engineer (Banking)
Senior Site Reliability Engineer (Banking)

Capco • Warszawa

Hybrid
PLN 180,000 - 260,000
B2B contract
Hybrid collaboration
Senior Platform Engineer (Python)
Senior Platform Engineer (Python)

Intellias • Poland

On-site
PLN 180,000 - 280,000
Senior Site Reliability Engineer (SRE) – Kubernetes
Senior Site Reliability Engineer (SRE) – Kubernetes

Software Mind • Kraków

Remote
PLN 180,000 - 240,000
Private healthcare and insurance
Multisport card
Language classes
+3
Senior Site Reliability Engineer
Senior Site Reliability Engineer

Akamai Technologies • Województwo małopolskie

On-site