Production Support & Operations Engineer Splunk, Sysdig, Prometheus, Grafana, Kubernetes, Docker Warszawa, Gdańsk

Diverse CG Sp. z o.o. Sp.k.

Województwo pomorskie

Hybrid

PLN 180,000 - 240,000

Full time

14 days+
Application generator

Don’t send a generic resume — generate a resume and cover letter tailored to this exact role.

Get past ATS filters

Benefits offered by this job

Private medical care
Co-financing for the sports card
Constant support of dedicated consult

Job summary

Diverse CG Sp. z o.o. Sp.k. is looking for a Production Support & Operations Engineer to lead RCA for production incidents, monitor services, and improve CI/CD pipelines. You will collaborate with Dev and platform teams to resolve issues, automate deployment, and strengthen observability with Grafana and Prometheus.

You will work in an ITIL-enabled, Agile environment, handling incident/problem/change management and driving automation across Docker, Kubernetes, Jenkins, and GitHub Actions.

Qualifications

  • 5+ years in IT operations, application support or similar production-facing role.
  • End-to-end incident management experience from alert to RCA and prevention.
  • 2+ years with ITIL processes: incident, problem, change management.
  • Experience in Agile delivery with development teams.
  • Strong Jenkins knowledge for building and debugging pipelines.
  • Hands-on CI/CD, automation and continuous improvement mindset.
  • Proficient in Splunk and Sysdig for log analysis and alerting.
  • Knowledge of Prometheus/Grafana dashboards and alert tuning.
  • Experience with Kubernetes including pod health checks and restarts.
  • Ansible for controlled configuration changes; Docker/Docker Compose expert.
  • Scripting: Bash and Python for automation and reconciliation.

Responsibilities

  • Own RCA activities for production incidents, including diagnosis, resolution, and preventive actions.
  • Monitor production services, identify anomalies, and proactively address issues.
  • Design, implement, and improve CI/CD pipelines using GitHub Actions.
  • Automate build, test, security scanning, and deployment processes.
  • Oversee day-to-day stability of Pre-Production and Production environments.
  • Maintain operational docs, runbooks, and resolution procedures.
  • Collaborate with development and platform teams to troubleshoot operational needs.
  • Identify recurring pain points and propose automation or tooling improvements.
  • Enhance observability with dashboards, alerts, and log queries.
  • Support service continuity initiatives and participate in DR exercises.

Skills

Jenkins
CI/CD
Kubernetes
Prometheus
Grafana
Splunk
Sysdig
Ansible
Docker
Docker Compose
Bash
Python
ITIL
Incident management
Agile
Monitoring
Log analysis

Tools

GitHub Actions
IBM Datastage
Pega
Airflow
Oracle
DB2
Kafka

Job description

Production Support & Operations Engineer
Position: Production Support & Operations Engineer
Production Support & Operations Engineer

Responsibilities:

  • Own RCA activities for production incidents, including diagnosis, resolution, and preventive actions
  • Monitor production services, identify anomalies, and proactively address issues before they become incidents
  • Design, implement, maintain, and improve CI/CD pipelines using GitHub Actions and related tooling
  • Automate build, test, security scanning, and deployment processes
  • Oversee day-to-day stability and alignment of Pre-Production and Production environments
  • Maintain operational documentation, runbooks, known issues, and resolution procedures
  • Collaborate with development and platform teams to troubleshoot issues and clarify operational requirements
  • Identify recurring operational pain points and propose automation or tooling improvements
  • Enhance observability through dashboards, alerts, and log queries
  • Support service continuity initiatives and participate in disaster recovery exercises

Requirements:

  • Minimum 5 years of experience in IT operations, application support (2nd/3rd line), or a similar production-facing role
  • Proven experience managing incidents end-to-end, from alerting and troubleshooting through RCA and prevention
  • Minimum 2 years of experience working within ITIL processes including incident, problem, and change management
  • Experience working in Agile delivery environments alongside development teams
  • Strong knowledge of Jenkins for building, maintaining, and troubleshooting deployment pipelines
  • Hands-on experience with CI/CD pipelines, automation, and continuous improvement initiatives
  • Excellent troubleshooting and problem-solving skills for complex production issues
  • Proficiency with Splunk and Sysdig for log analysis and alerting
  • Strong understanding of Prometheus and Grafana for monitoring, dashboards, and alert tuning
  • Practical experience operating services running on Kubernetes, including pod health checks, log analysis, and service restarts
  • Expertise in Ansible for controlled configuration changes in operational environments
  • Strong knowledge of Docker and Docker Compose
  • Basic scripting skills in Bash and Python for operational automation and data reconciliation

Nice to have:

  • Experience with IBM Datastage operations
  • Awareness of or willingness to learn Pega and Airflow
  • Experience with Oracle and DB2, including querying, execution plan interpretation, and data incident analysis
  • Understanding of ETL application behavior and REST API communication
  • Experience supporting distributed systems and Kafka-based message flows
  • Java or development background supporting understanding of solutions and integrations

Offer:

  • Private medical care
  • Co-financing for the sports card
  • Constant support of dedicated consultant
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Production Support & Operations Engineer Splunk, Sysdig, Prometheus, Grafana, Kubernetes, Docker Warszawa, Gdańsk
Production Support & Operations Engineer Splunk, Sysdig, Prometheus, Grafana, Kubernetes, Docker Warszawa, Gdańsk

DCG Poland • Województwo pomorskie

Hybrid
PLN 180,000 - 300,000
Private medical care
Co-financing for the sports card
Dedicated consultant
Production Support & Operations Engineer
Production Support & Operations Engineer

DCG Poland • Warszawa

On-site
PLN 180,000 - 240,000
Private medical care
Co-financing for sports card
Dedicated consultant support
Senior Application and Operation Support Engineer
Senior Application and Operation Support Engineer

OPTIVEUM sp. z o.o. • Łódź

Hybrid
PLN 91,000 - 152,000
IT Operations Engineer
IT Operations Engineer

Experis ManpowerGroup Sp. z o.o. • Warszawa

Hybrid
PLN 182,000 - 259,000
Multisport Card
Life insurance
Private healthcare
+1
Senior Application and Operation Support Engineer
Senior Application and Operation Support Engineer

Optiveum • Łódź

Hybrid
PLN 125,000 - 178,000
Production Support Engineer
Production Support Engineer

HeadHR • Kraków

Hybrid
PLN 120,000 - 160,000
Production Support Engineer (10/952)
Production Support Engineer (10/952)

INFOLET SP. Z O.O. • Kraków

On-site
PLN 120,000 - 180,000
Relocation package
Extended medical care
Multisport Benefit card
+1
Senior Application and Operation Support Engineer
Senior Application and Operation Support Engineer

OPTIVEUM SPÓŁKA Z OGRANICZONĄ ODPOWIEDZIALNOŚCIĄ • Województwo pomorskie

Hybrid
PLN 91,000 - 152,000
Remote work opportunities
Integration events
Senior DevOps Engineer Linux, Bamboo, Docker, Podman, Kubernetes, Python, Bash Gdynia
Senior DevOps Engineer Linux, Bamboo, Docker, Podman, Kubernetes, Python, Bash Gdynia

Diverse CG Sp. z o.o. Sp.k. • Gdynia

Hybrid
PLN 180,000 - 240,000
Private medical care
Co-financing for sports card
Dedicated consultant support
Senior Production Support Engineer (10/953)
Senior Production Support Engineer (10/953)

INFOLET SP. Z O.O. • Kraków

On-site
PLN 180,000 - 260,000
Relocation package
Extended medical care
Multisport card
+1