Infrastructure Analyst – Page Diversity

Jobtailor

Deutschland

Remote

EUR 65.000 - 95.000

Vollzeit

Vor 4 Tagen
Sei unter den ersten Bewerbenden
Bewerbungsgenerator

Mach aus dieser Rolle ein Vorstellungsgespräch — ein Lebenslauf und ein Anschreiben, die darauf ausgerichtet sind, was dieser Arbeitgeber sucht.

Schaffe es an den ATS-Filtern vorbei

Zusammenfassung

Jobtailor in Germany seeks a DevOps/SRE specialist to ensure cloud environments are reliable and cost-efficient. You will automate provisioning with Terraform and Helm, build dashboards in Grafana, and improve monitoring with Prometheus.

You will respond to incidents and drive RCAs, shaping SLOs and SLIs, while collaborating with teams to optimize performance. Join a dynamic team focused on scalable infrastructure, incident management, and continuous delivery across cloud platforms.

Qualifikationen

  • Experience running Kubernetes in production across multiple clusters.

Aufgaben

  • Ensure availability, reliability, and efficiency of cloud environments.
  • Improve monitoring and observability for production environments.
  • Identify and mitigate operational risks proactively.
  • Define metrics reflecting real user experience.
  • Improve alert quality by reducing false positives and noise.
  • Create and maintain dashboards for technical analysis and decision-making.
  • Automate manual and repetitive tasks.
  • Advance infrastructure as code using Terraform and Helm.
  • Develop Python and Shell Script automations.
  • Triage, mitigate, and resolve incidents in production.
  • Contribute to root cause analyses and preventive actions.
  • Track infrastructure utilization and cost indicators.
  • Automate provisioning and deprovisioning of environments and access.

Kenntnisse

AWS Cloud
Kubernetes
Python
Shell Script
Terraform
Helm
Observability
Grafana
Prometheus
Linux Troubleshooting
Networking
REST APIs
CI/CD
Ansible

Tools

Terraform
Helm
Kubernetes
Prometheus
Grafana
Elasticsearch
OpenSearch
Loki
Ansible
CI/CD pipelines

Jobbeschreibung

  • Ensure the availability, reliability, and efficiency of cloud environments
  • Enhance monitoring and observability processes for production environments
  • Identify and proactively mitigate operational risks
  • Define metrics and indicators that reflect the real user experience
  • Improve alert quality by reducing false positives and operational noise
  • Create and maintain dashboards for technical analysis and decision-making
  • Automate manual and repetitive tasks
  • Advance infrastructure as code using Terraform and Helm
  • Develop scripts and automations in Python and Shell Script
  • Triage, mitigate, and resolve incidents in production environments
  • Participate in root cause analyses and implement preventive actions
  • Track infrastructure utilization and cost indicators
  • Automate the provisioning and deprovisioning of environments and access
Requirements
  • Experience with Kubernetes in production environments
  • Knowledge of cloud computing, preferably AWS
  • Experience with Infrastructure as Code (IaC), using Terraform and Helm
  • Knowledge of observability and monitoring with Prometheus and Grafana
  • Hands-on experience troubleshooting Linux environments
  • Experience developing automations using Python and/or Shell Script
  • Experience responding to incidents in production environments
  • Knowledge of networking, DNS, TLS, VPNs, and diagnostic and debugging tools
  • Preferred: experience with REST APIs and API gateways
  • Preferred: knowledge of Elasticsearch, OpenSearch, or Loki
  • Preferred: experience building and maintaining CI/CD pipelines
  • Preferred: knowledge of SLIs, SLOs, and Error Budgets
  • Preferred: experience with Ansible or similar tools
  • Preferred: English proficiency for reading and consulting technical documentation
Core Competencies

Demonstrates expertise in cloud environment management, focusing on availability, reliability, and efficiency. Proficient in automating processes and enhancing observability using tools like Terraform, Helm, Prometheus, and Grafana.

Highest-signal resume keywords
  • Cloud Computing (AWS)
  • Infrastructure as Code (IaC)
  • Kubernetes in Production
  • Python and Shell Script Automation
  • Observability and Monitoring (Prometheus, Grafana)
Hard Skills
  • Terraform
  • Helm
  • Python
  • Shell Script
  • Kubernetes
  • Linux Troubleshooting
  • Networking
  • REST APIs
  • CI/CD Pipelines
  • Ansible
Industry Keywords
  • Infrastructure Utilization
  • Operational Risks
  • Incident Response
  • SLIs
  • SLOs
  • Error Budgets
  • Monitoring Processes
  • Automated Provisioning
  • Technical Dashboards
  • Root Cause Analysis
Tools & Technologies
  • Prometheus
  • Grafana
  • Elasticsearch
  • OpenSearch
  • Loki
Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Cloud Engineer – Senior
Cloud Engineer – Senior

Jobtailor • Deutschland

Remote
EUR 90.000 - 120.000
Senior Infrastructure Engineer
Senior Infrastructure Engineer

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000
DevOps Specialist
DevOps Specialist

Jobtailor • Deutschland

Remote
EUR 70.000 - 120.000
Senior Cloud Engineer – Infrastructure Services, FedRAMP
Senior Cloud Engineer – Infrastructure Services, FedRAMP

Jobtailor • Deutschland

Hybrid
EUR 110.000 - 160.000
DevOps Engineer, Blockchain Infra
DevOps Engineer, Blockchain Infra

Jobtailor • Deutschland

Remote
EUR 70.000 - 110.000
Mid-Level SRE Analyst
Mid-Level SRE Analyst

Jobtailor • Deutschland

Remote
EUR 90.000 - 130.000
Software Engineer, DevOps
Software Engineer, DevOps

Jobtailor • Deutschland

Remote
EUR 70.000 - 95.000
Senior Cloud Engineer – Linux Specialist
Senior Cloud Engineer – Linux Specialist

Jobtailor • Berlin

Vor Ort
EUR 70.000 - 100.000
Senior Cloud Engineer – IaC
Senior Cloud Engineer – IaC

Jobtailor • Deutschland

Remote
EUR 85.000 - 120.000
DevOps Engineer – On-Premises Cloud
DevOps Engineer – On-Premises Cloud

Jobtailor • Bonn

Vor Ort
EUR 90.000 - 130.000