Turn this role into an interview — a resume and cover letter built around what this employer wants.
HireHi ищет специалиста по наблюдаемости и инфраструктуре для проектирования, внедрения и поддержки масштабируемых платформ мониторинга и логирования. Вы будете работать с Prometheus, Alertmanager, Grafana и стеком OpenSearch, обеспечивая надёжное хранение и быстрый доступ к данным, а также развивать self-service observability для инженеров.
Требуется 3+ лет опыта администрирования Linux и работы с контейнеризацией (Docker/Kubernetes), знание Ansible/Terraform и навыки автоматизации.
Betby powers the iGaming industry with a premium sportsbook featuring risk management and omni-channel support, reaching millions of players across multiple markets.
Design, deploy, configure, and maintain scalable monitoring, logging, and alerting platforms for production and beta/test/dev environments Operate Prometheus, Alertmanager, Grafana, Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards, including upgrades, reliability, availability, capacity, and retention planning Build and maintain reliable metrics and log collection pipelines for infrastructure and business-critical services Create dashboards that provide visibility into service health, performance, capacity, and operational risks Design, tune, and maintain actionable alert rules and notification routing; reduce alert noise and improve incident response Monitor infrastructure and application metrics and logs, troubleshoot issues, and improve stability and performance under heavy loads Manage metric cardinality, log volume, retention, storage consumption, and query performance to keep observability platforms scalable and cost-effective Establish high-availability and recovery approaches for observability services and validate operational readiness Automate configuration management and standardize observability configuration through Ansible, Terraform, Python, and bash Develop self-service observability patterns, reusable dashboards, alert templates, and documentation for engineering teams Support production incidents, investigate root causes with telemetry, and improve dashboards, alerts, and runbooks after incidents Evaluate new technologies and their implementation in existing infrastructure Work with Kubernetes, Linux systems, networking, databases, and message brokers to ensure meaningful observability coverage Maintain and write documentation of observability architecture, configurations, standards, and operational procedures
At least 3 years of experience administering Linux systems and operating monitoring, logging, or observability systems Experience with Debian-based systems Experience with Docker and Kubernetes Hands-on experience with Prometheus, Alertmanager, and Grafana, including metric collection, alert rules, routing, and dashboards Experience with Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards for log collection, transport, processing, storage, search, and visualization Understanding of metrics and log pipeline design, including reliability, scalability, data retention, capacity planning, and cardinality management Experience designing actionable alerts, reducing alert noise, and troubleshooting infrastructure and application issues using metrics and logs Proficiency in shell command line usage, scripting, and automation tools such as Ansible and Terraform Python and bash scripting skills Understanding of networking concepts, including TCP/IP, DNS, VPN, firewalls, and configuring and troubleshooting network settings Experience operating highly available services and planning capacity for production workloads Experience with configuration-as-code, Git-based workflows, and enabling self-service observability for engineering teams Будет плюсом: VictoriaMetrics experience