infrastructure engineer in iGaming

HireHi

United States

Remote

USD 120,000 - 180,000

Full time

40 hours ago
Be an early applicant
Application generator

Turn this role into an interview — a resume and cover letter built around what this employer wants.

Get past ATS filters

Benefits offered by this job

Health insurance
Paid leave
Vacation
Learning opportunities
Language learning
Modern hardware
International team
Relocation opportunities

Job summary

HireHi ищет специалиста по наблюдаемости и инфраструктуре для проектирования, внедрения и поддержки масштабируемых платформ мониторинга и логирования. Вы будете работать с Prometheus, Alertmanager, Grafana и стеком OpenSearch, обеспечивая надёжное хранение и быстрый доступ к данным, а также развивать self-service observability для инженеров.

Требуется 3+ лет опыта администрирования Linux и работы с контейнеризацией (Docker/Kubernetes), знание Ansible/Terraform и навыки автоматизации.

Qualifications

  • 3+ года опыта администрирования Linux и мониторинга/об observability систем
  • Опыт работы с системами на базе Debian
  • Практический опыт Prometheus, Alertmanager и Grafana
  • Опыт работы Fluent Bit, Kafka, OpenSearch и OpenSearch Dashboards
  • Понимание проектирования пайплайнов метрик и логов, надёжности и масштабируемости
  • Умение проектировать алерты и снижать шум уведомлений
  • Уверенное владение shell, скриптингом и автоматизацией (Ansible, Terraform)
  • Понимание сетевых концепций (TCP/IP, DNS, VPN, firewalls)
  • Опыт эксплуатации высокодоступных сервисов и планирования ресурсов
  • Опыт работы с инфраструктурой как кодом и Git-воркфлоу
  • Готовность обеспечить самообслуживание observability в командах

Responsibilities

  • Разработка и внедрение масштабируемых платформ мониторинга, логирования и алертинга
  • Операция Prometheus, Alertmanager, Grafana, Fluent Bit, Kafka, OpenSearch и Dashboards
  • Построение и поддержка потоков метрик и логов для критически важных сервисов
  • Создание дашбордов для видимости здоровья сервиса, производительности и рисков
  • Проектирование действующих правил оповещений и маршрутизации, снижение шума
  • Мониторинг метрик и логов, устранение проблем, повышение устойчивости под нагрузкой
  • Управление размерностью метрик, хранением, retention и производительностью запросов
  • Обеспечение HA и восстановления для observability-сервисов
  • Автоматизация конфигураций через Ansible, Terraform, Python и Bash
  • Разработка паттернов observability и документации
  • Поддержка инцидентов, анализ причин и улучшение дашбордов/рунабуков
  • Оценка новых технологий и взаимодействие с командами Kubernetes/Linux/сетями
  • Поддержка документации по архитектуре observability

Skills

Linux admin
Networking basics
Shell scripting
Git workflows
Security basics

Tools

Prometheus
Alertmanager
Grafana
Fluent Bit
Kafka
OpenSearch
OpenSearch Dashboards
Docker
Kubernetes
Ansible
Terraform
Python
Bash

Job description

Описание:

Betby powers the iGaming industry with a premium sportsbook featuring risk management and omni-channel support, reaching millions of players across multiple markets.

Задачи:

Design, deploy, configure, and maintain scalable monitoring, logging, and alerting platforms for production and beta/test/dev environments Operate Prometheus, Alertmanager, Grafana, Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards, including upgrades, reliability, availability, capacity, and retention planning Build and maintain reliable metrics and log collection pipelines for infrastructure and business-critical services Create dashboards that provide visibility into service health, performance, capacity, and operational risks Design, tune, and maintain actionable alert rules and notification routing; reduce alert noise and improve incident response Monitor infrastructure and application metrics and logs, troubleshoot issues, and improve stability and performance under heavy loads Manage metric cardinality, log volume, retention, storage consumption, and query performance to keep observability platforms scalable and cost-effective Establish high-availability and recovery approaches for observability services and validate operational readiness Automate configuration management and standardize observability configuration through Ansible, Terraform, Python, and bash Develop self-service observability patterns, reusable dashboards, alert templates, and documentation for engineering teams Support production incidents, investigate root causes with telemetry, and improve dashboards, alerts, and runbooks after incidents Evaluate new technologies and their implementation in existing infrastructure Work with Kubernetes, Linux systems, networking, databases, and message brokers to ensure meaningful observability coverage Maintain and write documentation of observability architecture, configurations, standards, and operational procedures

Требования:

At least 3 years of experience administering Linux systems and operating monitoring, logging, or observability systems Experience with Debian-based systems Experience with Docker and Kubernetes Hands-on experience with Prometheus, Alertmanager, and Grafana, including metric collection, alert rules, routing, and dashboards Experience with Fluent Bit, Kafka, Fluentd, OpenSearch, and OpenSearch Dashboards for log collection, transport, processing, storage, search, and visualization Understanding of metrics and log pipeline design, including reliability, scalability, data retention, capacity planning, and cardinality management Experience designing actionable alerts, reducing alert noise, and troubleshooting infrastructure and application issues using metrics and logs Proficiency in shell command line usage, scripting, and automation tools such as Ansible and Terraform Python and bash scripting skills Understanding of networking concepts, including TCP/IP, DNS, VPN, firewalls, and configuring and troubleshooting network settings Experience operating highly available services and planning capacity for production workloads Experience with configuration-as-code, Git-based workflows, and enabling self-service observability for engineering teams Будет плюсом: VictoriaMetrics experience

Условия:
  • Comprehensive health insurance
  • Up to 10 days of paid sick leave without a medical certificate
  • 20 Days of paid vacation plus additional leave for important life events
  • Learning and growth opportunities with support for professional development
  • Language learning support
  • Modern work hardware provided
  • International team environment across multiple countries
  • Corporate events and team activities
  • Welfare support program for critical situations
  • Gifts and support for major life milestones
  • Relocation opportunities
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

support engineer for B2B iGaming platforms
support engineer for B2B iGaming platforms

HireHi • United States

Remote
USD 55,000 - 65,000
Annual Benefits Budget
Maternity / Paternity leave
20+ Vacation days
+1
Software Engineer (Ruby) – Senior in iGaming
Software Engineer (Ruby) – Senior in iGaming

HireHi • United States

Remote
USD 120,000 - 180,000
Private health insurance
Sports benefits
Comprehensive Mental Health Program
+6
технический директор в iGaming
технический директор в iGaming

HireHi • United States

Remote
USD 150,000 - 210,000
Remote work
Flexible schedule
28 days paid vacation
+2
data engineer in fintech
data engineer in fintech

Enfint • United States

On-site
USD 120,000 - 180,000
Provident Fund
Annual learning budget
€150 Monthly Wolt allowance
+4
devops engineer in Linux infrastructure
devops engineer in Linux infrastructure

HireHi • United States

Remote
USD 140,000 - 210,000
Гибкий график
Оплачиваемый отпуск 24 дня в год
Медицинское страхование
Человек на стыке сисадмина и DevOps
Человек на стыке сисадмина и DevOps

FirstMessage • United States

On-site
USD 33,000 - 43,000
Компенсация спорта и изучения языков
Бизнес-трипы за счёт компании
Тимбилдинги и дружелюбная атмосфера
+1
go developer for the insurance and assistance industry
go developer for the insurance and assistance industry

HireHi • United States

Remote
USD 120,000 - 170,000
Private health insurance
Training portal access
Performance-based bonus
media buyer in Gambling/iGaming
media buyer in Gambling/iGaming

HireHi • United States

Remote
USD 90,000 - 130,000
Modern equipment
Health insurance
Paid time off
+9
Solution Architect payment platform
Solution Architect payment platform

HireHi • United States

Remote
USD 120,000 - 180,000
Private health insurance
Sports compensation
Certification compensation (AWS, PMP,等
head of analytics в iGaming
head of analytics в iGaming

HireHi • United States

Remote
USD 120,000 - 180,000
Remote work
Flexible schedule
Growth opportunities
+1