platform engineer for observability platforms

HireHi

United States

Remote

USD 140,000 - 230,000

Full time

40 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Flexible working hours
Private medical insurance compensation
Education budget
Co-working and gym reimbursement
Vacation and holidays

Job summary

HireHi ищет опытного инженера по инфраструктуре для поддержки observability-платформы и GitLab CI. Вы будете отвечать за устойчивость сервисов, обновления и автоматизацию через IaC, а также развивать runbooks и документацию. В команде ценится умение сотрудничать с разработчиками и управлять инцидентами.

Требуется опыт работы с Kubernetes, GitLab CI, Prometheus и Grafana; английский не ниже upper-intermediate; акцент на самообслуживание и эффективное взаимодействие с командами.

Qualifications

  • Старший уровень опыта в инфраструктуре, платформах или SRE.
  • Опыт работы с Kubernetes в продакшене через GitOps и обновления кластеров.
  • IaC с использованием Ansible и Terraform/OpenTofu; изменения через merge-запросы.
  • Администрирование GitLab и GitLab CI, self-hosted или SaaS.
  • Знание Prometheus и Grafana: настройка алертов и дашбордов.
  • Умение писать технические пояснения, runbooks и уведомления.
  • Сильные коммуникативные навыки и работа с заинтересованными сторонами.
  • Продвинутые навыки работы с AI-ассистентами (Claude, Codex); проверка результатов.
  • Уровень английского не ниже upper-intermediate.
  • Не подходит кандидатам, ориентированным только на тикет-queue; важны bare metal/ВМ.

Responsibilities

  • Эксплуатация и поддержка observability-платформы, мониторинг расходов и емкости, настройка алертинга.
  • Управление GitLab и CI-флоу, обновлениями, емкостью, доступом, бэкапами и учениями восстановления.
  • Поддержка других сервисов через мониторинг и документацию runbooks.
  • Развертывание новых сервисов как код, выбор дизайна и обеспечение управляемости.
  • Обработка запросов разработчиков: доступ, пайплайны, экспортёры и дашборды; перевод повторяющихся задач в self-service.
  • Ответ на инциденты, диагностика, восстановление и пост-мортем-расследования; внедрение профилактики.
  • Деление изменений как код через MR, планирование и проверку изменений.
  • Написание инструкций, уведомлений и статусов для инженеров за пределами команды.
  • Работа с ИИ-агентами: делегирование сбора, черновиков и обучение команды.

Skills

Kubernetes
GitLab
CI/CD
GitOps
Ansible
Terraform/OpenTofu
Prometheus
Grafana
Runbooks
English

Tools

GitLab CI
OpenTofu
Terraform
Ansible
Prometheus
Grafana

Job description

Описание

CloudLinux builds Linux infrastructure and security products for hosting companies and data centers. Its Infrastructure Department runs an observability platform, GitLab, CI runners, engineering services, and provisioning and configuration automation.

Задачи
  • Run the observability platform, keep it healthy, onboard teams, monitor cost and capacity, and maintain alerting
  • Run GitLab and the CI runner fleet, including upgrades, capacity, access, backups, and restore drills
  • Keep other services healthy with production monitoring and runbooks
  • Deploy requested services from scratch by researching options, choosing designs, and ensuring they are managed as code, monitored, backed up, and documented
  • Handle developers’ requests related to access, onboarding, pipeline issues, exporters, and dashboards, and turn recurring requests into self-service
  • Respond to incidents, diagnose and mitigate impact, restore services safely, complete root-cause analyses and post-mortems, and implement prevention or detection improvements
  • Ship changes as code through reviewed merge requests, planning and checking each change
  • Write runbooks, onboarding guides, maintenance notices, and status updates for engineers outside the team
  • Work with AI agents by delegating collection and drafting, reviewing their output, and recording lessons for the team
Требования
  • Senior-level experience in infrastructure, platform, or site reliability engineering, including responsibility for keeping at least one production service running Linux systems administration and debugging on bare metal and virtual machines
  • Production Kubernetes experience delivered through GitOps, including personally performing cluster upgrades
  • Infrastructure as code experience using Ansible and Terraform or OpenTofu, with changes reviewed in merge requests
  • Production GitLab administration and GitLab CI experience, self-hosted or SaaS; equivalent depth with another CI system is acceptable
  • Working knowledge of Prometheus and Grafana, including running them for a team, writing alert rules and dashboards, and reading PromQL
  • Ability to write technical explanations for engineers outside the team, including runbooks, notices, and responses to requests
  • Strong communication and interpersonal skills; able to understand product-team needs, agree on scope, priority, and timing, push back politely, and keep stakeholders informed
  • Advanced use of AI engineering assistants such as Claude and Codex, including providing context, breaking down tasks, designing agent loops, and delegating scoped end-to-end execution with clear stop conditions; able to explain, debug, and test automation and verify generated commands, scripts, and conclusions before production use
  • Upper-intermediate or higher English
  • Not suited to candidates seeking ticket-queue operations, a pure cloud or Kubernetes role, or responsibility as a DBA, network engineer, or security engineer. Recurring requests are expected to become self-service; bare metal and virtual machines are a significant part of the infrastructure, and other teams run their own systems
  • Будет плюсом: SLO and burn-rate alert design with data-sized thresholds, Kata Containers, Firecracker or gVisor, S3-compatible object storage such as Ceph RGW, AWS cost work, self-hosted Sentry or Kafka-, ClickHouse- and Redis-backed applications kept running under load, Python or Go for exporters and small internal services
Условия
  • Flexible working hours 24
  • Paid vacation days per year, 10 national holidays, and unlimited sick leave
  • Private medical insurance compensation
  • Co-working and gym/sports reimbursement
  • Education budget
  • Opportunity to receive a reward for an idea the company can patent
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

devops engineer in Linux infrastructure
devops engineer in Linux infrastructure

HireHi • United States

Remote
USD 140,000 - 210,000
Гибкий график
Оплачиваемый отпуск 24 дня в год
Медицинское страхование
team lead platform engineering
team lead platform engineering

HireHi • United States

Remote
USD 124,000 - 180,000
Performance bonus
Private health insurance
Corporate pension
+5
platform operations engineer
platform operations engineer

HireHi • United States

Remote
USD 120,000 - 170,000
Vacation days
English lessons
platform engineer in cloud infrastructure
platform engineer in cloud infrastructure

HireHi • United States

Remote
USD 90,000 - 130,000
Health insurance
Vacation policy
Sick leave
+3
Platform Engineer (remote work)
Platform Engineer (remote work)

CloudLinux Inc. • United States

Remote
USD 120,000 - 160,000
Fully remote work with flexible hours
Education budget
Private medical insurance
+1
devops engineer in cybersecurity
devops engineer in cybersecurity

HireHi • United States

Remote
USD 68,000 - 113,000
Private health insurance
Training и менторство
Годовой отпуск и больничные
+3
devops инженер tools team
devops инженер tools team

HireHi • United States

Remote
USD 90,000 - 130,000
Обучение и развитие
ДМС со стоматологией
Корпоративный спорт
+5
Senior DevOps Engineer
Senior DevOps Engineer

DevOps Jobs - работа и аналитика • United States

On-site
USD 35,000 - 42,000
DevOps - Centicore
DevOps - Centicore

DevOps Jobs - работа и аналитика • United States

On-site
USD 43,000 - 47,000
devops engineer for hybrid infrastructure
devops engineer for hybrid infrastructure

HireHi • United States

Remote
USD 120,000 - 160,000
AWS/GCP stack
Microsoft 365 stack