platform operations engineer

HireHi

United States

Remote

USD 120,000 - 170,000

Full time

40 hours ago
Be an early applicant
Application generator

Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.

Get past ATS filters

Benefits offered by this job

Vacation days
English lessons

Job summary

HireHi ищет специалиста по инцидентам и поддержке платформы, работающего в США. Вы будете отвечать за реагирование на оповещения, анализ причин с использованием логов и мониторинга, а также восстановление сервиса по утвержденным runbooks.

Роль требует владения Linux, навыков диагностики и умения работать с англоязычной документацией. Команда строит AI-продукты с акцентом на приватность и on-premises развёртывание; ожидается эффективная коммуникация в межфункциональных группах и умение быстро

Qualifications

  • Используйте Linux и CLI для навигации по файловой системе
  • Устранение проблем: собирайте факты, тестируйте гипотезы
  • Поиск и корреляция событий в логах
  • Опыт Grafana или аналогичных инструментов мониторинга
  • Понимание DNS, IP, портов, HTTP и API; диагностика ошибок
  • Чтение YAML и JSON; проверка окружения и конфигураций
  • Следование runbooks и умение работать автономно
  • Коммуникация на английском языке по техническим вопросам
  • Будет плюсом: Docker, Kubernetes, Helm, Git, CI/CD, Python/Bash, SQL, очереди сообщений
  • Опыт в микросервисной архитектуре и AI/LLM будет преимуществом

Responsibilities

  • Реагировать на оповещения, результаты проверки и запросы
  • Выяснять симптомы, оценивать влияние на пользователей, анализировать логи и конфигурации
  • Восстанавливать работу платформы по утверждённым runbooks, перезапуск и откат
  • Эскалировать инциденты к специалистам и передавать факты и выводы
  • Владеть инцидентами от сигнала до закрытия и держать команду в курсе статуса
  • Писать постмортемы, расследовать причины и перевести выводы в баг-репорты и задачи ремонта
  • Поддерживать актуальные runbooks и автоматизировать проверки
  • Обеспечивать ясную коммуникацию и инициативу в команде

Skills

Linux
CLI
Troubleshooting
Log correlation
Grafana
DNS, HTTP, APIs
YAML/JSON
Runbooks
English communication

Tools

Grafana
curl
Postman
Docker
Kubernetes
Helm
Git
CI/CD
Python/Bash
n8n

Job description

Описание

The team builds AI products for technology corporations, including devices and voice assistants, using speech technologies, NLP, generative AI, and voice-first agentic architecture with privacy-first and on-premises deployment.

Задачи
  • Respond to alerts, E2E check results, and requests
  • Clarify symptoms, assess user impact, analyze logs, service state, and configuration, test hypotheses, and identify failures
  • Restore the platform by executing approved runbooks, including restarts, redeployments, and failover to backup resources; confirm that the product is working again
  • Escalate to DevOps, SRE, SIP engineers, and developers when specialist expertise is needed, no suitable runbook exists, or a procedure does not produce the expected result; hand over gathered facts, actions, and diagnostic findings
  • Own incidents from the first signal to closure; record timelines and action outcomes, keep status current, track next steps, and keep everyone involved informed
  • Write postmortems, investigate causes and consequences, assess diagnosis and recovery, turn findings into bug reports and fix tasks, update runbooks, and help automate checks and recurring operations
Требования
  • Use Linux and the command line to navigate filesystems, inspect files and processes, and check resources and environment variables
  • Troubleshoot applications by reproducing problems, gathering facts, testing hypotheses, and distinguishing bugs from misconfiguration, unavailable dependencies, or resource shortages
  • Find and correlate log events across services using timestamps, request IDs, and other markers
  • Have hands-on experience with Grafana or similar monitoring tools; understand availability, error rate, response time, and resource consumption
  • Understand DNS, IP addressing, ports, HTTP, and APIs; check service reachability and diagnose status codes, timeouts, and authorization errors using curl or Postman
  • Read YAML and JSON, check environment variables, compare settings across environments, and understand how configuration affects connections and application behavior
  • Follow runbooks, verify preconditions and results, understand access limits, and know when to stop, roll back, or escape
  • Show technical curiosity, ask clear questions, communicate findings and status effectively, work independently, recognize when to involve specialists, prioritize calmly during incidents, take initiative, and collaborate
  • Read technical documentation, handle written communication, and discuss technical matters in English
  • Будет плюсом: Docker and Kubernetes, Helm, ConfigMaps and Secrets, Git, CI/CD, Python or Bash scripting, n8n automation, Claude Code or similar AI tools, SQL, microservices, message queues, telephony, and AI inference services
Условия
  • The team has built award-winning AI products for tech corporations
  • The team uses a cutting-edge stack including Speech Technologies, NLP, Generative AI (LLMs, diffusion models), and voice-first agentic architecture with privacy-first and on-premises deployment
  • High engineering standards, real ownership, and direct impact on production
  • Fast career progression in a senior-heavy team with a high volume of real problems
  • Startup pace with enterprise stability, real clients, real revenue, and no bureaucracy
  • 21 Vacation days, public holidays, and 5 sick days
  • Private English lessons via Preply
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

ai engineer for internal AI tools
ai engineer for internal AI tools

HireHi • United States

Hybrid
USD 101,000 - 169,000
Learning budget
Mental health sessions
On-site workshops
+1
fullstack software engineer ops ai platform
fullstack software engineer ops ai platform

HireHi • United States

Remote
USD 120,000 - 180,000
Equity package
Home office equipment sponsorship
Flexible vacation policy
+1
full stack developer for stablecoin payments
full stack developer for stablecoin payments

HireHi • United States

Remote
USD 120,000 - 180,000
Fully remote work in internationalteam
AI tooling paid by the company
Competitive salary based on review
IT Operations Manager
IT Operations Manager

HireHi • United States

Remote
USD 70,000 - 110,000
Remote work
Internal R&D opportunities
Paid time off
+1
application engineer for automation platform
application engineer for automation platform

HireHi • United States

Remote
USD 150,000 - 190,000
Medical, dental, and vision insurance
Flexible paid time off
Employee stock options
+1
devops engineer for hybrid cloud Data & AI platforms
devops engineer for hybrid cloud Data & AI platforms

HireHi • United States

Remote
USD 120,000 - 180,000
Competitive salary package
Career growth opportunities
Professional development: mentorship,技
+6
technical service operations tso
technical service operations tso

HireHi • United States

Remote
USD 120,000 - 190,000
Unlimited Flexible Time Off
Gym membership
Monthly train ticket
qa engineer (manual) for cybersecurity products
qa engineer (manual) for cybersecurity products

HireHi • United States

Remote
USD 65,000 - 100,000
Paid time off and sick leave
Direct impact on a large-scale globalB
devops engineer in Linux infrastructure
devops engineer in Linux infrastructure

HireHi • United States

Remote
USD 140,000 - 210,000
Гибкий график
Оплачиваемый отпуск 24 дня в год
Медицинское страхование
team lead platform engineering
team lead platform engineering

HireHi • United States

Remote
USD 124,000 - 180,000
Performance bonus
Private health insurance
Corporate pension
+5