Описание
The team builds AI products for technology corporations, including devices and voice assistants, using speech technologies, NLP, generative AI, and voice-first agentic architecture with privacy-first and on-premises deployment.
Задачи
- Respond to alerts, E2E check results, and requests
- Clarify symptoms, assess user impact, analyze logs, service state, and configuration, test hypotheses, and identify failures
- Restore the platform by executing approved runbooks, including restarts, redeployments, and failover to backup resources; confirm that the product is working again
- Escalate to DevOps, SRE, SIP engineers, and developers when specialist expertise is needed, no suitable runbook exists, or a procedure does not produce the expected result; hand over gathered facts, actions, and diagnostic findings
- Own incidents from the first signal to closure; record timelines and action outcomes, keep status current, track next steps, and keep everyone involved informed
- Write postmortems, investigate causes and consequences, assess diagnosis and recovery, turn findings into bug reports and fix tasks, update runbooks, and help automate checks and recurring operations
Требования
- Use Linux and the command line to navigate filesystems, inspect files and processes, and check resources and environment variables
- Troubleshoot applications by reproducing problems, gathering facts, testing hypotheses, and distinguishing bugs from misconfiguration, unavailable dependencies, or resource shortages
- Find and correlate log events across services using timestamps, request IDs, and other markers
- Have hands-on experience with Grafana or similar monitoring tools; understand availability, error rate, response time, and resource consumption
- Understand DNS, IP addressing, ports, HTTP, and APIs; check service reachability and diagnose status codes, timeouts, and authorization errors using curl or Postman
- Read YAML and JSON, check environment variables, compare settings across environments, and understand how configuration affects connections and application behavior
- Follow runbooks, verify preconditions and results, understand access limits, and know when to stop, roll back, or escape
- Show technical curiosity, ask clear questions, communicate findings and status effectively, work independently, recognize when to involve specialists, prioritize calmly during incidents, take initiative, and collaborate
- Read technical documentation, handle written communication, and discuss technical matters in English
- Будет плюсом: Docker and Kubernetes, Helm, ConfigMaps and Secrets, Git, CI/CD, Python or Bash scripting, n8n automation, Claude Code or similar AI tools, SQL, microservices, message queues, telephony, and AI inference services
Условия
- The team has built award-winning AI products for tech corporations
- The team uses a cutting-edge stack including Speech Technologies, NLP, Generative AI (LLMs, diffusion models), and voice-first agentic architecture with privacy-first and on-premises deployment
- High engineering standards, real ownership, and direct impact on production
- Fast career progression in a senior-heavy team with a high volume of real problems
- Startup pace with enterprise stability, real clients, real revenue, and no bureaucracy
- 21 Vacation days, public holidays, and 5 sick days
- Private English lessons via Preply