No envíes un currículum genérico — crea un currículum y una carta de presentación adaptados a este puesto concreto.
Factorial ищет инженера по инфраструктуре, ответственного за CI/CD и Kubernetes, нацеленного на масштабируемость и надёжность.
Вы будете поддерживать self-hosted GitHub Actions, тестовые сервисы (MySQL, Redis, ClickHouse) и дизайн кэширования сборки. Ваша роль — взаимодействовать с инженерами и документировать системы.
Ожидается 5+ лет опыта, знание IaC, Terraform, OpenTelemetry и Prometheus; гибридное место работы в Барселоне.
Factorial provides AI for businesses that connects information about company operations and turns that context into action, supporting tasks such as candidate screening, expense management, answering questions, and report creation. It serves more than 16,000 companies across 112 countries.
Own the CI and CD Kubernetes clusters end to end, including capacity, reliability, upgrades, security, and cost; Run the self-hosted GitHub Actions platform at scale using Actions Runner Controller, runner scale sets, ephemeral pods, Docker-in-Docker, and runner images; Keep test data services fast and healthy, including MySQL, Redis, and ClickHouse running per job alongside Rails application containers; Design and tune build caching, including node-local overlay, custom image and artifact caches, and pull-through registry mirrors; Measure and reduce queue wait and build duration by instrumenting the platform, setting targets, and demonstrating improvements; Manage infrastructure as code using Terraform pull requests, GitOps delivery with Flux and Argo CD, and Kustomize and Helm manifests; Provision and operate bare metal, Linux, and networking across datacentre segments, troubleshooting down to disk or kernel level; Lead platform incident response, run blameless postmortems, and turn findings into alerts, guardrails, or runbooks; Treat the platform as a product for engineers by talking to users, observing where they get stuck, and building easy-to-follow paths; Right-size runner tiers using CPU and memory data and submit pull requests to change them; Trace flaky jobs from failed checks to the container runtime, kernel, or clock drift; Add fleet machines by installing, enrolling, networking, verifying, and documenting them; Reduce p95 queue wait by identifying the blocking stage; Upgrade clusters or controllers without disrupting users; Pair with product engineers to investigate slow workflows.