Senior DevOps Engineer, CI/CD Platform (Developer Experience)

Factorial HR

Madrid

Híbrido

EUR 90.000 - 130.000

Jornada completa

14 días+
Generador de candidaturas

No envíes un currículum genérico: crea un currículum y una carta de presentación adaptados a este puesto concreto.

Supera los filtros ATS

Ventajas ofrecidas por este puesto de trabajo

Alan health insurance
Wellhub fitness benefits
Cobee perks
Language classes
Office breakfast

Descripción de la vacante

Factorial is looking for a senior Kubernetes platform engineer to own the CI/CD Kubernetes clusters end to end. You will manage capacity, reliability, upgrades, security, and cost across a large self-hosted runner fleet in European datacentres.

You will scale and run Actions Runner Controller with runner scale sets, episodic pods, and image caches. A strong Linux, IaC, and observability background is essential to reduce queue wait and improve build efficiency.

Formación

  • 5+ years running production infrastructure with Kubernetes at the centre.
  • Deep Kubernetes knowledge: scheduling, limits, evictions, node pressure, DaemonSets, controllers and operators.
  • Strong Linux and container fundamentals (containerd/Docker, cgroups, namespaces).
  • Terraform and IaC with good module design and state hygiene.
  • CI/CD at the platform level: shared workflows, templates, runner architecture, build caching.
  • Comfortable with GitHub Actions and self-hosted runners.
  • Scripting: Bash, plus Python or Go.
  • Observability: metrics, logs, traces, dashboards; OpenTelemetry and Prometheus tooling.
  • Operational maturity: on-call, incident command, postmortems.
  • Capacity and cost awareness; size a fleet and explain the bill.
  • Clear written English; document what you build.

Responsabilidades

  • Rightsize runner tiers based on CPU/memory data.
  • Track flaky job from red check to container/runtime.
  • Add machines to the fleet: install, enrol, network, verify, document.
  • Reduce queue wait by identifying bottlenecks.
  • Upgrade a cluster or controller without disruption.
  • Collaborate with a product engineer on slow workflows.

Conocimientos

Kubernetes
Production Infra
Linux Fundamentals
Terraform
CI/CD
GitHub Actions
Scripting Bash
Python
Observability
On-call
Cost Awareness

Herramientas

Flux
Argo CD
Kustomize
Helm
Terraform

Descripción del empleo

Hey there, Kubernetes people!

Every engineer here waits on CI several times a day. We want someone who finds that wait personally annoying.

Why this role exists

Hundreds of developers push to a large Ruby on Rails monorepo, and every push lands on a runner fleet that the Developer Experience team builds, runs and keeps fast.

That fleet is self-hosted Kubernetes on bare metal in European datacentres, in the low hundreds of machines, running episodic GitHub Actions runners that peak above a thousand concurrent pods. We own all of it: the hardware, the clusters, the runner images, the caches and the delivery pipelines on top. We add nodes by hand rather than letting an autoscaler do it, so capacity planning is genuinely part of the job.

The leverage is unusual. Shave a minute off the average build and every engineer in the company gets that minute back, several times a day.

The mission
  • Own the CI and CD Kubernetes clusters end to end: capacity, reliability, upgrades, security and cost.
  • Run the self-hosted GitHub Actions platform at scale with Actions Runner Controller. Runner scale sets, episodic pods, docker-in-docker, and the runner images themselves.
  • Keep the data services our tests depend on fast and healthy. MySQL, Redis and ClickHouse come up per job, alongside Rails application containers.
  • Design and tune the caching that makes builds quick: node-local overlay, image and artifact caches we build ourselves, plus pull-through registry mirrors.
  • Bring queue wait and build duration down, with measurement behind it. Instrument the platform, set targets, and show the improvement.
  • Manage it all as code. Terraform planned and applied from pull requests, GitOps delivery with Flux and Argo CD, Kustomize and Helm for manifests.
  • Provision and operate bare metal. Linux, networking across datacentre segments, and troubleshooting that sometimes ends up at the disk or the kernel.
  • Lead incident response for the platform, run blameless postmortems, and turn each one into an alert, a guardrail or a runbook.
  • Treat the platform as a product whose users are engineers. Talk to them, watch where they get stuck, and build paths that make the right thing easy.
Your day to day
  • Rightsizing runner tiers against a week of CPU and memory data, then opening the pull request that changes them.
  • Tracking a flaky job from a red check down to the container runtime, the kernel, or a clock that drifted.
  • Adding machines to the fleet: install, enrol, network, verify, document.
  • Cutting p95 queue wait by working out which stage actually blocks.
  • Upgrading a cluster or a controller without anyone noticing.
  • Pairing with a product engineer on a workflow that has been slow for a reason nobody has looked at yet.
Your profile
  • 5+ years running production infrastructure, with Kubernetes at the centre of it.
  • Real depth in Kubernetes: scheduling, requests and limits, evictions, node pressure, DaemonSets, controllers and operators. You have debugged a cluster that was lying to you.
  • Strong Linux and container fundamentals. containerd or Docker internals, cgroups, namespaces, storage drivers, networking.
  • Terraform and infrastructure as code, with a feel for module design and state hygiene.
  • CI/CD at the platform level: shared workflows, reusable templates, runner architecture, build caching.
  • Comfortable with GitHub Actions, including self-hosted runners.
  • You script to automate. Bash, plus at least one of Python or Go.
  • Observability as a working method. Metrics, logs, traces, dashboards and alerts that people trust, with OpenTelemetry and Prometheus style tooling.
  • Operational maturity. On-call, incident command, postmortems, and the discipline to chase causes.
  • Capacity and cost awareness. You can size a fleet and explain the bill.
  • Clear written English. You document what you build, because the next person on call will not be you.
Bonus points
  • Large Ruby on Rails test suites, and what it takes to parallelise them honestly.
  • Operating MySQL, Redis or ClickHouse yourself.
  • Bare metal: provisioning, hardware failure, and the different rhythm it has compared to cloud.
  • Overlay and WireGuard style mesh networking, VLANs, and Kubernetes CNIs such as Cilium.
  • Building internal developer tooling, portals or platform APIs.
  • Supply chain and build security. Runner isolation, secret scoping, image provenance.
  • Working across more than one cloud. We use AWS, Azure and Cloudflare alongside our own hardware.
  • Contributions to open source infrastructure projects.
How we work

At Factorial, we believe the best products are built when people come together, in person, to collaborate, challenge ideas, and move fast. That's why our Engineering Team follows an office-first, flexible approach.

We work on‑site several days a week (80%), using that time to connect, align, and innovate as a team. However, we also support remote work when it makes sense (20%) for deep focus or personal needs.

At Factorial we don't evaluate you by years of experience but by your knowledge and skills.

We use our Career Path with a rubric framework where we define what is expected for each experience level and skill. This framework allows you to know your current level and what you need to keep growing. We love transparency, so our Career Path includes the salaries for each level, which we share during the first interview to ensure alignment.

Our Hiring Process
  • Intro Call: Chat with our Talent Partner about your journey and goals.
  • Hiring Manager Interview: Deep dive into your product mindset, teamwork, and problem-solving.
  • Technical Interview: A collaborative technical conversation with the team based on real-life problems.
  • Final Coffee Chat: A relaxed conversation with our CTO & VP of Engineering to explore our vision, culture, and your growth.

And that’s it! Feel free to request any other conversations you want, with team members or to address specific concerns at any time in the process. The whole process is remote, using videoconferencing tools!

About Factorial

At Factorial, we’re building the leading AI Business Management Software for companies of all sizes. Our platform centralizes key workflows across HR, Finance, Talent, Operations, and IT, freeing teams from manual processes so they can focus on what really matters: leading, growing, and taking care of their people. With over 1,500 employees across Europe, Asia, Africa, and America. We are one of Europe’s fastest‑growing SaaS companies, proudly headquartered in Barcelona. If you're excited to shape the future of business management technology, we’d love to meet you.

We believe in diverse talent: we welcome applicants from all backgrounds and strongly encourage people of diverse experiences and identities to apply.

We believe in inclusion: we are committed to equal opportunities and actively promote workplace inclusion of people with disabilities. If you would like to learn more about our inclusive recruitment processes, you are welcome to indicate so optionally and we will share additional information with you.

Our Values
  • We own it: We take responsibility for every project. We make decisions, not excuses.
  • We learn and teach: We're dedicated to learning something new every day and, above all, share it.
  • We partner: Every decision is a team decision. We trust each other.
  • We grow fast: We act fast. We think that the worst mistake is not learning from them.

Wanna learn more about us? Check our website!

Perks of being part of Factorial
  • High growth, multicultural and friendly environment
  • Alan as private health insurance
  • Healthy life with Wellhub (Gyms, pools, outdoor classes)
  • Save expenses with Cobee
  • Language classes
  • Breakfast in the office and organic fruit
  • Food discounts with Nora
  • Pet Friendly
  • Much more that we'll spill during the interview process!
Consigue la evaluación confidencial y gratuita de tu currículum.
o arrastra y suelta tu archivo aquí
Similar jobs

Puestos de trabajo similares que vale la pena comparar

Senior DevOps Engineer, CI/CD Platform (Developer Experience)
Senior DevOps Engineer, CI/CD Platform (Developer Experience)

Factorial HR • Oleiros

Híbrido
EUR 70.000 - 110.000
Private health insurance
Wellhub access
Cobee perks
+5
Senior DevOps Engineer, CI/CD Platform (Developer Experience)
Senior DevOps Engineer, CI/CD Platform (Developer Experience)

Factorial HR • Barcelona

Híbrido
EUR 90.000 - 130.000
Private health insurance
Wellhub gym benefits
Cobee benefits
+4
Engineering Team Lead
Engineering Team Lead

Factorial HR • Barcelona

Presencial
EUR 75.000 - 110.000
High growth team
Private health insurance
Wellhub access
+5
Technical Project Manager - Infrastructure Team
Technical Project Manager - Infrastructure Team

Factorial HR • Barcelona

Presencial
EUR 50.000 - 70.000
Alan as private health insurance
Wellhub gym access
Cobee
+4
Staff Ai Engineer - Api & Integrations Team (Barcelona)
Staff Ai Engineer - Api & Integrations Team (Barcelona)

Factorial Hr • Madrid

Híbrido
EUR 65.000 - 90.000
High growth environment
Private health insurance
Wellbeing benefits
+1
Senior Product Software Engineer - Operations Domain
Senior Product Software Engineer - Operations Domain

Factorial HR • Barcelona

Híbrido
EUR 90.000 - 120.000
Alan as private health insurance
Wellhub gym membership
Cobee benefits platform
+4
AI Software Developer - Founding role
AI Software Developer - Founding role

Factorial HR • Barcelona

Híbrido
EUR 90.000 - 120.000
Private health insurance
WellHub gym access
Cobee benefits
+4
Senior Product Engineer - Operations Domain
Senior Product Engineer - Operations Domain

Factorial • Barcelona

Presencial
EUR 70.000 - 110.000
Alan private health insurance
Cobee benefits platform
Language classes
+2
Engineering Manager
Engineering Manager

Factorial • Oleiros

Híbrido
EUR 90.000 - 135.000
High growth environment
Private health insurance
Wellhub gym benefits
+3
Founding Engineer
Founding Engineer

Factorial • Barcelona

Presencial
EUR 70.000 - 110.000
Private health insurance
Wellhub gym
Cobee
+2