Site Reliability Engineer

Cacheflow

Milano

In loco

EUR 90.000 - 130.000

Tempo pieno

14 giorni+

Ricevi più risposte dai datori di lavoro

Invia un CV specifico per questa offerta in pochi minuti.

Descrizione del lavoro

Kong Inc. is seeking an experienced Site Reliability Engineer to join our SRE team. You will design and operate scalable, reliable cloud infrastructure using Terraform, Ansible, and modern observability stacks.

You’ll implement monitoring and incident response processes, drive blameless post-mortems, automate toil away, and collaborate with developers to bake reliability into the software lifecycle. A strong focus on uptime and security will guide your daily work.

Competenze

  • Experience operating production workloads on a major cloud provider (AWS, GCP, or Azure).
  • Proficiency in Golang, Python, or Bash.
  • Hands-on experience with Docker and Kubernetes.
  • Knowledge of Infrastructure as Code principles and Terraform.
  • Familiarity with CI/CD concepts and pipelines (GitLab CI, Jenkins).
  • Understanding of modern observability stacks (Prometheus, Grafana, ELK).

Mansioni

  • Build and maintain core infrastructure as code using Terraform and Ansible.
  • Implement robust monitoring, logging, and alerting to meet 99.99% uptime.
  • Resolve production incidents with blameless post-mortems.
  • Write automation to reduce toil and enable self-service for engineering teams.
  • Collaborate with developers to embed reliability into the software lifecycle.
  • Contribute to capacity planning, disaster recovery drills, and security hardening.
  • Participate in fair on-call rotation to keep the platform available.

Conoscenze

Cloud provider experience (AWS/GCP/AZ)
Programming in Golang/Python/Bash
Docker/Kubernetes
Infrastructure as Code (Terraform)
CI/CD concepts & pipelines
Observability stacks (Prometheus/Graf/

Strumenti

Terraform
Docker
Kubernetes
GitLab CI
Jenkins

Descrizione del lavoro

Are you ready to unlock intelligence?

If you don’t think you meet all of the criteria below but are still interested in the job, please apply. Nobody checks every box - we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.

The Mission

The Site Reliability Engineering team is the backbone of Kong's cloud services, responsible for architecting and operating the large-scale infrastructure that powers our customers' most critical applications. Our mission is to achieve world-class reliability and performance, enabling our product engineering teams to ship features with velocity and confidence. We are the guardians of uptime and the champions of developer delight.

What You’ll Do
  • Build and maintain our core infrastructure as code using tools like Terraform and Ansible.

  • Implement robust monitoring, logging, and alerting systems to ensure our services meet and exceed 99.99% uptime.

  • Resolve production incidents through systematic debugging, and drive the blameless post-mortem process to prevent recurrence.

  • Write automation to reduce operational toil, improve system efficiency, and enable self-service for engineering teams.

  • Collaborate with developers to embed reliability and scalability best practices directly into the application lifecycle.

  • Contribute to our capacity planning, disaster recovery drills, and security hardening processes.

  • Participate in a fair and sustainable on-call rotation to ensure our platform is always available.

What You’ll Bring
  • Experience operating production workloads on a major cloud provider (AWS, GCP, Azure).

  • Proficiency in at least one programming or scripting language, such as Golang, Python, or Bash.

  • Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes).

  • Knowledge of Infrastructure as Code principles and tools (Terraform is a plus).

  • Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins).

  • An understanding of modern observability stacks (e.g., Prometheus, Grafana, ELK).

About Kong:

Kong Inc., a leading developer of API and AI connectivity technologies, is building the infrastructure that powers the agentic era. Trusted by the Fortune 500 and startups alike, Kong's unified API and AI platform, Kong Konnect, enables organizations to secure, manage, accelerate, govern, and monetize the flow of intelligence across APIs and AI models. For more information, visit www.konghq.com.

Ottieni la revisione del curriculum gratis e riservata.
o trascina qui il file.
Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Site Reliability Engineer
Site Reliability Engineer

KONG • Milano

In loco
EUR 70.000 - 110.000
Senior Software Engineer, Konnect Control Plane
Senior Software Engineer, Konnect Control Plane

Kong • Italia

In loco
EUR 70.000 - 95.000
Senior Incident Response Engineer
Senior Incident Response Engineer

Kong • Turbigo

In loco
EUR 60.000 - 80.000
Senior Engineering Manager - Test Function
Senior Engineering Manager - Test Function

Kong • Milano

In loco
EUR 130.000 - 180.000
Senior Engineering Manager - Test Function
Senior Engineering Manager - Test Function

Cacheflow • Milano

In loco
EUR 150.000 - 190.000
Site Reliability Engineer: Build Fault-Tolerant Cloud Infra
Site Reliability Engineer: Build Fault-Tolerant Cloud Infra

KONG • Milano

In loco
EUR 70.000 - 110.000
SRE: Build Scalable, Reliable Cloud Infra
SRE: Build Scalable, Reliable Cloud Infra

Cacheflow • Milano

In loco
EUR 90.000 - 130.000
Senior Security Engineer
Senior Security Engineer

Kong Inc • Italia

Ibrido
EUR 70.000 - 90.000
Software Engineer AI Gateway
Software Engineer AI Gateway

Kong • Turbigo

Ibrido
EUR 50.000 - 70.000
Software Engineer AI Gateway
Software Engineer AI Gateway

Kong • Milano

Ibrido
EUR 45.000 - 65.000