Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes

Cisco Systems Inc

City of Westminster

In loco

GBP 80.000 - 105.000

Tempo pieno

3 giorni fa
Candidati tra i primi
Generatore di candidature

Una candidatura completa in un minuto — curriculum e lettera di presentazione personalizzati, pronti da inviare.

Supera i filtri ATS

Descrizione del lavoro

Cisco ThousandEyes is seeking a Site Reliability Engineer (DevSecOps) to join Production Engineering. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating with application teams to enhance reliability, performance, and security of the platform.

Key activities include deploying elastic AWS services, implementing scalable operations tooling, and participating in 24x7 incident response.

Competenze

  • 3+ years of experience in a related SRE/DevSecOps role.
  • Hands-on with production Kubernetes environments and containerized workloads.
  • Experience with Python or Go automation tooling and scripting.
  • Familiarity with security scanning tools and CI/CD security integration.
  • Strong communication and documentation skills.

Mansioni

  • Collaborate with software engineers to optimize architecture for availability, latency and reliability.
  • Design, deploy and maintain cloud-native services (AWS) for scalability and resilience.
  • Participate in 24x7 incident response and on-call rotations.
  • Expand CNCF tooling like Kubernetes, Prometheus, OpenTelemetry, and ArgoCD to boost reliability.
  • Automate production operations and implement infrastructure-as-code practices.
  • Develop automation for scalable service deployment, scale testing, and chaos testing.

Strumenti

Kubernetes
Service Mesh
Prometheus
OpenTelemetry
ArgoCD
Python
Go
Linux/Unix

Descrizione del lavoro

We are seeking a skilled Site Reliability Engineer (DevSecOps) in Production Engineering with a strong background in SaaS, operations and security. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.,

  • Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
  • Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
  • Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
  • Participate in and improve our 24x7 incident response and on-call rotation.
  • Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
  • Automate production operations to provide guardrails and continuous platform operation.
  • Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
  • Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
  • Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
  • Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
  • Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
  • Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.
    Hands-on experience deploying, operating, and troubleshooting containerized workloads in production Kubernetes environments.
  • Professional experience diagnosing and administrating Linux/Unix systems, including process management, file systems, and networking protocols (TCP/IP, DNS, HTTP).
  • Professional experience developing automation tooling, operational scripts, or backend services using Python or Go.
  • Practical experience building hardened container images and integrating automated security scanning tools (SAST, DAST, or container vulnerability scanners) into CI/CD pipelines., Familiarity with best practices for operating a large-scale, highly available enterprise platform.
  • 3+ years of experience in a related role.
  • Excellent communication and documentation skills.
  • Strong sense of ownership, drive, and attention to detail.
    Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network-even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences.
  • Cisco is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco's Networking, Security, Collaboration, and Observability portfolios., At Cisco, we're revolutionizing how data and infrastructure connect and protect organizations in the AI era - and beyond. We've been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint. Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you'll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. We are Cisco, and our power starts with you. Cisco is an Affir
Ottieni la revisione del curriculum gratis e riservata.

o trascina qui il file.

Similar jobs

Offerte di lavoro simili che vale la pena confrontare

Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes
Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes

Cisco • Greater London

Ibrido
GBP 90.000 - 130.000
Site Reliability Engineer, Infrastructure - ThousandEyes
Site Reliability Engineer, Infrastructure - ThousandEyes

Cisco • Greater London

In loco
GBP 90.000 - 130.000
Principal Software Engineer - ThousandEyes
Principal Software Engineer - ThousandEyes

Remote Worker LTD. • Greater London

In loco
GBP 95.000 - 130.000
Principal Software Engineer - ThousandEyes
Principal Software Engineer - ThousandEyes

Cisco Systems, Inc • Greater London

Ibrido
GBP 90.000 - 130.000
Principal Software Engineer
Principal Software Engineer

Cisco Systems, Inc. • Greater London

In loco
GBP 110.000 - 170.000
Senior Software Engineer - ThousandEyes
Senior Software Engineer - ThousandEyes

624 Cisco International Limited • Greater London

In loco
GBP 60.000 - 110.000
Principal Software Engineer
Principal Software Engineer

Cisco • Greater London

In loco
GBP 90.000 - 130.000
Senior Software Engineer - ThousandEyes
Senior Software Engineer - ThousandEyes

Cisco Systems, Inc • Greater London

In loco
GBP 80.000 - 110.000
Software Engineering Manager - ThousandEyes
Software Engineering Manager - ThousandEyes

Cisco • Greater London

In loco
GBP 90.000 - 130.000
Lead Platform Engineer — AI-Driven, Scale & Automation
Lead Platform Engineer — AI-Driven, Scale & Automation

Cisco Systems, Inc • Greater London

Ibrido
GBP 90.000 - 130.000