Get a reply from this employer — a resume and cover letter tailored to exactly what they’re hiring for.
Okta is seeking a Senior Site Reliability Engineer to evolve our internal Platform as a Service across the Dublin/EMEA region. You will build and operate a Kubernetes-based platform that hosts internal workflows and AI-driven pipelines, collaborating with global SRE teams.
You will mentor engineers, drive CI/CD improvements, implement observability, and manage day-2 operations, security posture, and runbooks while advancing platform thinking across the organization.
Comfortable working in a fast-moving, evolving environment with ambiguity around process and toolingEffective verbal and written communication skills, with the ability to collaborate across time zonesSolid experience with infrastructure-as-code (Terraform) and CI/CD pipelinesStrong hands-on experience with Kubernetes in production — deployment, networking, security, and troubleshootingExposure to or interest in building infrastructure that hosts AI/ML or agentic workflowsExperience operating large-scale internal infrastructure platforms in a public cloud, preferably AWS5+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure EngineeringExperience with observability platforms and monitoring tools (Grafana, Splunk, APM, or equivalent)Computer Science degree or related field, or equivalent experienceExposure to internal developer platform (IDP) concepts or “platform as a product” thinkingHands-on experience with AI agent orchestration, vector databases, model routing, or inference infrastructureMulti-cloud experience (AWS plus Azure or GCP)Service mesh technologies (Istio, Linkerd) or GitOps tooling (ArgoCD, Flux)Kubernetes certifications (CKA, CKS, CKAD) or equivalent cloud certificationsExperience administering an enterprise-scale SCM platform (GitHub, GitLab, or equivalent)