Get a reply from this recruiter — a resume and cover letter tailored to exactly what they’re hiring for.
Recruitment Intelligence is seeking a senior cloud infrastructure engineer to own the Kubernetes and cloud foundation for the agentic platform, focusing on scalability, isolation, observability, and cost-per-request control.
You will define topologies for development, test and production, drive GitOps, IaC, identity and secret management, and partner with security and DevOps to ensure resilience and cost efficiency.
Own the Kubernetes and cloud foundation for the agentic platform, including scalability, isolation, observability, identity integration and cost-per-request control.
Define the target Kubernetes and cloud topology for development, test and production environments.
Own autoscaling, workload placement, network boundaries and runtime isolation for agent, tool and LLM-serving workloads.
Set infrastructure-as-code, GitOps, identity, secrets, gateway and certificate-management standards.
Define service-level indicators, operational dashboards, alerting and incident-response expectations for the platform.
Establish capacity and FinOps controls, including cost-per-request budgets and resource-consumption attribution.
Review platform changes for security, resilience, recoverability and vendor decoupling.
Minimum experience: 8+ years in cloud/platform infrastructure, including 3+ years leading production Kubernetes architecture or operations.
Production architecture and operations experience with Kubernetes, including autoscaling, networking, storage and workload isolation.
Strong Terraform or equivalent infrastructure-as-code capability and practical GitOps delivery experience.
Hands-on cloud platform experience, preferably Azure, covering compute, networking, managed identity and secure service exposure.
Observability engineering across metrics, logs and traces, with actionable service and cost dashboards.
Experience integrating identity providers, API gateways, secrets management and policy controls.
Understanding of LLM inference or model-serving workloads, including GPU/CPU scheduling and endpoint reliability.
AKS, Azure Monitor/Application Insights, managed identity and private networking.
Policy-as-code, service mesh, multi-cluster operations and disaster-recovery design.
FinOps practices for shared AI platforms and usage-based cost allocation.
Approved cloud/Kubernetes reference architecture and environment topology.
Version-controlled IaC and GitOps baseline with security and rollback controls.
Capacity, availability and cost-per-request dashboards with alert thresholds.
Operational runbooks for deployment, failure recovery and platform incidents.
Works with the Agent Runtime team, DevOps, Architecture Office, security/identity owners and the data-platform infrastructure lead.