Senior Platform Engineer – Google Cloud Platform (DevOps & Agentic AI Infrastructure)

Whiztek Corp

Schaumburg (IL)

Hybrid

USD 140,000 - 170,000

Full time

4 days ago
Be an early applicant
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Whiztek Corp is seeking a Senior Platform Engineer to own the Google Cloud Platform infrastructure for agentic AI workloads. You will design, build, and operate the end-to-end platform, from browser to AI agents and services, setting architectural direction and balancing cost, security, and latency.

You will collaborate with AI engineers to secure, observe, and run agents in production, creating reusable Terraform/OpenTofu modules and robust CI/CD pipelines for zero-downtime releases in a hybrid

Qualifications

  • 6 years in platform, DevOps, cloud, or site reliability engineering.
  • 3+ years working deeply on Google Cloud Platform.
  • Hands-on production experience with Cloud Run including autoscaling, concurrency tuning, container optimization, min-instances and cold-start mitigation.
  • Practical experience configuring Global HTTP(S) Load Balancers with path-based routing, Identity-Aware Proxy, and an API management layer for authentication and rate limiting.
  • Strong command of Google Cloud Platform IAM: service account design, least privilege, Workload Identity, and authenticated service-to-service calls using ID tokens.
  • Advanced Terraform or OpenTofu skills, including reusable modules, remote state management, and multi-environment structures.
  • Experience building CI/CD pipelines with GitHub Actions and/or Cloud Build, including progressive delivery and automated rollback.
  • Experience with Secret Manager and runtime configuration patterns that avoid rebuilding containers per environment.
  • Solid observability experience with Cloud Logging, Cloud Monitoring, and Cloud Trace (or OpenTelemetry).
  • Working knowledge of Python, enough to read, containerize, and troubleshoot agent code.
  • A proven track record of making and defending architectural decisions, with clear communication for both technical and non-technical audiences.

Responsibilities

  • Own the architecture and lifecycle of our Google Cloud Platform infrastructure, including design, provisioning, security, cost management, and ongoing operations, without needing step-by-step direction.
  • Design deployment strategies for Python multi-agent systems built with Google ADK, choosing between the managed Agent Runtime/Agent Engine and Cloud Run based on workload needs.
  • Architect the full request path for agentic products: frontend hosting, Global HTTP(S) Load Balancing, Identity-Aware Proxy, API Gateway/Apigee, backend orchestrators, and agent services.
  • Enable real-time user experiences by supporting streaming responses from agents to the frontend, and reduce cold-start impact on latency-sensitive paths.
  • Build a zero-trust security model with least-privilege IAM, keyless authentication through Workload Identity, and ID-token-based authentication for agent-to-agent and service-to-service calls.
  • Design secure agent state and memory with Memory Bank and vector databases, keeping sessions persistent and tenant data strictly isolated.
  • Implement rate limiting, secret rotation, and audit logging across the agent call graph and external API integrations.
  • Build and maintain modular Terraform/OpenTofu for isolated Dev, QA, and Prod environments, with remote state, drift prevention, and safeguards against accidental production changes.
  • Create CI/CD pipelines (GitHub Actions, Cloud Build) with zero-downtime releases, canary and blue-green traffic splitting, and fast rollback.
  • Set up observability for agent workloads, including distributed tracing with Cloud Trace, structured logging, and custom metrics and alerts for agent reasoning loops, tool-call latency, and failures.
  • Write and defend architecture decision records, and translate technical trade-offs into clear recommendations for leadership.

Skills

Platform engineering
DevOps
GCP
Cloud Run
Terraform/OpenTofu
CI/CD
GitHub Actions
Cloud IAM
Observability
Python
Containerization
Architecture decisions

Tools

Cloud Run
API Gateway/Apigee
Terraform
OpenTofu
GitHub Actions
Cloud Build
Cloud Logging
Cloud Monitoring
Cloud Trace/OpenTelemetry
Vertex AI/ADK

Job description

Senior Platform Engineer -- Google Cloud Platform (DevOps & Agentic AI Infrastructure) Schaumburg, IL (Hybrid)
About the Role

We are launching customer-facing products powered by multi-agent AI systems, and we need a senior platform engineer to own the infrastructure that runs them. You will design, build, and operate our Google Cloud Platform platform end to end, from the user's browser, through our edge and API layers, down to the AI agents and the services they call. This is an ownership role: you will set architectural direction, make the trade-off calls on cost, security, scalability, and latency, and explain those decisions to both engineering and leadership.

You'll work closely with our AI engineers, who build agents in Python using Google's Agent Development Kit (ADK), and you'll be responsible for making those agents secure, observable, and reliable in production.

What You’ll Do
  • Own the architecture and lifecycle of our Google Cloud Platform infrastructure, including design, provisioning, security, cost management, and ongoing operations, without needing step-by-step direction.
  • Design deployment strategies for Python multi-agent systems built with Google ADK, choosing between the managed Agent Runtime/Agent Engine and Cloud Run based on each workload's needs.
  • Architect the full request path for agentic products: frontend hosting, Global HTTP(S) Load Balancing, Identity-Aware Proxy, API Gateway/Apigee, backend orchestrators, and agent services.
  • Enable real-time user experiences by supporting streaming responses (Server-Sent Events, WebSockets) from agents to the frontend, and reduce cold-start impact on latency-sensitive paths.
  • Build a zero-trust security model with least-privilege IAM, keyless authentication through Workload Identity, and ID-token-based authentication for agent-to-agent and service-to-service calls.
  • Design secure agent state and memory with Memory Bank and vector databases, keeping sessions persistent and tenant data strictly isolated.
  • Implement rate limiting, secret rotation, and audit logging across the agent call graph and external API integrations.
  • Build and maintain modular Terraform/OpenTofu for isolated Dev, QA, and Prod environments, with remote state, drift prevention, and safeguards against accidental production changes.
  • Create CI/CD pipelines (GitHub Actions, Cloud Build) with zero-downtime releases, canary and blue-green traffic splitting, and fast rollback.
  • Set up observability for agent workloads, including distributed tracing with Cloud Trace, structured logging, and custom metrics and alerts for agent reasoning loops, tool-call latency, and failures.
  • Write and defend architecture decision records, and translate technical trade-offs into clear recommendations for leadership.
What You’ll Bring
Required
  • 6 years in platform, DevOps, cloud, or site reliability engineering, including at least 3 years working deeply on Google Cloud Platform.
  • Hands-on production experience with Cloud Run, including autoscaling, concurrency tuning, container optimization, min-instances and cold-start mitigation, and serverless Network Endpoint Groups.
  • Practical experience configuring Global HTTP(S) Load Balancers with path-based routing, Identity-Aware Proxy, and an API management layer (Google Cloud Platform API Gateway or Apigee) for authentication and rate limiting.
  • Strong command of Google Cloud Platform IAM: service account design, least privilege, Workload Identity/Workload Identity Federation, and authenticated service-to-service calls using ID tokens.
  • Advanced Terraform or OpenTofu skills, including reusable modules, remote state management, and multi-environment structures.
  • Experience building CI/CD pipelines with GitHub Actions and/or Cloud Build, including progressive delivery (canary or blue-green) and automated rollback.
  • Experience with Secret Manager and runtime configuration patterns that avoid rebuilding containers per environment.
  • Solid observability experience with Cloud Logging, Cloud Monitoring, and Cloud Trace (or OpenTelemetry), including trace-context propagation across services.
  • Working knowledge of Python, enough to read, containerize, and troubleshoot agent code.
  • A proven track record of making and defending architectural decisions, with clear communication for both technical and non-technical audiences.
Strongly Preferred
  • Hands-on experience deploying LLM or AI agent workloads to production, ideally with Google's Gemini Enterprise Agent Platform (formerly Vertex AI), including Agent Engine, Agent Runtime, Agent Registry, or Memory Bank.
  • Familiarity with Google ADK or similar agent frameworks (LangGraph, CrewAI, and so on) and how multi-agent orchestration works.
  • Experience with vector databases (such as Vertex AI Vector Search, AlloyDB/pgvector, or similar) and multi-tenant data isolation.
  • Experience with streaming architectures such as SSE and WebSockets through load balancers and serverless platforms.
  • GKE experience, and the judgment to know when it's the right choice over serverless.
  • A Google Cloud Platform Professional Cloud Architect, Cloud DevOps Engineer, or Cloud Security Engineer certification
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

AI Platform Engineer
AI Platform Engineer

EPAM Systems Inc • United States

Remote
USD 120,000 - 190,000
Senior Staff Software Engineer, Cloud AI, Agent Governance
Senior Staff Software Engineer, Cloud AI, Agent Governance

Google • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
AI Platform Engineer
AI Platform Engineer

Aspen Dental • Chicago (IL)

Hybrid
USD 111,000 - 135,000
Health insurance
Dental insurance
Vision insurance
+3
Senior Staff Software Engineer, Cloud AI, Agent Governance
Senior Staff Software Engineer, Cloud AI, Agent Governance

Google Inc. • Sunnyvale (CA)

On-site
USD 262,000 - 364,000
AI Architect – AI Agent Platform
AI Architect – AI Agent Platform

Stellar Consulting Solutions, LLC • United States

On-site
USD 150,000 - 190,000
Senior Staff Software Engineer, Agent Protocol, Cloud AI
Senior Staff Software Engineer, Agent Protocol, Cloud AI

Google • Sunnyvale (CA)

On-site
USD 248,000 - 349,000
Bonus and equity options
Health and wellness benefits
Opportunities for growth and learning
Software Engineer - Level 3
Software Engineer - Level 3

Motion Recruitment Partners LLC • Richardson (TX)

On-site
USD 140,000 - 190,000
Forward Deployed Engineer IV, GenAI, Google Cloud
Forward Deployed Engineer IV, GenAI, Google Cloud

Google • Ann Arbor (MI)

On-site
USD 207,000 - 300,000
Health insurance
Dental insurance
Vision insurance
+8
GenAI Engineer
GenAI Engineer

Magicforce • Denver (CO)

On-site
USD 140,000 - 230,000
AI Architect
AI Architect

Cognizant • United States

On-site
USD 180,000 - 240,000