Platform Engineer (Agent Runtime) Taguig, Taguig, Philippines

Netskope, Inc.

Taguig

On-site

PHP 1,800,000 - 3,120,000

Full time

14 days+
Application generator

Stand out for this role — generate a tailored resume and cover letter in about a minute.

Get past ATS filters

Job summary

Netskope, Inc. is seeking a Senior Agent Runtime Engineer to own the health of the agent execution environment, ensuring strict session isolation and efficient resource use.

You’ll collaborate with engineers building the agents and act as the first line in production when issues emerge in real environments. Responsibilities include tuning cold-start and concurrency settings, investigating production anomalies, and maintaining observability to distinguish between “up” and “working correctly.”

Qualifications

  • 5+ years in platform, SRE, or infrastructure role with production workloads on AWS.
  • Hands-on experience with AWS observability tooling to diagnose root causes from traces and logs.
  • Proficient deploying infrastructure via IaC end-to-end (Terraform, CDK or similar).
  • Knowledge of compute isolation concepts and edge cases of platform guarantees.

Responsibilities

  • Own health of the agent execution environment, including session isolation and resource limits.
  • Investigate production vs testing discrepancies and partner with engineers to resolve root causes.

Skills

Platform SRE experience
AWS observability tooling
IaC deployment
Cost analysis
Incident response
Root-cause tracing

Tools

CloudWatch
X-Ray
Terraform
CDK
Kubernetes
CI/CD tooling

Job description

As a Senior Agent Runtime Engineer, you’ll own the health of the systems agents run on: keeping sessions isolated from each other, catching the failure modes that only show up under real concurrency, and making sure a slow or expensive agent gets caught before it becomes everyone’s problem. You’ll work closely with the engineers building the agents themselves, and you’ll be the first call when something behaves strangely in a real environment but not in a test one.

Skills and Competencies
  • Own the health of the agent execution environment day to day — session isolation, resource limits, and catching the failure modes that only appear once real concurrency and real traffic patterns show up, not just in a clean test run.
  • Investigate and resolve cases where an agent behaves differently in production than it did in testing, working directly with the Agent Engineer or Quality Engineer who built it to figure out whether the problem is the agent’s design or the environment it’s running in.
  • Tune cold-start and concurrency settings for the platform’s critical-path functions, and review them on a regular cadence as usage patterns shift rather than setting them once and forgetting them.
  • Understand the platform’s session isolation model well enough to reason about its limits — where isolation is guaranteed by the underlying compute layer, and where the platform has to add its own controls because that guarantee doesn’t fully hold.
  • Build and maintain the observability that lets someone answer "is this agent actually working correctly," not just "is it technically up" — tracing, behavioral drift detection, and quality signals sitting alongside the usual metrics and logs.
  • Track per‑agent and per‑session cost and efficiency, and flag agents that are burning more tokens, calling more tools, or running longer than the task should reasonably require.
  • Deploy and manage runtime‑layer infrastructure resources through existing CI/CD pipelines as needed — provisioned concurrency settings, runtime‑specific IAM roles, observability configurations, etc.
  • Run and improve the tests that validate isolation actually holds — for example, confirming that one agent’s session genuinely can’t reach or affect another’s, not just assuming it because the platform is supposed to guarantee it.
Must‑Have
  • At least 5 years in a platform, SRE, or infrastructure engineering role, with real production experience running serverless or containerized workloads on AWS (Lambda, Fargate, Docker, Kubernetes or equivalent) at meaningful scale.
  • Hands‑on experience with AWS observability tooling (CloudWatch, X‑Ray, or a comparable distributed tracing stack) — able to go from "something’s wrong" to a root cause using traces and logs, not just dashboards.
  • Real experience deploying and managing infrastructure through CI/CD independently — comfortable owning IaC changes (Terraform, CDK, or similar) end to end rather than handing them off to someone else.
  • Working understanding of compute isolation concepts (containers, microVMs, or similar sandboxing models) and where their guarantees actually stop, since a lot of this role is reasoning about the edges of what a platform promises versus what it might not fully cover on its own.
  • Comfort investigating cost and performance problems at a granular level — able to trace an unexpectedly expensive or slow workload back to a specific cause, not just flag that costs went up.
  • Strong incident response instincts: staying calm and methodical while root‑causing a live production issue, and following through with an actual fix rather than a workaround.
Strong Advantage
  • Direct experience with AWS Bedrock, SageMaker, or any managed AI/agent runtime platform.
  • Experience with cold‑start or concurrency tuning specifically (Lambda provisioned concurrency, container warm pools, or similar).
  • Exposure to a security‑conscious or regulated environment where infrastructure changes go through a formal review or approval process.
  • Familiarity with LLM‑specific cost drivers (token pricing, tool‑call volume, model tiering) even if it wasn’t the primary focus of a past role.

Netskope is committed to implementing equal employment opportunities for all employees and applicants for employment. Netskope does not discriminate in employment opportunities or practices based on religion, race, color, sex, marital or veteran statues, age, national origin, ancestry, physical or mental disability, medical condition, sexual orientation, gender identity/expression, genetic information, pregnancy (including childbirth, lactation and related medical conditions), or any other characteristic protected by the laws or regulations of any jurisdiction in which we operate.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Platform Engineer (Agent Runtime)
Platform Engineer (Agent Runtime)

Netskope • Taguig

On-site
PHP 1,200,000 - 1,800,000
Platform Engineer (Agent Release)
Platform Engineer (Agent Release)

Netskope • Taguig

On-site
PHP 1,200,000 - 2,000,000
AI Agent Developer
AI Agent Developer

Netskope • Taguig

On-site
PHP 1,000,000 - 1,800,000
Senior Platform Runtime Engineer
Senior Platform Runtime Engineer

Netskope • Taguig

On-site
PHP 1,200,000 - 1,800,000
Senior Platform Engineer – Agent Runtime
Senior Platform Engineer – Agent Runtime

Netskope, Inc. • Taguig

On-site
PHP 1,800,000 - 3,120,000
Head of Platform Engineering (Remote)
Head of Platform Engineering (Remote)

Cyberbacker Careers • Cebu City

On-site
PHP 1,200,000 - 1,800,000
Platform Security Engineer
Platform Security Engineer

Causa Prima • España

On-site
PHP 3,075,000 - 4,613,000
Head of Platform Engineering (Fully Remote)
Head of Platform Engineering (Fully Remote)

Cyberbacker Careers • Iligan

On-site
PHP 1,800,000 - 3,000,000
Head of Platform Engineering (Work From Home)
Head of Platform Engineering (Work From Home)

Cyberbacker Careers • Metro Manila

On-site
PHP 2,500,000 - 4,500,000
Head of Platform Engineering (WFH)
Head of Platform Engineering (WFH)

Cyberbacker Careers • Mandaluyong

On-site
PHP 1,800,000 - 3,000,000