Staff AI Platform Engineer, Infrastructure Services

SentinelOne

United States

On-site

USD 156,000 - 215,000

Full time

10 days ago

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

RSUs
ESPP
Flexible time off
Paid holidays
Parental leave
Medical, dental, vision
401(k) with match
Life and disability insurance
Home office allowance
Mobile phone reimbursement
Wellness reimbursement
Wellness coach

Job summary

SentinelOne seeks a Staff AI Platform Engineer in Infrastructure Services to own the AI Gateway infrastructure, including Kong AI Gateway, authentication, rate-limiting, and observability. You will drive reliability and architectural direction across the platform stack.

Collaborating with security, DevEx, and product engineering, you will shape CI/CD, GitOps, and artifact pipelines, while enabling self-hosted model serving and AI tooling adoption at scale.

Qualifications

  • 8+ years in platform, infrastructure, or DevOps engineering with end-to-end ownership.
  • Hands-on with API gateway tech (Kong/Envoy/Apigee) and AI/LLM gateway patterns.
  • Kubernetes and GitOps experience across dev/gov/prod environments.
  • Strong CI/CD and build infra know-how (Jenkins, runners, GitHub Actions).
  • Artifact/package mgmt and source control (Artifactory, Xray, GitHub Enterprise).
  • Infra-as-code (Terraform) and cloud fluency (AWS/EKS).
  • Experience deploying self-hosted LLM inference stacks and GPU scheduling.

Responsibilities

  • Architect, harden, and scale Kong AI Gateway deployment with auth, tiering, and caching.
  • Drive reliability, incident response, and observability across gateway issues.
  • Design platform solutions beyond the gateway with CI/CD and GitOps tooling.
  • Evaluate AI developer tooling and build buy/build recommendations.
  • Mentor engineers; set architectural standards for AI infrastructure.
  • Collaborate cross-functionally with security, DevEx, and product teams.
  • Operate self-hosted model serving infra and manage GPU capacity/costs.
  • Support model lifecycle: versioning, evaluation, and safe rollouts.

Skills

Platform engineering
DevOps
Kubernetes
CI/CD
GitOps
Okta/OIDC
AI tooling
Communication

Tools

Kong
Envoy
Apigee
Jenkins
ArgoCD
Artifactory
Xray
GitHub Enterprise
Terraform
AWS EKS
vLLM
NVIDIA Triton
Ollama

Job description

Our Purpose

At SentinelOne, we are driven by a clear purpose: to give the advantage to those who secure our future. As AI reshapes how organizations build, operate, and innovate, the responsibility to protect them becomes more critical than ever. When you join SentinelOne, your work helps protect global enterprises, critical infrastructure, and the technologies shaping tomorrow. If you are motivated by meaningful challenges and want your impact to be real, measurable, and global, you will find purpose here.

About Us

SentinelOne is a company at the intersection of AI and security, pioneering a new operating model for cybersecurity. Our AI-native platform unifies protection across endpoint, cloud, identity, data, and AI systems to deliver autonomous detection and response with clarity and speed. By combining real-time analytics, intelligent automation, and a unified data foundation, we reduce noise, simplify complexity, and empower security teams to focus on what truly matters.

Our teams are builders, problem-solvers, and innovators committed to shaping the future of security. If you are excited to solve hard problems alongside talented, mission-driven people, we invite you to help us build a safer future for humanity.

What Are We Looking For?

We’re looking for people who are relentlessly curious and committed to continuous learning. AI is reshaping every function across our business, and we enable every team member, regardless of role or level, to build fluency in AI tools and concepts. Those who thrive here actively seek out new solutions, experiment thoughtfully, and apply what they learn to drive better, faster, smarter outcomes.

As a Staff AI Platform Engineer, Infrastructure Services, you will be tasked with taking ownership of our AI Gateway infrastructure (built on Kong AI Gateway), the system that authenticates, routes, rate-limits, and monitors AI coding assistant traffic org-wide, while also being fluent enough across our broader platform stack to design solutions that span the two. This is a high-autonomy, high-scope role: you will set technical direction for AI infrastructure, drive incident response and reliability work, and partner closely with the engineers who own our CI/CD, GitOps, and artifact systems rather than working in isolation from them.

What Will You Do?

Primary responsibilities include:

  • Work on the AI Gateway platform: architect, harden, and scale our Kong AI Gateway deployment (Konnect Hybrid on KCP/EKS), including auth (Okta/OIDC), consumer tiers and budgets, rate limiting, semantic caching, and observability.
  • Lead reliability and incident response: drive root-cause analysis and remediation for gateway issues (timeouts, latency, capacity, failover) and build the monitoring/alerting needed to catch them before users do.
  • Design across the platform, not just the gateway: work fluently with our CI/CD (Jenkins, JPAAS), GitOps and Kubernetes deployment tooling (ArgoCD across dev/gov/prod), artifact management (Artifactory/Xray), GitHub Enterprise administration, and GitHub Actions runner fleet, so that AI infrastructure decisions account for how the rest of the platform actually works.
  • Evaluate and roll out AI developer tooling: run structured pilots and adoption efforts for tools like AI-assisted PR review (Qodo) and engineering metrics platforms (LinearB), and make clear build-vs-buy recommendations.
  • Set technical direction and mentor: define architecture and standards for AI infrastructure, review designs across the team, and raise the bar for other engineers working in this space.
  • Partner cross-functionally: work directly with security, DevEx, and product engineering teams consuming the gateway to translate their needs into platform capabilities.
  • Host and serve local models: stand up and operate self-hosted/open-weight model serving infrastructure (e.g. vLLM, NVIDIA Triton/NIM, TGI, Ollama) for workloads where routing to an external provider isn't the right fit, including GPU capacity planning, autoscaling, and cost/performance tuning.
  • Support the broader model lifecycle: help build LLMOps practices such as model versioning, evaluation, and safe rollout, plus supporting infrastructure for retrieval-augmented generation (vector stores, embedding pipelines) as use cases mature.
  • Track usage and cost: build observability into token usage, latency, and spend across both API-based and self-hosted models so the business can see what AI infrastructure actually costs.
What Skills and Knowledge Will You Bring?

Ideal candidates will have:

  • 8 or more years of experience in platform, infrastructure, or DevOps engineering, with a track record of owning systems end-to-end in production.
  • Hands-on experience with API gateway technologies (Kong, Envoy, Apigee, or similar); direct experience with AI/LLM gateway patterns (rate limiting, semantic caching, prompt/response observability) is a strong plus.
  • Strong Kubernetes and GitOps experience (ArgoCD or comparable), and comfort operating across multiple environments (dev, gov, prod).
  • Solid CI/CD background: Jenkins pipeline design and administration, build infrastructure, and runner/agent fleet management (GitHub Actions runners or equivalent).
  • Experience with artifact and package management systems (Artifactory, Xray, or similar) and source control platform administration (GitHub Enterprise).
  • Working knowledge of infrastructure-as-code (Terraform) and cloud platforms (AWS/EKS).
  • Experience deploying and operating self-hosted LLM inference stacks (vLLM, NVIDIA Triton/NIM, TGI, Ollama, or similar) and GPU-backed infrastructure, including Kubernetes GPU scheduling and autoscaling.
  • Familiarity with LLMOps practices: model versioning, evaluation harnesses, and usage/cost observability across API-based and self-hosted models.
  • Track record of setting technical direction, driving cross-team initiatives, and mentoring other engineers; this role has significant scope and minimal day-to-day oversight.
  • Clear, proactive communicator who can explain infrastructure trade-offs to both engineers and non-technical stakeholders.
  • Experience operating LLM/AI-assisted developer tooling at scale (Claude Code, Copilot, or similar) inside an enterprise is preferred.
  • Familiarity with Okta/OIDC and enterprise auth patterns for internal platforms is preferred.
  • Experience with engineering productivity metrics tooling (LinearB or similar) and AI-based code review tooling (Qodo or similar) is preferred.
  • Experience with vector databases and RAG pipelines (e.g. Milvus, Pinecone, pgvector, or similar) in a production setting is preferred.
  • Exposure to model fine-tuning or lightweight training pipelines (LoRA/QLoRA or similar) for domain-specific model adaptation is preferred.
Why SentinelOne?

AI is redefining how the world operates and rewriting the rules of security in real time, and SentinelOne was built for this moment. From day one, we architected an AI-native platform designed to operate at machine speed, not as an add-on to legacy systems but as the foundation itself. If you want to build where innovation and impact move together, this is that place.

We invest in our Sentinels with comprehensive, competitive benefits designed to support you and your family:

Equity & Rewards
  • Restricted Stock Units (RSUs)
  • Employee Stock Purchase Plan (ESPP)
Time Off & Wellbeing
  • Flexible time off
  • Paid company holidays and paid sick time
  • Gender-neutral parental leave
  • Grandparent leave
Insurance & Financial Security
  • Medical, dental, and vision coverage
  • 401(k) retirement plan with company match
  • Life and disability insurance
  • Health and dependent care FSA
  • Voluntary benefits (hospital, accident, critical illness)
  • Employee Assistance Program (EAP)
  • ARAG pre-paid legal
  • Nationwide pet insurance
  • Cancer Care program
  • Global business travel medical insurance
Work Perks & Flexibility
  • Home office allowance
  • Mobile phone reimbursement
Wellness & Lifestyle
  • Wellness coach
  • Wellness/gym reimbursement
  • Fertility coverage
  • Adoption & surrogacy reimbursement

This U.S. role has a base pay range that will vary based on the location of the candidate. For some locations, a different pay range may apply. If so, this range will be provided to you during the recruiting process. You can also reach out to the recruiter with any questions.

Base Salary Range

$156,000—$215,000 USD

SentinelOne is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics.

SentinelOne participates in the E-Verify Program for all U.S. based roles.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior AI Platform Engineer, Infrastructure Services
Senior AI Platform Engineer, Infrastructure Services

SentinelOne • United States

On-site
USD 132,000 - 182,000
RSUs
ESPP
Flexible time off
+1
Senior Staff AI Platform Engineer
Senior Staff AI Platform Engineer

SentinelOne • United States

Hybrid
USD 184,000 - 253,000
Restricted Stock Units (RSUs)
Employee Stock Purchase Plan (ESPP)
Flexible time off
+2
Sr. Manager, AI Software Engineering
Sr. Manager, AI Software Engineering

SentinelOne • United States

On-site
USD 200,000 - 275,000
Restricted Stock Units (RSUs)
Employee Stock Purchase Plan (ESPP)
Flexible time off
+1
Staff Forward Deployed Engineer, AI
Staff Forward Deployed Engineer, AI

SentinelOne • United States

On-site
USD 156,000 - 215,000
RSUs
ESPP
Home office allowance
+1
Principal Software Engineer, AI SIEM
Principal Software Engineer, AI SIEM

SentinelOne • United States

On-site
USD 216,000 - 297,000
Restricted Stock Units (RSUs)
Employee Stock Purchase Plan (ESPP)
Flexible time off
+15
Senior Staff Forward Deployed Engineer, AI
Senior Staff Forward Deployed Engineer, AI

SentinelOne • United States

On-site
USD 184,000 - 253,000
Equity & Rewards
Time Off & Wellbeing
Insurance & Financial Security
+3
Senior Backend Software Engineer - Agent Platform
Senior Backend Software Engineer - Agent Platform

SentinelOne • United States

On-site
USD 132,000 - 182,000
RSUs
ESPP
Flexible time off
+9
Senior Software Engineer - C++ Linux & Cloud Workload Security
Senior Software Engineer - C++ Linux & Cloud Workload Security

SentinelOne • United States

On-site
USD 128,000 - 176,000
Restricted Stock Units (RSUs)
Employee Stock Purchase Plan (ESPP)
Flexible time off
+5
Staff Endpoint Software Engineer, Prompt (Python & OS Internals)
Staff Endpoint Software Engineer, Prompt (Python & OS Internals)

SentinelOne • Northern (KY)

Hybrid
USD 156,000 - 215,000
RSUs
ESPP
Flexible time off
+8
Sr Manager, Solutions Engineering
Sr Manager, Solutions Engineering

SentinelOne • Town of Texas (WI)

On-site
USD 232,000 - 319,000
RSUs
ESPP
Flexible time off
+13