Staff AI Platform Engineer: Agent & Retrieval Infrastructure

Bedrock Ocean Exploration

New York (NY)

On-site

USD 160,000 - 220,000

Full time

12 days ago
Application generator

A complete application in a minute — tailored resume and cover letter, ready to send.

Get past ATS filters

Job summary

Bedrock Ocean is seeking a Staff Platform Engineer to lead our AI architecture, building the infrastructure for AI agents, including orchestration, data and retrieval pipelines, and the security model required to work with production data.

This role blends software engineering, data engineering, and infrastructure operations, overseeing the Amazon Bedrock implementation from data chunking to IAM access controls, expanding pipelines, and ensuring secure, scalable operations across stages.

Qualifications

  • 8+ years in software and infrastructure engineering, with senior backend experience and staff-level ownership.
  • Production backend experience in Python or TypeScript (Go allowed).
  • Hands-on experience standing up Amazon Bedrock in production: agents, knowledge bases, guardrails, and model access.
  • Containerized services (ECS/EKS/Lambda) with CI/CD you own.
  • Practical RAG and vector search experience with production vector DBs.
  • Strong AWS IAM, VPC networking, PrivateLink, and IaC (Terraform/CDK/CloudFormation).
  • Experience designing developer-facing APIs, SDKs, or platform services.
  • Experience building multi-tenant services with isolation guarantees.

Responsibilities

  • Architect Agent Orchestration: design Bedrock integration, agent configurations, APIs, model access, throughput, and deployment.
  • Manage Retrieval Data Plane: end-to-end retrieval pipeline, ingestion to embedding and storage in OpenSearch Serverless; optimize index design, cost, capacity.
  • Extend Data Pipelines: adapt ingestion for internal knowledge, ocean data, and customer platforms; handle geospatial and large-binary data challenges.
  • Secure AI Infrastructure: implement Bedrock Guardrails, VPC and PrivateLink boundaries, least-privilege IAM, audit trails.
  • Define Agent Governance: enforce approval boundaries for autonomous actions and monitoring.
  • Establish LLMOps & Observability: monitoring for tracing, tool calls, and retrieval performance (CloudWatch, Langfuse/Phoenix).
  • Build Evaluation Frameworks: infrastructure to run automated evaluations and manage release gates for model accuracy.
  • Enable Engineering Productivity: provide abstractions, SDKs, and self-service environments for engineers.

Skills

Python/TypeScript
Go
AWS Bedrock
LLM/AI platforms
CI/CD
OpenSearch Serverless
Terraform/CDK/CloudFormation
Security IAM/VPC
Multi-tenant platforms

Tools

OpenSearch Serverless
Langfuse
Phoenix
Terraform
CDK
CloudFormation
PrivateLink

Job description

About Bedrock Ocean

Bedrock Ocean builds and operates autonomous underwater vehicles (AUVs) that collect georeferenced ocean-floor data at commercial scale. We deliver bathymetric and imagery data products to customers through our own platform, and we're scaling toward continuous, around-the-clock data collection campaigns spanning months at a time.

We are building AI agents on Amazon Bedrock to support our ocean data, internal operations, and customer platform. This role owns that architecture.

(One note on names. Amazon Bedrock is the AWS service. Bedrock Ocean is us. They are unrelated, and we are aware it is confusing.)

The Role

We are looking for a Staff Platform Engineer to lead our AI architecture. This role goes beyond building agents on existing platforms; you will create the infrastructure itself, including the orchestration layer, the data and retrieval pipeline, and the security model required to work with production data. You will also build the tools and abstractions that allow our engineering team to implement AI features independently.

This position combines software engineering, data engineering, and infrastructure operations. You will manage the full lifecycle of our Amazon Bedrock implementation, from initial data chunking to IAM access controls. While some of our data pipelines are already in place, they will require significant expansion, and others will need to be built from scratch.

Security is core to this role, not an afterthought. Because agents with tool access represent a new kind of system actor, you will define their operational boundaries, including what they can access, the actions they can perform autonomously, and the monitoring required to detect issues.

Our roadmap prioritizes internal engineering and operational systems first to ensure a fast feedback loop, followed by our ocean and survey data products. Customer-facing retrieval is the final, high-stakes phase. You will play a key role in defining this sequence. While you will not be responsible for the core data transport design (store-and-forward or hub-and-spoke), you will work closely with that team to ensure it meets our retrieval and data freshness requirements.

What You'll Do
  • Architect Agent Orchestration: Design the Amazon Bedrock integration, including agent and action group configuration, backend APIs, model access, throughput, and cross-environment deployment.

  • Manage Retrieval Data Plane: Own the end-to-end retrieval pipeline from ingestion and chunking to embedding and storage in Amazon OpenSearch Serverless. Focus on optimizing for index design, cost, and capacity.

  • Extend Data Pipelines: Adapt ingestion pipelines for internal knowledge, ocean data, and customer platforms, addressing challenges specific to geospatial and large-binary datasets.

  • Secure AI Infrastructure: Implement robust security including Bedrock Guardrails, VPC and PrivateLink network boundaries, least-privilege IAM, and audit trails to ensure data isolation.

  • Define Agent Governance: Build the mechanisms to enforce approval boundaries for autonomous actions, ensuring agents are safe and monitored.

  • Establish LLMOps & Observability: Implement comprehensive monitoring for tracing, tool calls, and retrieval performance, using CloudWatch and LLM-specific tools like Langfuse or Phoenix.

  • Build Evaluation Frameworks: Create the infrastructure to run automated evaluations, track results, and manage release gates for model accuracy.

  • Enable Engineering Productivity: Provide the team with abstraction layers, SDKs, and self-service environments that allow engineers to ship AI features independently.

  • Operational Excellence: Manage the environment as code across all stages, ensuring deployment safety and participating in incident reviews.

What We're Looking For
  • 8+ years in software and infrastructure engineering, including deep production backend experience (Python or TypeScript preferred, Go fine) and staff-level ownership of technical direction.

  • Hands-on experience standing up Amazon Bedrock in production: agents, knowledge bases, guardrails, model access, and the throughput and quota decisions that come with them.

  • Containerized service deployment on ECS, EKS, or Lambda, with CI/CD you have owned rather than inherited. The models are managed, but the backend APIs, tool endpoints, and ingestion jobs still run somewhere real.

  • Practical RAG and vector search experience: embeddings, chunking strategies, semantic search quality, and operating a managed vector database (OpenSearch Serverless, Pinecone, pgvector, or similar) at production scale and cost.

  • Real data engineering: you have built or substantially extended ingestion pipelines over messy, heterogeneous, unstructured sources, and you think about freshness and correctness as SLAs rather than afterthoughts.

  • Strong AWS ecosystem expertise: IAM roles and least privilege for machine identities, VPC networking and PrivateLink, Lambda, S3, KMS, CloudWatch, and provisioning safely through infrastructure as code (Terraform, CDK, or CloudFormation).

  • Production LLM exposure: you have moved LLM features or autonomous agents past the prototype stage into environments other people depend on.

  • A working point of view on securing agentic systems: scoping tool permissions, prompt injection and exfiltration risk, sensitive data handling in retrieval, and where a human belongs in the loop.

  • Experience designing developer-facing APIs, SDKs, or platform services with an API-first mindset. The interface is the product for the engineers who consume it.

  • Experience building and operating multi-tenant services, with isolation guarantees that hold when the data belongs to customers rather than to us.

  • Platform instinct: you build the abstraction other engineers stand on, and you measure yourself by what they ship rather than by what you ship directly.

  • Demonstrated technical leadership and system design judgment at staff level: you have driven an architectural direction across pods or teams you do not manage, and made it stick through influence rather than authority.

  • A pragmatic builder's bias. You reach for boring, fully managed infrastructure before complex self-hosted alternatives, and you can tell the difference between the two in an architecture review.

  • Comfort wearing several hats on a small team, and the discipline to write things down so the system runs without you.

Nice to Have
  • Experience deploying LLM evaluations to measure accuracy over time and treating eval results as a release gate.

  • Involvement in AI red-teaming or the AI security community.

  • Experience with GraphRAG or knowledge graphs.

  • Experience running retrieval over geospatial, scientific, or large-binary datasets.

  • Experience moving data across intermittent or unreliable links: store-and-forward, hub-and-spoke topologies, offload from disconnected or edge systems, and reconciliation once a link comes back.

  • Compliance experience such as SOC 2, or handling government or defense customer data.

  • Background supporting data platforms, autonomous systems, or field operations.

Not a Fit If
  • Your AI work has been prototypes and notebooks rather than systems other people depend on in production.

  • You want to be a model researcher, a prompt engineer, or to spend your time fine-tuning models. This role owns the platform underneath agents and partners closely with the people building them.

  • You treat security as a gate at the end of a project rather than something designed from the start.

  • You would rather self-host and build from scratch than adopt a managed service that already works. We are optimizing for a small team shipping, not for architectural purity.

  • You want a mature platform team and a narrow, well-bounded scope. This is an early build with a lot of surface area and few existing answers.

Why This Role Matters

At Bedrock Ocean, our mission is to make the ocean transparent. We are building more than just a survey service; we are creating a source of deep ocean intelligence that grows with every mission.

To realize this, we need to make our data accessible and actionable. This role is about building the infrastructure that allows our teams and eventually our customers to directly query and learn from our findings. Because this data is strategically sensitive, security is woven into the foundation of your work, not added on later.

Ultimately, an AI agent with access to our production data is a powerful tool, but it requires careful design to be a success. You will build a secure, reliable platform that ensures these systems empower our mission safely and effectively.

The base compensation for this role is expected to be $160,000- $220,000 annually plus equity.

Bedrock Ocean is an equal opportunity employer.

Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Staff AI Platform Engineer: Agent & Retrieval Infrastructure
Staff AI Platform Engineer: Agent & Retrieval Infrastructure

Worky • New York (NY)

On-site
USD 180,000 - 220,000
Staff AI Platform Engineer – Agents, Retrieval & Security
Staff AI Platform Engineer – Agents, Retrieval & Security

Bedrock Ocean Exploration • New York (NY)

On-site
USD 160,000 - 220,000
Staff Robotics Engineer
Staff Robotics Engineer

Bedrock Ocean Exploration • New York (NY)

On-site
USD 190,000 - 235,000
Staff Robotics Engineer
Staff Robotics Engineer

Bedrock • Richmond (CA), Northern (KY)

Hybrid
USD 190,000 - 235,000
Staff Robotics Engineer
Staff Robotics Engineer

bedrockocean • Richmond (CA)

On-site
USD 190,000 - 235,000
Staff AI Platform Engineer: Agent Orchestration & Retrieval
Staff AI Platform Engineer: Agent Orchestration & Retrieval

Quiet Capital • United States

Remote
USD 160,000 - 220,000
Applied AI Engineer
Applied AI Engineer

Vation Ventures • Denver (CO)

On-site
USD 150,000 - 210,000
Staff Platform Engineer - AI Orchestration & Retrieval
Staff Platform Engineer - AI Orchestration & Retrieval

Worky • New York (NY)

On-site
USD 180,000 - 220,000
Staff AI Architect
Staff AI Architect

DigitalOcean • Boston (MA)

Hybrid
USD 191,000 - 239,000
Sr. Software Development Engineer, Bedrock AgentCore Knowledge Bases
Sr. Software Development Engineer, Bedrock AgentCore Knowledge Bases

Amazon Web Services (AWS) • New York (NY)

On-site
USD 185,000 - 250,000
Health insurance
401(k) matching
Paid time off
+2