Role Overview
We are seeking aVice President – AI SafetyPlatformsto build and lead our enterprise AI safety engineering initiatives. As generative AI in financial services evolves from simple prompt-response workflows to autonomous agentic systems that execute multi-step plans, call APIs, and interact directly with internal systems, establishingrobust safety mechanisms and standardized evaluation protocols is essential.
Key Responsibilities
1. Unified Agentic Evaluation Framework
- Company-Wide Architecture:Design, build, and deploy a single, company-wide agentic evaluation frameworkthat standardizes how teams across all business lines benchmark, test, and measure AI agent performance prior to production deployment.
- Trajectory & Multi-Step Reasoning Assessment:Implement evaluation methodologies that score autonomous planning quality, tool-calling precision, multi-turn state retention, trajectory efficiency, and error-recovery behaviors.
- Continuous Monitoring & Production Drift:Integrate automated evaluation pipelines into runtime environments to continuously audit agent execution traces, detecting reasoning drift, tool failure modes, and unexpected trajectory shifts in production.
- Domain-Specific Benchmarking:Establishstandardized test suites and synthetic evaluation benchmarks tailored to complex financial workflows, such as automated research, risk assessment, and operational task execution.
2. LLM Guardrails Infrastructure & Real-Time Controls
- Low-Latency Guardrail Engine:Architect and scale enterprise guardrail microservices that inspect prompt inputs, retrieved context, and model outputs in real time to prevent data leakage, policy violations, and unvalidated execution.
- Tool-Use & Action Control:Implement runtime policy gateways that inspect and authorize tool calls before execution, ensuring agentsoperatewithin authorized data boundaries and action scopes.
- Human-in-the-Loop (HITL) Triggers:Build configurable escalation workflows and approval gates that automatically pause execution for high-risk operations (e.g., money movement, client record modifications, or external communications) until human authorization is granted.
3. Core AI Platform Enhancements & Governance Integration
- Drive Platform Enhancements:Partner directly with the core AI Platform team to drive the implementation of safety APIs, telemetry hooks, developer SDKs, andMLOps/LLMOpspipeline integrations.
- Auditability & Execution Telemetry:Define and enforce technical standards for immutable audit logging, execution tracing (e.g.,OpenTelemetrystandards), and principal identity propagation across all agentic workflows.
- Regulatory & Model Risk Alignment:Translate model risk management standards (e.g., SR 11-7 / SR 26-2 guidance, FINRA supervision requirements) into automated engineering safeguards and policy checks.
4. Engineering Leadership & Strategic Oversight
- Team Building & Mentorship:Hire, develop, and mentor high-performing engineering teams specializing in applied machine learning, AI safety, and enterprise platform engineering.
- Strategic Roadmap:Own the technical roadmap for enterprise AI safety infrastructure, setting clear milestones for evaluation framework adoption, runtime latency optimization, and governance automation.
- Stakeholder Collaboration:Articulate technical risk profiles,evaluationmetrics, and safety architecture to risk committees, model validation teams, and executive leadership.
Key Qualifications
Basic Qualifications
- Role Level:Vice President experience (or equivalent senior engineering leadership) in financial services or large-scale enterprise software environments.
- Education:Bachelor’s orMaster’s degree in Computer Science, Artificial Intelligence, Systems Engineering, or a related quantitative field.
- Engineering Leadership:4+ years leading applied ML or software engineering teams in building platform infrastructure or microservices.
- Software Engineering Depth:8+ years of hands-on software development experience (Python, Go, Java, or C++) building microservices, high-throughput APIs, or enterprise platform services.
- AI & Agentic Expertise:Technical fluency with Large Language Models (LLMs), RAG systems, function calling / tool integration, and agentic execution paradigms (e.g.,LangChain,AutoGen,CrewAI, MCP server architectures).
Preferred Experience & Technical Skills
- Agentic Evaluation:Direct experience building agent evaluation frameworks and metrics (e.g., LLM-as-a-Judge, G-Eval, trajectory trace evaluation, task completion scoring).
- Guardrail Frameworks:Hands-on experience integrating low-latency guardrail tools and runtime filters (e.g.,NeMoGuardrails, Guardrails AI, Llama Guard).
- AI Observability & Tracing:Experience with LLM and agent tracing tools (e.g.,LangSmith,OpenTelemetry, Phoenix,MLflow) and structured audit logging infrastructure.
- Platform Engineering Alignment:Proven ability to partner across teams and drive key governance capabilities into core shared platforms.
Salary Range
The expected base salary for this New York, NY, United States-based position is $130000-$250000. In addition, you may be eligible for a discretionary bonus if you are an active employee as of fiscal year-end.
Benefits
Goldman Sachs is committed to providing our people with valuable and competitive benefits and wellness offerings, as it is a core part of providing a strong overall employee experience. A summary of these offerings, which are generally available to active, non-temporary, full-time and part-time US employees who work at least 20 hours per week, can be found here.