Senior Platform AI Engineer

Cacheflow

San Francisco (CA)

Hybrid

USD 192,000 - 259,800

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Benefits offered by this job

Health & Wellness benefits
Generous paid time off
Stock equity options
Professional development stipends

Job summary

Cacheflow is seeking a skilled individual to join their AI Platform team in San Francisco. You'll take charge of the development of infrastructure that powers AI features within the compliance platform.

This hybrid role involves collaborating closely with cross-functional teams to enhance agent orchestration and manage the production AI stack. Ideal candidates will have extensive experience in software engineering and AI infrastructure support.

The position offers a competitive base salary, comprehensive health benefits, and stock options.

Qualifications

  • 7+ years of software engineering experience, with 2+ years building AI/ML infrastructure.
  • Experience with LLM APIs, vector databases, and AI orchestration platforms.
  • Ability to debug prompt templates and design orchestration frameworks.

Responsibilities

  • Design and build MCP servers and API architectures for AI agents.
  • Manage agent orchestration and workflow infrastructure.
  • Operate and evolve production AI stacks and implement RAG systems.

Skills

Python
TypeScript/Node.js
CI/CD pipeline design
API design
Cloud infrastructure (AWS)

Tools

Terraform
Docker
LLM APIs
Vector databases

Job description

Location

Hybrid - San Francisco

Employment Type

Full time

Location Type

Hybrid

Department

Engineering

Job Summary

Drata's AI Platform team builds the production infrastructure that powers AI features across our compliance platform — from MCP servers that make Drata's data available to AI agents, to LLM workflow orchestration that automates SOC 2, TPRM, and policy analysis. You'll own the systems that sit between our AI models and our customers: tool definitions that agents actually understand, deployment pipelines that handle model upgrades without breaking output quality, and orchestration layers that manage multi-step agent workflows with persistent state.

This is not a traditional infrastructure role. You'll debug prompt templates alongside Terraform modules. You'll design API schemas optimized for LLM token budgets, not just HTTP throughput. When a model upgrade changes behavior across 15 workflows, you'll assess quality impact — not just confirm the containers are healthy.

You'll work closely with our agent developers, product engineers, and an embedded SRE partner, sitting at the intersection of AI development and production reliability.

Our north star is simple: minimize the time it takes to launch a new agent in production. You're someone who asks are we solving the right problem? before writing the first line of code, who builds systems that make five other engineers faster, not just yourself, and who's equally proud of what they chose not to build.

What you'll do
MCP Server Development & AI-Optimized API Design
  • Design and build MCP (Model Context Protocol) servers that expose Drata's platform to AI agents. This means making architectural decisions about tool granularity, naming conventions for agent disambiguation, response compression for LLM context windows, and workspace isolation for multi-tenant access. You'll own the protocol layer that determines whether agents can reliably find and use the right tools — writing semantic parameter descriptions, contextual hints, and tool schemas that optimize for model comprehension, not just developer ergonomics.
Agent Orchestration & Workflow Infrastructure
  • Build and operate the infrastructure for deploying multi-step agent workflows — state management across complex reasoning chains, tool routing and execution runtimes, and long-running agentic processes that persist over time. Own the orchestration layer that coordinates agent planning, tool calls, and human-in-the-loop patterns. Design systems that handle agent failure modes gracefully: retries on ambiguous tool outputs, fallback strategies when models produce unexpected results, and observability into multi-step execution traces.
LLM Operations & Model Lifecycle Management
  • Own the operational side of our LLM workflows: model upgrades across production pipelines (assessing behavior changes, not just version bumps), prompt versioning and A/B testing, AI workflow deployment with custom container compatibility, and output quality monitoring.
  • Manage token capacity planning — understanding model costs, context limits, batching strategies, and rate governance across workflows. When an AI workflow fails, you'll investigate whether it's a prompt template issue, a model behavior change, or an infrastructure problem. Making that distinction requires understanding both systems.
Production AI Infrastructure & RAG Systems
  • Operate and evolve our production AI stack: vector storage and indexing (designing chunking strategies and metadata schemas for retrieval quality), document parsing pipelines, multi-region deployment, and cost optimization across LLM providers. You'll make RAG architecture decisions — embedding strategies, retrieval filtering, data model coordination — where the engineering challenge is search quality, not just system uptime. Implement caching layers and token-aware request routing to manage spend as AI workloads scale.
Platform Enablement & Developer Experience
  • Build CI/CD patterns specific to AI workflows (reproducible deployments, SDK version compatibility, workflow rollback semantics). Own AI-specific observability — token usage dashboards, response quality metrics, agent execution traces, and cost-per-workflow tracking alongside traditional infrastructure monitoring. Enable product engineering teams to ship AI features faster by providing reliable, well-documented platform primitives.
What you'll bring

7+ years of software engineering experience, with 2+ years building or operating AI/ML infrastructure in production. You're strong in Python (our AI services are built in Python), with TypeScript/Node.js a nice-to-have. You've worked with LLM APIs, vector databases, or AI orchestration platforms and understand the difference between "the service is up" and "the model output is good." You're comfortable across the stack: writing Terraform one day, debugging a prompt template the next, and designing an agent orchestration framework the day after.

Specifically, you bring experience in several of these areas: cloud infrastructure (AWS preferred — ECS, S3, Bedrock), container orchestration, infrastructure-as-code, CI/CD pipeline design, API design, workflow orchestration engines, and distributed systems. You've worked with at least some AI-specific tooling: LLM APIs (Claude, OpenAI, etc), model serving frameworks (vLLM, SageMaker etc), vector databases, embedding pipelines, prompt management platforms, or agent frameworks.

You communicate clearly about technical tradeoffs, especially when explaining AI-specific infrastructure decisions to stakeholders who think in terms of traditional reliability engineering. You own what you see broken, not just what's assigned to you, and you can spot when an architecture decision will fail at scale and say so early, clearly, and with an alternative.

How we support you
  • Shared Success: We provide stock equity to ensure that as the company grows, you share directly in that success. Equity gives every employee a sense of ownership and the opportunity to celebrate our wins together—because your contributions don’t just support our progress; they help drive our collective success.
  • Health & Wellness: Up to 100% employer-paid premiums for medical, dental, and vision coverage for employees and their dependents, along with comprehensive wellness benefits and healthcare concierge services designed to support your needs beyond traditional insurance.
  • Financial Well-being: A comprehensive suite of financial benefits, including a 401(k) plan, company-paid life and disability insurance, tax-advantaged spending accounts, and a range of discounted voluntary offerings to help you customize and strengthen your overall financial position.
  • Family Support: We want to support you in life's most important moments, so we offer a paid Parental Leave policy, after six months of employment. Employees also receive access to Kindbody fertility and family-building benefits and dedicated leave specialists who help guide you through the entire process.
  • Growth & Development: Generous annual stipends for both professional and personal development, empowering you to invest in your continued growth. You’ll also have access to a wide range of internal learning opportunities, ensuring you can build new skills, deepen your expertise, and advance your career with confidence.
  • Time Off & Flexibility: We believe that to do your best work, you should get the time you need for rest, rejuvenation and recovery. Drata offers a flexible vacation policy, paid holidays, and other perks to recharge.

This role will receive a competitive base salary, benefits, and stock, typically in the form of Restricted Stock Units (RSUs). The applicable salary range for this role is: $192,000 - $259,800.

A variety of factors are considered when determining someone’s leveling and compensation–including a candidate’s professional background and experience. These ranges may be modified in the future and final offer amounts may vary from the amounts listed above.

Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Senior Platform AI Engineer
Senior Platform AI Engineer

Drata • San Francisco (CA)

On-site
USD 192,000 - 260,000
Stock equity
Up to 100% employer-paid health insurance
401(k) plan
+2
Senior Platform AI Engineer
Senior Platform AI Engineer

Careers at Drata • San Francisco (CA)

Hybrid
USD 192,000 - 260,000
100% employer-paid premiums for medical, dental, and vision coverage
Flexible vacation policy
401(k) plan with company-paid life and disability insurance
+1
Senior Platform Engineer 2, AI Tooling
Senior Platform Engineer 2, AI Tooling

Cacheflow • San Francisco (CA)

Hybrid
USD 174,000 - 237,000
Stock equity
Full medical, dental, and vision coverage
401(k) plan
Senior Platform Engineer 2, AI Tooling
Senior Platform Engineer 2, AI Tooling

Drata • San Francisco (CA)

On-site
USD 174,000 - 237,000
Stock equity
100% employer-paid health premiums
401(k) plan
+2
Platform Engineer, AI Tooling
Platform Engineer, AI Tooling

Drata • San Francisco (CA)

Hybrid
USD 131,000 - 179,000
Stock equity
100% employer-paid medical, dental, and vision coverage
401(k) plan
+2
Senior AI Product Engineer 2, Audit
Senior AI Product Engineer 2, Audit

Drata • San Francisco (CA)

Hybrid
USD 192,000 - 260,000
Stock equity
100% employer-paid health premiums
Flexible vacation policy
+1
Senior AI Engineer
Senior AI Engineer

Cacheflow • United States

Hybrid
USD 192,000 - 238,000
Stock equity
100% employer-paid health premiums
401(k) plan
+2
Senior Platform Engineer 2, AI Tooling
Senior Platform Engineer 2, AI Tooling

Careers at Drata • San Francisco (CA)

Hybrid
USD 174,000 - 237,000
Stock equity
100% employer-paid health coverage
401(k) plan
+2
AI Product Engineer, Evidence
AI Product Engineer, Evidence

Cacheflow • San Francisco (CA)

Hybrid
USD 145,000 - 197,000
Stock equity
100% employer-paid health, dental, and vision coverage
401(k) plan
+2
Platform Engineer, AI Tooling
Platform Engineer, AI Tooling

Careers at Drata • San Francisco (CA)

Hybrid
USD 131,000 - 179,000
Up to 100% employer-paid health premiums
401(k) plan
Paid parental leave
+1