Founding AI Engineer

Worky

San Francisco (CA)

On-site

USD 225,000 - 255,000

Full time

14 days+

Get more replies from employers

Send a job-specific resume in minutes.

Job summary

Worky is seeking an experienced software engineer to advance its AI pricing engine. You will build evaluation systems, automate pricing personas, and own LLM infrastructure across models with attention to cost, latency, and quality.

You will ensure data residency boundaries, extend our MCP-based tooling, and land model outputs into structured, auditable forms within a typed ontology. You’ll mentor engineers and engage closely with customers.

Qualifications

  • 8+ years of engineering experience with strong, recent production LLM depth.
  • Shipped LLM-powered product features to production and owned them after launch.
  • Built evals and observability for LLM systems.
  • Strong communication bridging technical and non-technical stakeholders.
  • Product engineer instincts: you own scope and prefer pragmatic options.
  • Experience with AI coding tools in daily work.

Responsibilities

  • Build evaluation for AI pricing recommendations with eval harnesses.
  • Automate persona training and simulate B2B buying committees.
  • Own LLM infrastructure across models with cost, latency, and quality tradeoffs.
  • Maintain data residency boundaries for EU and other regions.
  • Extend MCP server used by LLM agents and customer agents.
  • Land model outputs into structured, auditable forms within a typed ontology.

Skills

LLM depth
Production features
Observability
Communication
Product instincts
AI coding tools
US work authorization

Tools

LangChain
LlamaIndex
Braintrust
OpenRouter

Job description

About the role

Mondrio's AI recommends prices, and expert Pricing Architects stay in the loop on the high-stakes calls. Your mandate is to build the evaluation systems and feedback loops that let the AI earn more of that trust.

This is applied AI on a problem where quality is measurable in customer revenue.

What you'll do
  • Build evaluation for AI pricing recommendations: eval harnesses and benchmarks that use tracked pricing outcomes as ground truth. Expert review is manual today, and you make it systematic.
  • Take AI personas further. They simulate B2B buying committees and behavioral effects such as new versus existing customers, grounded in usage data and call transcripts. Automate the parts of persona training that are still manual.
  • Own LLM infrastructure: routing across current (Anthropic and Google) and future models, with explicit cost, latency, and quality tradeoffs.
  • Maintain infra and data residency boundaries (e.g. model calls for EU customers must remain within the EU) as we add providers and scale up operations.
  • Extend the MCP server that LLM agents, including our customers' own agents, use to drive the platform. A feature is done when an agent can drive it through MCP, not when the React component renders.
  • Work within our typed ontology of pricing entities (Pydantic models for SKU, Proposition, Persona, Quote) so model outputs land in structured, auditable form.
Your first 90 days

First 30 Days: Foundation & Guardrails

  • Model routing across Anthropic and Google has explicit cost and latency budgets, and the fail-closed EU residency guarantee covers every model call.
  • Pinpoint systemic latency, data drift, or cold-start issues in the continuous pricing loop.
  • Baseline current prompt and model outputs against our typed ontology to prepare for release-gating evals.
  • Conduct code reviews and lead a technical session on advanced AI/ML patterns for the team.

By Day 60: Trust & Automation

  • An eval harness runs on every model or prompt change, and the team trusts its benchmarks enough to gate a release on them.
  • Manual steps in persona training now run as an automated pipeline built on the same usage data and call transcripts.

By Day 90: Closed-Loop Impact

  • Tracked pricing outcomes feed back into recommendation quality, so evals measure revenue impact, not proxy scores.
  • Participate actively in interview loops to scale the engineering team and mentor mid-level engineers.
What we're looking for
  • 8+ years of engineering experience with strong and recent production LLM depth.
  • You have shipped LLM-powered product features to production and owned them after launch.
  • You have built evals and observability for LLM systems yourself. Running someone else's dashboard does not count.
  • Strong communication skills to bridge the technical gap around non-deterministic engineering to less savvy clients and partners
  • Product engineer instincts: you pick your own scope and choose the pragmatic option over the interesting one. This is not a research-lab role.
  • You can show how AI coding tools fit into your work today. We weigh that over where you studied or previous role.
  • Work authorization: You must be authorized to work in the US. We're unable to sponsor visas at this time.
Nice to have:
  • Experience with MCP or building tools for LLM agents.
  • Experience with platforms such as LangChain, LlamaIndex, Braintrust, OpenRouter
  • Work in a domain where correctness is audited, such as pricing, billing, or payments.
  • Familiarity with data residency or compliance constraints. SOC2 and GDPR shape what you build against.
Our stack

Client: Typescript, React, Vercel Chat SDK

Server: Python, FastAPI

Data: Mongo, Atlas

Infra: GCP, Pulumi, Cloudflare Pages

AI: FastMCP, Langfuse, Claude Code, Cursor, Vercel Eve

Security: SOC2 Type 1/Type 2, GDPR compliant, EU and US data residency

How we work

We are under ten people, and everyone ships and talks to customers. Engineers are product engineers: you own outcomes, scope your own work, and demo every week.

We appreciate fast feedback loops, real collaboration, and a team that genuinely enjoys working together. If that's how you do your best work, you'll fit right in.

There is no separate PM layer. Our SDLC is AI-native as we move to a robust software factory. Claude Code and Cursor are standard kit, and LLM agents write and review pull requests behind automated review gates. We keep pull requests small, around 500 lines, because the research on small batches holds up and we act on it.

The architecture rules are short and enforced: the API is the product so it evolves additively, business logic stays server-side, and writes are audited by default for compliance. Your evals are how we stay honest about whether the AI actually works.

What we offer
  • Compensation: $225,000 - $255,000
    Plus meaningful equity through our employee stock option plan (ESOP). We aim to be highly competitive on cash and generous on equity. Joining at this stage means a real stake in what we build together.
  • Benefits: A market-conform package, including health coverage, pension/retirement provision, generous paid time off, paid parental leave.
  • Direct customer contact from early on: you hear how your recommendations land.
  • Real influence on the architecture while the big decisions are still open. The AI layer is young, and you set its shape.
  • A team where your impact is visible.
  • An opportunity to be part of a fast-growing company.
  • Daily work with, and real influence over, a cutting-edge AI pricing engine.
  • Continuous exposure to the newest AI tools and techniques.
Get your free, confidential resume review.
or drag and drop your file here.
Similar jobs

Similar jobs worth comparing

Founding AI Engineer
Founding AI Engineer

Mondrio • San Francisco (CA)

On-site
USD 225,000 - 255,000
Equity/ESOP
Competitive cash compensation
Founding Product Engineer
Founding Product Engineer

Worky • San Francisco (CA)

On-site
USD 200,000 - 250,000
Health insurance
Pension/retirement plan
Generous paid time off
+1
Founding Design Engineer
Founding Design Engineer

Worky • San Francisco (CA)

On-site
USD 120,000 - 180,000
Founding Design Engineer
Founding Design Engineer

Mondrio • San Francisco (CA)

On-site
USD 180,000 - 220,000
Equity
Health benefits
Pension/retirement
+2
AI Engineer
AI Engineer

Fluency • San Francisco (CA)

On-site
USD 180,000 - 250,000
US$1,000 per month food and commuting allowance
Laptop of choice
ESOP available
Founding Product Engineer
Founding Product Engineer

Mondrio • San Francisco (CA)

On-site
USD 180,000 - 260,000
Senior AI Engineer – Agents
Senior AI Engineer – Agents

Doist • San Francisco (CA)

On-site
USD 150,000 - 210,000
Product Engineer - Frontend Leaning
Product Engineer - Frontend Leaning

Alma • Palo Alto (CA)

On-site
USD 180,000 - 230,000
Health coverage
PTO + Holidays
Wellness allowance
+1
Staff AI Engineer
Staff AI Engineer

Empathy Talent • San Francisco (CA)

On-site
USD 279,000 - 341,000
Wellness benefits starting day 1
Flexible time off
Technology reimbursements
+3
Founding AI Engineer
Founding AI Engineer

VoiceOps • New York (NY)

On-site
USD 120,000 - 160,000
100% employer-paid health insurance premiums
Flexible PTO
Seed-stage equity grant
+2