Location: Remote in the United States (ET/CT) | Travel: Occasional travel to NYC or Atlanta
Employees: 50 | Industry: SaaS, Hospitality, Media, Marketing | Reports To: Vice President, Product
Responsibilities
- Tracking how often, how favorably, and how accurately clients appear in AI answers versus a defined competitive set - with rigorous baselines, sampling, and stated uncertainty so real change can be distinguished from noise.
- Maintaining the tracked-prompt panels (across engines, topics, markets) this runs on, and report movement over time as the core output.
- Identifying which sources AI systems cite for a given topic and characterize that mix
- Tracking whether and how published content later surfaces as cited evidence, how long that takes, and how long it lasts - surfacing where a client's evidence base is thin and what's missing
- Testing which content traits (source, framing, format, specificity, recency, structure) make material more likely to get picked up by AI systems.
- Connecting AI visibility to real outcomes — referral traffic, engagement, bookings/conversions and build an honest instrumentation view of what's measurable today, what isn't, and where the recommendation-to-revenue chain breaks
- Evaluating AI answers for accuracy, relevance, and completeness; catalog failure modes by engine over time; maintaining the rubrics/benchmarks/rater-agreement process
- Continuously refining how prompts are constructed - applying the same evaluation discipline to any AI features built internally
- Staying ahead of how AI search systems retrieve/cite sources, validating claims via research (not industry commentary), defining product metrics with product/engineering, ensuring every customer-facing figure is fully re-derivable from documented inputs
Requirements
- 5+ years of experience in analytics, data science, research, and/or search products
- 2+ years of experience with AI search, SEO/GEO/AEO, content performance, or LLM evaluation
- Hands-on fluency with LLMs + agentic tooling – prompt design, evaluation harnesses, agents that collect and process data at volume
- Strong SQL and Python skills – experience with Databricks is preferred (or similar tools)
- Proven experience in experiment design applied to observational data – comparison group construction, difference-in-differences or matched comparison, and honest treatment of statistical power
- Strong analytical judgement with unstructured text: designing rubrics, coding qualitative signals into structured data, and reasoning carefully about ambiguous cases
- Startup experience is required – fast-moving, small teams, high visibility/ownership
- 401K with 4% match, monthly wellness stipend, comprehensive medical and dental insurance plans, long-term disability, and more