Backend Engineer, AI Systems & Evals

Renaissance Geek

San Francisco (CA)

Hybrid

USD 180,000 - 240,000

Full time

11 days ago
Application generator

An application made for this job — a tailored resume and cover letter that speak straight to the posting.

Get past ATS filters

Benefits offered by this job

Health coverage
Dental coverage
Vision coverage

Job summary

Renaissance Geek seeks a senior backend engineer to own the review pipeline and the production backend. You will manage durable pipelines, integrations, data observability, and the evaluations that determine review quality.

You will optimize asynchronous workflows, improve prompts and model routing, and ensure scalability as more repositories come online. Strong experience with TypeScript, APIs, and distributed systems is expected.

Qualifications

  • Senior-level judgment with a track record of shipping outcomes.
  • Experience using AI tools in production and evaluating their outputs.
  • Proven ability to own backend systems end-to-end, including debugging and recovery.

Responsibilities

  • Own backend services and the review pipeline from events to feedback.
  • Make asynchronous workflows reliable with retries, idempotency, and recovery.
  • Build/maintain evaluation datasets, regression checks, and experiments to measure quality.

Skills

Senior judgment
AI tool usage
Backend systems
TypeScript
APIs
Relational data
Distributed systems

Tools

TypeScript

Job description

Impeccable PRs brings design review into the pull request. It combines deterministic checks, visual evidence, and model judgment to help teams catch interface problems before they ship.

You’ll own the systems that make that work in production: durable pipelines, integrations, data, observability, and the evaluations that tell us whether a review actually helps.

The hard part is making good judgments repeatable. A fast review that misses the problem is useless. A clever critique full of false positives loses trust. You’ll make the tradeoffs, measure the results, and keep improving the system.

What you’ll own
  • Own backend services and the review pipeline from GitHub events through analysis, critique, and posted feedback.
  • Make asynchronous workflows reliable: retries, idempotency, failure recovery, and useful operational visibility.
  • Build and maintain evaluation datasets, regression checks, and experiments that measure critique quality and false positives.
  • Improve prompts, model routing, and the boundary between deterministic checks and model judgment.
  • Own performance, inference cost, tenant isolation, and scaling as more repositories and customers come online.
  • Build the full-stack tools you need to inspect results and make the next decision with evidence.
What we’re looking for
  • You bring senior-level judgment and a track record of owning work from an ambiguous problem through a shipped result. You set priorities, make sound tradeoffs, and know when to involve others.
  • You have deep expertise in something and curiosity that extends well beyond it.
  • You see what needs doing, exercise judgment, and follow through without waiting for a perfectly defined brief.
  • You already use AI tools in your work, question their output, and stay responsible for the result.
  • You’ve built and operated production backend systems, including the less glamorous work of debugging and recovery.
  • You can set technical priorities, make architecture and quality tradeoffs, and own their consequences in production.
  • You’re comfortable with TypeScript, APIs, relational data, and asynchronous or distributed systems.
  • You’ve evaluated an AI system beyond trying a few prompts, or can demonstrate a rigorous approach to measuring quality.
  • You can connect infrastructure decisions to product quality and explain where the evidence is still uncertain.
Location and travel

San Francisco preferred US remote by exception. Approximately once per quarter.

What we offer
  • Ownership of a meaningful part of the product or community, with direct access to the people building it.
  • Comprehensive health, dental, and vision coverage.
  • Flexible time off. Put care into the work, and take the time you need to rest.
Get your free, confidential resume review.

or drag and drop your file here.

Similar jobs

Similar jobs worth comparing

Frontend Engineer
Frontend Engineer

Renaissance Geek • San Francisco (CA)

On-site
USD 140,000 - 220,000
Health, dental, and vision coverage
Flexible time off
Direct access to product team
Forward Deployed Engineer
Forward Deployed Engineer

Renaissance Geek • San Francisco (CA)

Remote
USD 180,000 - 260,000
Health coverage
Dental coverage
Vision coverage
+1
Backend / Systems Engineer
Backend / Systems Engineer

QofAI • San Francisco (CA), Northern (KY)

On-site
USD 140,000 - 210,000
Backend Software Engineer (Evals)
Backend Software Engineer (Evals)

OpenAI • Los Angeles (CA)

On-site
USD 230,000 - 385,000
Head of Evals, AI Red Teaming
Head of Evals, AI Red Teaming

Trajectory Labs, PBC • Berkeley (CA), Northern (KY)

On-site
USD 200,000 - 400,000
Health coverage stipend
401(k)
Visa sponsorship
+1
Design Engineer
Design Engineer

AI Chopping Block • San Francisco (CA), Northern (KY)

On-site
USD 120,000 - 180,000
Research Engineer – Evals
Research Engineer – Evals

Firecrawl • San Francisco (CA)

On-site
USD 160,000 - 240,000
Unlimited PTO
12 weeks fully paid parental leave
Wellness stipend
+5
Senior AI Engineer
Senior AI Engineer

People In AI • San Francisco (CA)

On-site
USD 180,000 - 240,000
Software Engineer, Backend
Software Engineer, Backend

krea.ai • San Francisco (CA)

On-site
USD 140,000 - 210,000
Team
Impact
Competitive compensation
+6
Product Engineer
Product Engineer

Harper • San Francisco (CA)

On-site
USD 140,000 - 280,000
Uber commuter benefits
Breakfast, lunch, dinner provided
Snacks, drinks and coffee daily
+2