Labs
Skip to content

Eval report · 2 suites · 42 cases · Last run 2026-09-22

The first thing these evals caught was a stale index, not a bad answer

Two evals score the AI features here: whether the garden’s chat answers from the right documents, and whether the router reads a prompt into the right task. Both scorers are deterministic — no model judges another model — and both runs are started by hand. The numbers below are dated for that reason.

Suites
2
Cases
42
Last run
2026-09-22
Scorers
deterministic
Model as judge
not used
In CI
no — run by hand

A What a failing eval should look like

The first time the chat eval ran against a real model it did not score. It stopped and said that one document the dataset expects was not in the embedding index, and that a recall of zero for a document which was never embedded says nothing about the model.

That refusal is the design. An empty index scores every grounded question zero and every refusal question one — a plausible-looking 24% that reads like a regression and is really a setup mistake. A number that low would have sent someone to the prompt. Stopping sends them to the index, which is where the problem was.

The rule it follows

A setup failure should read as a setup failure. Anything an eval cannot legitimately score, it should decline to score — loudly, and before the spend.

B Grounded chat: did it use the right sources?

The garden’s chat may only answer from documents it retrieved. So the eval scores two things and neither needs an opinion about prose: source recall — were the documents a good answer must rest on actually among the ones it cited — and refusal, which is structural: the no-sources state, not a sentence that happens to sound uncertain.

13 questions have an expected set of sources; 4 are about things the garden has nothing to say about, and must be refused. Most of the documents involved are private or internal, so they are counted here rather than named. One case is public at both ends, and it is the worked example:

A case, in fulltest/evals/ai-chat.dataset.ts
AskedMust be grounded in
Can a ZIP code be filled in from the browser's location?/proofs/zip-from-geolocation

C The router classifier: where it slips

Ask for something and a small model reads the request into one of six tasks, which is what decides the model that answers. 21 prompts here carry a labeled task; 4 are deliberately vague and should land on the fallback rather than a confident guess. Scoring is an exact match on the task id.

It gets most of them. Two it gets wrong the same way every time — and those are worth more than the score, because each one names something the brief does not say:

Stable missestest/evals/ai-router.dataset.ts
PromptExpectedRead asWhat it tells us
Write three sentences for the About page of a designer's personal site.chat-balancedchat-fastShort output reads as a cheap task. Length is not difficulty — the brief does not say so.
Make this better.chat-balanced (the fallback)codeAn undecidable prompt should fall back, not guess. “This” reads as source to a classifier that has seen a lot of source.

D Every run, dated

Every score a run produced, never the average. A pass that reads every labeled prompt correctly and still guesses at a vague one is not the same result as the reverse, and one blended number hides which happened.

Grounded chattest/evals/ai-chat.eval.ts
DateCommitModelSource recallRefusal
2026-09-21812c0654direct / claude-sonnet-5 over text-embedding-3-large13/134/4
2026-09-22aa7eff79direct / claude-sonnet-5 over text-embedding-3-large13/134/4
Router auto-classifytest/evals/ai-router.eval.ts
DateCommitModelLabeledAmbiguous
2026-09-22aa7eff79claude-haiku-4-5 (Router chat-fast)19/213/4

E What these numbers do not say

  • Nothing about what a member can get an answer to. The eval retrieves as the owner, so every document is visible to it. A run shaped like a member would refuse nearly every question and would be measuring the visibility rules, which are tested exhaustively elsewhere.
  • Nothing about whether an answer reads well. Deterministic scorers can tell you the right documents were used. They cannot tell you the sentences were worth reading. That would take a model judging a model, which is not turned on.
  • Nothing about right now. These runs are started by hand and cost real tokens. No schedule, no CI job. A row is exactly as current as its date.
Last run 2026-09-22Generated from shared/data/eval-runs.tsThe same file writes docs/agents/ai-evals.mdNo model was called to render this page
  • a: Appearance
  • ?: Keyboard shortcuts