rememberoReading recall labdeterministic pipeline · no model executes on this page

The reader contract · paired-run replay

Watch the reading pipeline assemble — then see what the notes bought.

Retrieval, re-ranking, context tiering, and the computed-notes block below are real code running in your browser over a fixed fictional history, asked on 2026-07-20. The two reader answers are labeled replays of a recorded paired run: same reader, same evidence, with and without the notes. That pair is the unit the research scores.

Deterministic · liveReplayed answersNo weights served
1

Retrieve the shortlist

A lexical score counts the question's content words in each session; ties break toward the more recent session.

1
s3 · 2026-04-02lexical score 2 · carries a dated expression
lexical order
2
s1 · 2026-03-03lexical score 2 · carries a dated expression
lexical order
3
s4 · 2026-05-21lexical score 0 · carries a dated expression
moved from #4
4
s5 · 2026-07-15lexical score 0
moved from #3
5
s2 · 2026-03-18lexical score 0
lexical order

Deterministic stand-in for the typed re-ranker: for temporal questions it prefers sessions that carry dated expressions. The measured TypeSafe lane reorders the real shortlist and lifted answer turns in context from 84% to 95%.

2

Tier the context

Every retrieved session gets a code-built abstract; only the highest-ranked two keep their full text. Tiering replaces cutting every session to the same sliver.

full texts3 · 2026-04-02

Ran my first 10K yesterday in 54:30. My knees held up fine.

full texts1 · 2026-03-03

I moved to the Marina two weeks ago. The new place is 72 square meters and my commute is forty-five minutes now.

abstracts4 · 2026-05-21

We adopted a cat last Saturday and named him Milo.

abstracts5 · 2026-07-15

I'm selling the road bike I bought in March for $1,100 — asking $700 for it now.

abstracts2 · 2026-03-18

The pottery studio charges $40 per session and I signed up for Monday evenings.

3

Write the computed notes

Rebuilt for this question in front of you: relative dates resolved against the day they were said, the gap between the best-matching events, and the quantities that belong to the question.

inserted before the reader reads0 model calls · identical on every run
COMPUTED NOTES — written by code, not a model

Dated events (the user's own words, resolved)
  "I moved to the Marina two weeks ago."
      said 2026-03-03 (s1) → 2026-02-17 · 153 days before the question
  "Ran my first 10K yesterday in 54:30."
      said 2026-04-02 (s3) → 2026-04-01 · 110 days before the question
  "We adopted a cat last Saturday and named him Milo."
      said 2026-05-21 (s4) → 2026-05-16 · 65 days before the question

Gaps
  Moved Marina → Ran first 10K 54:30: 43 days (~6 weeks)

Quantities (each line quotes its sentence)
  no figures in sentences belonging to this question

The parser in this lab covers this fixture's constructions — a closed set of relative dates, currency, and a few units. The shipping engine in the repository covers the full set the reader is measured with.

4

The reader, paired

Gold for this question: 43 days. Reader answers are replays of recorded paired-run arms: the same reader, the same retrieved evidence, with and without the notes. No model executes on this page.

Small reader, no notes — replayReplay

You moved to the Marina in March and ran the 10K in April, so about two weeks later.

✗ judged wrong

Both sessions were in context. The reader never resolved “two weeks ago” against the day it was said, so it anchored the move to March and guessed the gap.

Small reader, notes + working — replayReplay
Marina move: 2026-02-17 (“two weeks ago”, said 2026-03-03).
First 10K: 2026-04-01 (“yesterday”, said 2026-04-02).
2026-04-01 − 2026-02-17 = 43 days.

43 days — I moved on February 17 and ran the 10K on April 1.

✓ matches gold

The dated events and the gap were already computed in the notes; the thinking step copies them and subtracts nothing itself.