Every claim on this site
comes with its measurement.

The benchmark is LongMemEval-S: 500 questions asked against long, simulated chat histories — questions that span sessions, depend on dates, track updates, and punish guessing. Answers are graded by an automated grader, and every grader is named with its score, because graders differ in leniency. Each improvement below was measured as a controlled pair: the same model, the same retrieved chats, exactly one change. Run the same setup twice and scores move by about ±7 — any claim smaller than that is noise, and we do not make it.

Raw text · chat, 12 January

“We signed Norsk Dental as a client yesterday. I'm the account lead. Their contract runs twelve months and is worth $84,000.”

  • Norsk Dental became a clientclient_of(norsk_dental, us) · from 2026-01-11“signed … as a client yesterday”
  • You led the account — superseded 1 Mayaccount_lead(norsk_dental, you) · 2026-01-12 → 2026-05-01“I'm the account lead.”
  • The contract is worth $84,000 over twelve monthscontract_value(norsk_dental, $84,000 / 12 mo)“worth $84,000”

A small model reads fine.
It computes wrong.

We handed a small open model every piece of evidence for each question — perfect retrieval — and it still lost half the questions that span sessions or dates. Reading the logs, the failures had one shape: the right sentences, the wrong arithmetic. This is what it got wrong, in its own answers:

  • Counts drift across sessions — three festivals counted for four.
  • Relative dates stay unanchored — “just got back,” said July 15, asked Aug 5, answered without the three weeks.
  • Gaps get mis-subtracted — “4:22 minus 4:10” answered as 17 minutes.
  • Sums drop an item — two of three road-trip legs; 50 lb of feed for 70.

A model can read a sentence; it cannot be trusted to subtract two timestamps or anchor “two weeks ago” to the day it was said. So we stopped asking it to.

Let code do the arithmetic.

Before the model reads, deterministic code scans the retrieved chats and writes a block of notes: every relative date resolved against the day it was said, gaps between the events the question asks about, and the quantities that belong to the question — each line quoting the sentence it came from.

The block is identical on every run, costs no model call, and never guesses. The model's job shrinks to what it is good at: reading.

318 → 359The same small model, the same retrieved chats, before and after adding the code-written notes — 41 more questions right out of 500. One controlled pair, one automated grader (GPT-4o), no training involved.
Inserted before the model reads0 model calls
COMPUTED NOTES — written by code, not a model

Dated events (the user's own words, resolved)
  "two weeks ago I moved to the Marina"
      said 2026-03-03 → 2026-02-17 · 153 days before the question
  "ran my first 10K yesterday"
      said 2026-04-02 → 2026-04-01 · 110 days before the question

Gaps
  Marina move → first 10K: 43 days (~6 weeks)

Quantities (each line quotes its sentence)
  $1,100  "bought the road bike in March"
  $700    "selling it for $700"
  exactly two figures → difference: $400

A real block from the reading lab's fixture — watch it being written live in your browser.

Then train on the structure.

With the notes proven, we distilled a small reader from its frontier teacher — trained on real sessions with the notes present, and taught to write its own working before answering. Four measured steps, each a controlled pair under one grader (DeepSeek), took it to within fifteen questions of the model that trained it.

How the small reader closed on its teacher

controlled pairs · one grader (DeepSeek) · noise ±7
Dates and arithmetic, computed by codeBefore the model reads, a code-written block resolves every relative date against the day it was said, states the gaps, and lists the quantities the question needs — each line quoting its sentence. No training, no extra model call.
+29▲
Trained to show its workingThe reader learned to write the dated facts it relies on and the arithmetic before answering — then one final answer line. Only that line is graded.
+10▲
Smarter selection of which chats to includeThe unit of retrieved chat is chosen from the question itself — whole sessions for some question types, single turns for others — with the date notes rebuilt to match.
+17▲
Re-ranked shortlistA typed re-ranker reorders the candidate chats so the ones carrying the answer actually reach the model. Answer-bearing chats in context rose from 84% to 95%, for $0.28 of calls.
+13▲
The small reader, full system — against its teacherAn open 4-billion-parameter model, distilled from GLM 5.3 Flash — the same frontier model that taught it, and the yardstick. Fifteen questions short on the 500.
425/500teacher 440

Each row is a measured pair from the method and every run — including the runs that failed: every negative result in the log is published with its configuration.

Two small open models. Measured, never served.

Both are fine-tunes of open-weights models, small enough to run on a laptop. Neither is hosted, downloadable, or required by this site — they live in the repository as research artifacts, with every run that measured them.

Model comparison
Not served here

This site ships no model weights and calls no model API. The playground and labs run deterministic code in your browser. Where model output appears, it is either a clearly labeled replay of a recorded run or the optional third-party open model your own browser loads via WebLLM.

Writer · 2.3B

The translator

Turns chat and documents into facts, authors the query, extracts policy claims. A fine-tune of Gemma 4 E2B, served as a 4.6 GiB quantized file under llama.cpp — a laptop scores what a cloud GPU scores.

Extraction benchmark
87/103
Query authoring
27/31
Policy decisions, unseen fixture
20/20
Unjustified approvals
0
Reader · 4B

The answerer

Reads retrieved chat history and answers, handed the code-written notes before it starts. A fine-tune of Gemma 4 E4B, distilled from GLM 5.3 Flash over real sessions — the teacher is the yardstick it is measured against.

Long-memory benchmark, 500 questions
425/500
Teacher, same grader
440/500
Code-written notes, no training
+29
Trained to show its working
+10
Smarter turn selection
+17
Re-ranked shortlist
+13
Notes before reading. Every relative date resolved against when it was said, every line quoting its sentence.Structure before chats. The memory's own facts, dated and deduplicated, placed before the raw conversations.One call to translate. None to phrase — recalled facts never leave the process.Empty results explain themselves. The engine names the goal that matched nothing and the swap that would return rows.

Local embeddings (nomic-embed-text) tie the hosted model on the semantic route. Method and every run.

Don't take our word for any of it.

Run the sixty-second demo, work a lab, or reproduce a benchmark from the repository — the claim and the check ship together.