LongMemEval-S
Published LongMemEval benchmark; 500 questions, with 470 scored after excluding 30 abstentions. SHA-256 d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442.
Firekeep retrieved at least one labeled evidence-bearing session in the first ten results for 459 of 470 scored questions. A benchmark is most useful when the number travels with its denominator, protocol, limits, and artifacts. They are all here.
For each question, Firekeep had ten chances to return a session containing the labeled evidence. It found one or more for 459 questions.
Open the machine-readable resultRead the headline
LongMemEval-S contains 500 questions over synthetic chat histories. Thirty abstention questions are excluded from this aggregate, leaving 470 scored questions. Each scored question names the session or sessions containing its evidence.
Firekeep received the question, returned its first ten retrieval results, and earned one hit when any of those results came from a labeled evidence session. It earned 459 hits and missed 11. The exact ratio is 97.65957446808511%, displayed as 97.7%.
This measures retrieval. It asks whether the labeled evidence reached the context window. It does not score the final answer a language model might generate from that evidence.
Complete result
Recall answers whether any labeled evidence appeared. Coverage asks how much of all labeled evidence appeared. Mean reciprocal rank rewards finding the first relevant result early. NDCG rewards relevant results throughout the top ten, with more weight near the top.
| Question type | Hits | Scored | Recall@10 |
|---|---|---|---|
| Knowledge update | 72 | 72 | 100% |
| Multi-session | 121 | 121 | 100% |
| Single-session assistant | 56 | 56 | 100% |
| Single-session preference | 29 | 30 | 96.6667% |
| Single-session user | 62 | 64 | 96.875% |
| Temporal reasoning | 119 | 127 | 93.7008% |
Frozen protocol
The run began with empty, isolated Qdrant, Neo4j, and Redis stores. Authentication was disabled to match the benchmark's single-workspace posture. The fixture's sessions were serialized inside each namespace, with up to eight namespaces in flight.
Published LongMemEval benchmark; 500 questions, with 470 scored after excluding 30 abstentions. SHA-256 d6f21ea9d60a0d56f34a05b609c79c88a451d2ae03597821ea3d5a9678c3a442.
mxbai-embed-large:latest, digest 468836162de7f81e041c43663fedbbba921dcea9b9fefea135685a39b2d83dd8, through Ollama 0.32.4 on an NVIDIA GeForce RTX 5080.
format=raw, top_k=10, token_budget=10000, with eight recall workers across independent namespaces.
Ingestion covered 23,867 fixture session rows and 124,366
/memory/learn calls, including 13 repeated session keys. HTTP 200 responses with
status=partial were rejected and retried rather than counted as completed.
Acceptance gates
Dated comparison
The 470 scored question IDs were paired against the previously published run. The current result adds 74 net hits, a change of 15.7447 percentage points. The comparison describes the measured difference; it does not assign that difference to a single cause.
| Outcome | Questions |
|---|---|
| Evidence found in both runs | 384 |
| Evidence missed in both runs | 10 |
| Previous miss → current hit | 75 |
| Previous hit → current miss | 1 |
The exact two-sided McNemar p-value is 2.0381730293306602e-21.
Reproduce and inspect
1d4405a256e922a1494f251bb8ef78cd81f9d382.top_k=10, and token_budget=10000 across all 500 questions.Interpretation
These boundaries make the number more useful: they say exactly where it applies.