Commit f527939
paper(working-memory-cliff): Phase 1B — FP32-weights control + arXiv-ready
Closes the quantization confound loop on the working memory cliff
finding and ships an arXiv-submission-ready tech report draft.
R4 (FP32-weights control, 6 trials): the cliff sits in the same
place when on-the-fly Q4 weight requantization is disabled.
Weight precision ctx=1024 ctx=1280
Q4 (default) 100% 0%
FP32 (TQ_NO_Q4=1) 100% 0%
Going from Q4 to FP32 weights eliminates any quantization artifact
but does not move the transition. The cliff is therefore a model
property — chat-template-anchored instruction-following robustness —
not a weight-quantization artifact and not a KV-cache artifact.
The above-cliff failure mode is identical between Q4 and FP32
weights: all six FP32 ctx=1280 trials produced wikitext continuation
("Doctors , followed by a role in How to Curse..."), the same
dominant failure mode as the Q4 grid.
CLI bug discovered during the seed-sweep attempt (R5): tools/quant.c
documents `-s <seed>` in --help but does not implement it. There is
no parser case for `-s`, no rng_seed config field, and the underlying
tq_sample_topp call hardcodes rng_state=42 per CLI invocation. The
result is that all 60 attempted seed-sweep trials degenerated to
"model path = <seed>" (e.g., `Loading model from 42... cannot open
'42'`). Documented in §5.6 as a limitation. Filing the CLI fix as a
separate quant.cpp issue. The seed-sweep artifacts are committed for
transparency (they are evidence of the bug discovery, not data).
Tech report draft v0.3 (docs/paper/working-memory-cliff.md, 292 lines):
- §1: TL;DR with concrete cliff numbers and the FP32 control finding
- §2: Related Work (KIVI, H2O, SnapKV, PyramidKV, NIAH, Lost in the
Middle, RULER, LongBench, MLC-LLM) with explicit comparison
- §3: Method — protocol, models, KV configs, grid
- §4: Results — six tables (1B, 3B, summary, neutrality, FP32 control)
+ failure mode taxonomy
- §5: Negative findings (prompt format trap, panic output, 8B problem,
single-language scope, single prompt format, seed-sweep CLI bug)
- §6: Discussion — what "long-context replaces RAG" means at the edge,
with the 0.4–0.78% effective-window numbers
- §7: Reproducibility — exact CLI commands, CSV file references, git
commit hash for fixed-version reproduction
- §8: Future work (8B+, mechanistic interpretability, cross-lingual)
- §9: References (9 citations, all open arXiv)
Submission package (docs/paper/):
- working-memory-cliff.md — single-source markdown
- working-memory-cliff.tex — auto-generated LaTeX (517 lines)
- md2tex.py — pure-Python markdown → arXiv LaTeX
converter, no pandoc dependency
- build.sh — pandoc-or-fallback build script
- arxiv-metadata.md — abstract (280 words), classification,
keywords, submission checklist
- hf-blog-draft.md — HuggingFace blog post (151 lines, friendly
tone, ready for publication)
- twitter-thread.md — 10-tweet launch thread + 5 anticipated
criticism responses
Next-step option for the user: the user can submit
working-memory-cliff.tex to arXiv directly, publish hf-blog-draft.md
to HuggingFace, and queue the twitter thread for simultaneous launch.
The CLI seed bug is a separate small fix tracked outside this commit.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>1 parent 56d750b commit f527939
14 files changed
Lines changed: 2226 additions & 11 deletions
File tree
- bench
- results/niah
- docs/paper
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
| 80 | + | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
| 94 | + | |
| 95 | + | |
| 96 | + | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
| 105 | + | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
| 109 | + | |
| 110 | + | |
| 111 | + | |
| 112 | + | |
| 113 | + | |
| 114 | + | |
| 115 | + | |
| 116 | + | |
| 117 | + | |
| 118 | + | |
| 119 | + | |
| 120 | + | |
| 121 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
| 80 | + | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
| 94 | + | |
| 95 | + | |
| 96 | + | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
| 105 | + | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
| 109 | + | |
| 110 | + | |
| 111 | + | |
| 112 | + | |
| 113 | + | |
| 114 | + | |
| 115 | + | |
| 116 | + | |
| 117 | + | |
| 118 | + | |
| 119 | + | |
| 120 | + | |
| 121 | + | |
| 122 | + | |
| 123 | + | |
| 124 | + | |
| 125 | + | |
| 126 | + | |
| 127 | + | |
| 128 | + | |
| 129 | + | |
| 130 | + | |
| 131 | + | |
| 132 | + | |
| 133 | + | |
| 134 | + | |
| 135 | + | |
| 136 | + | |
| 137 | + | |
| 138 | + | |
| 139 | + | |
| 140 | + | |
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
1 | | - | |
| 1 | + | |
2 | 2 | | |
3 | 3 | | |
4 | 4 | | |
5 | | - | |
| 5 | + | |
6 | 6 | | |
7 | | - | |
| 7 | + | |
8 | 8 | | |
9 | 9 | | |
10 | 10 | | |
| |||
60 | 60 | | |
61 | 61 | | |
62 | 62 | | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
63 | 80 | | |
64 | 81 | | |
65 | 82 | | |
| |||
0 commit comments