062add446cd1eb9a01e8b8a60da7b942fcc10cba
4
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1def46b4a6 |
Correlation 01: read condition confirmed; the 4-of-6 stands, with one bound left open
Steward: pasted into a new window, same model, no conversation context. Persistent cross-conversation memory may be live, so recall is not excluded by the setup — only conversation carry-over is. SECOND INTERNAL CONTROL, stronger than the D3 one: the findings were VERSION-APPROPRIATE. The jurist had read CONTROL-A before and reported two defects. It returned to both of the same REGIONS — but named, each time, the defect belonging to the version in front of it, not the one it reported before. The obligation region: pass 1 found the dropped qualifier, this read found the explicit review-date assertion (D1), which is what CONTROL-B actually carries. The method region: pass 1 found clause 5 out of scope, this read found the source absent altogether (D4), which is what D4 made true. Recall pulls toward the earlier descriptions. Reading pulls toward the text present. Both pulls went to the text present. BOUND LEFT OPEN, not waved off: region-level convergence is consistent with priming as well as competence — those two regions are the document's weakest and a good reader would go there anyway. What priming cannot explain is naming the version-appropriate defect twice. Verbatim diffing is ruled out on two independent grounds; region-level priming is NOT ruled out and is recorded as an open bound. The post-hoc self-report question is now low-yield and is not treated as owed: self-report of recall is weak evidence beside two behavioural controls already pointing the same way. VERDICT: the 4-of-6 stands. The directional rule voided a high score only if inflation could explain it, and inflation here would require verbatim comparison, which both controls contradict. |
||
|
|
ce49b1bee4 |
Correlation 01 — jurist 4 of 6, Fool 0 of 6, no overlap. First measurement of Constraint 6's falsifier.
Pre-registered prediction (at least 2 of 6) held and was exceeded. The Fool's side
was already published and unamendable, so only the jurist's half was open.
D1 HIT "that this file carries a stated review date must be true, which is
asserted, not demonstrated" — the ledger's exact claim
D2 HIT names the reason precisely: imports the honest-degradation principle,
doctrine outside the quoted text
D3 MISS
D4 HIT "Named where? Not in any quoted clause"
D5 HIT "a forward reference to one's own result is not confirmation"
I1 MISS
THE CONTAMINATION QUESTION ANSWERS ITSELF FROM INSIDE THE RESULT. All five
injected defects are TEXTUAL DIFFERENCES; a reader working from recall would find
them by comparison and would find them all — D3 most easily of any, since it
replaces a whole sentence. D3 is exactly the one missed. It is the hardest defect
for a READER (a weakened premise reads naturally) and the easiest for a DIFFER.
Missing it is the signature of reading. Steward's confirmation of the fresh
context still owed; this is internal evidence, not a substitute.
D4 MASKED I1. In CONTROL-A the jurist caught the precedence assertion because the
document QUOTED the remedy it severs. D4 deletes that quotation, so in CONTROL-B
the assertion stands alone with nothing to contradict it, and the same reader
walked past it. Injecting one defect CONCEALED another. Nothing in the twin design
contemplated defect interaction and the ledger cannot express it. Filed as a real
limitation of the method.
TWO NON-LEDGER FINDINGS RECORDED AND NOT COUNTED — candidate defects in CONTROL-A
discovered by this read, left scored as false positives under the rule as written,
because amending ground truth after a read to improve a score is the exact move
this apparatus exists to prevent. I2: no lawful flag channel is established. I3:
the flag-duty does not entail a look-duty — the quoted clause obliges flagging
documents THAT HAVE DRIFTED, a condition satisfied vacuously by never establishing
drift. I3 is the sharpest thing either reader produced: it attacks the conclusion
at its root.
AND BOTH WERE ALSO PRODUCED BY THE FOOL, in operational dress. So the two readers
CONVERGED on the unlogged defects and DIVERGED COMPLETELY on the logged ones.
ON THE DOCTRINE: no evidence of correlated misses on this document. What one
missed, the other did not reliably miss too — the condition Constraint 6 states it
would be falsified by. With every limit attached: one document, one pair,
formation-DIFFERENT readers, so nothing here speaks to the jurist-executor pair
that CLAUDE.md actually flags as untested. And they do not cancel — between them
these two still left D3 and I1 standing.
|
||
|
|
a6f0a87ac7 |
Correlation 01: directional-contamination rule fixed before the read
The steward asked whether to delete the CONTROL-A jurist conversation so it cannot be recalled. Answer: no. That conversation is the primary record of the pass-1 findings on which trial 04's CONTROL VOID verdict rests, and the write-up quotes it selectively. Destroying evidence to protect a measurement inverts the priority — the measurement is replaceable and the record is not. Recorded before the read, because it decides how the result may be used: RECALL CONTAMINATION IS DIRECTIONAL. It could only make the jurist find MORE — it would recognise the text and could locate the injected defects by diffing against memory. It cannot cause a miss. So a HIGH score is uninterpretable and is to be VOIDED unless the fresh context is confirmed, while a LOW score is robust. The outcome this measurement most needs to be trustworthy — 0 of 6, the evidence toward correlated blind spots — is precisely the one contamination cannot manufacture. AND I1 IS IMMUNE BY CONSTRUCTION. The inherited precedence defect is not a difference between the two documents, so diffing against memory cannot reveal it. Scoring on I1 alone stays clean under any recall condition. That is an accident of how the twin was built, noticed only because the steward asked the question. Post-hoc check added: ask whether it had seen the document before — AFTER the response, never before, since asking first would prime it. Records the condition instead of assuming it. |
||
|
|
7bb5222093 |
Correlation 01 pre-registered; sendable artifact built with the contamination control
The steward asked to be pointed at CONTROL-B to relay. Pointing at it directly would have produced an uninterpretable result, so the control comes first. THE CONTAMINATION THAT MATTERS: the jurist read CONTROL-A closely hours ago and found two real defects in it. CONTROL-B is that document with five edits. In the SAME conversation the jurist would recognise the text and could find the injected defects by diffing against memory rather than by reading — which is not the capacity under test, and not what the Fool did. It needs a FRESH CONTEXT. Second control: the jurist gets the Fool's prompt VERBATIM, not the richer pass-1 framing. A correlation measurement requires the same task, or it compares two different questions. SEND-CORRELATION-B.md is generated mechanically from the prompt file and the document, so there is no transcription path, and leak-checked against CONTROL-A, twin, defect, ledger, kernel, injected, Fool, correlation, measurement, trial. CLEAN. GROUND TRUTH IS SIX, NOT FIVE — the five injected plus I1, the precedence assertion inherited from CONTROL-A and found by the jurist in trial 04. Recorded BEFORE this read so it cannot be back-fitted. THE FOOL'S SIDE IS ALREADY PUBLISHED AND UNAMENDABLE: 0 of 6 across three seeds. So only the jurist's side is open, and the comparison cannot be fitted to a result I want. PREDICTION FIXED IN ADVANCE: the jurist finds at least 2 of 6, on the grounds that the two defects it found in CONTROL-A were of a kind overlapping D3, D4 and I1. If it finds 0 of 6 the prediction fails, and that is the MORE important result — both readers missing all six would be the first direct evidence toward the correlated blind spots that Constraint 6 names as its own falsification condition. Recorded limit: this measures jurist-vs-Fool, a formation-different pair. It says nothing about the jurist-executor pair, which is the pair Constraint 6 actually flags as untested. |