[FIX] Degraded guard: deliberation is two cases, not one
Filed in trial 04's tool review, now closed. The guard reported UNTAGGED
SCRATCHPAD ... "Do not grade this as the checker's findings" for both of the two
situations it can see, and they are opposite:
trial 03 — deliberation that ran into the CEILING. No answer ever existed. VOID,
and the absence of findings is NOT restraint.
trial 04 — deliberation that COMPLETED. The answer follows the scratchpad in the
same file. Perfectly gradeable once extracted. NOT void.
Collapsing them would have thrown away six good runs; not distinguishing them
would have graded trial 03's silence as restraint. The guard now branches on
hit_token_ceiling and says which case it is.
Controls added for all four shapes, including the two the trials actually
produced and a clean answer that merely hit the ceiling — truncation is reported
separately and is not a scratchpad problem.
The guard does NOT auto-extract the embedded answer. A heuristic split would be a
new failure mode in the instrument whose entire job is to not silently mis-report
what it has. It flags; a person extracts.
This commit is contained in:
@@ -88,6 +88,37 @@ check("clean answer not flagged", bool(UNTAGGED_SCRATCHPAD_RE.match(CLEAN_ANSWER
|
||||
for i, text in enumerate(NEAR_MISSES):
|
||||
check(f"near-miss {i}", bool(UNTAGGED_SCRATCHPAD_RE.match(text)), False, repr(text[:50]))
|
||||
|
||||
print("\nThe distinction trials 03 and 04 paid for — deliberation is not one case:")
|
||||
# Replicates the guard's branch logic without importing mlx-dependent code.
|
||||
def classify(answer: str, reasoning, generated_tokens, ceiling):
|
||||
untagged = reasoning is None and bool(UNTAGGED_SCRATCHPAD_RE.match(answer))
|
||||
hit = generated_tokens is not None and generated_tokens >= ceiling - 2
|
||||
if untagged and hit:
|
||||
return "VOID"
|
||||
if untagged:
|
||||
return "EMBEDDED"
|
||||
return "OK"
|
||||
|
||||
check(
|
||||
"trial 03 shape (deliberation + ceiling) → VOID",
|
||||
classify(SYNTHETIC_SCRATCHPAD, None, 4096, 4096), "VOID",
|
||||
"this is the run where no answer ever existed",
|
||||
)
|
||||
check(
|
||||
"trial 04 shape (deliberation, completed) → EMBEDDED, not void",
|
||||
classify(SYNTHETIC_SCRATCHPAD, None, 4428, 12000), "EMBEDDED",
|
||||
"collapsing this into VOID would have discarded six good runs",
|
||||
)
|
||||
check(
|
||||
"clean answer, completed → OK",
|
||||
classify(CLEAN_ANSWER, None, 900, 12000), "OK",
|
||||
)
|
||||
check(
|
||||
"clean answer that hit the ceiling → not EMBEDDED",
|
||||
classify(CLEAN_ANSWER, None, 12000, 12000), "OK",
|
||||
"truncation is reported separately; it is not a scratchpad problem",
|
||||
)
|
||||
|
||||
print("\nTagged output still splits correctly:")
|
||||
r, a = split_reasoning("<think>deliberating</think>\nThe answer.")
|
||||
check("reasoning extracted", r == "deliberating", True)
|
||||
|
||||
Reference in New Issue
Block a user