[FIX] fool trial 03 VOID; degraded-guard rebuilt with a positive control

Trial 03 ran and produced nothing gradeable. Recorded as VOID rather than
omitted, because an absent row reads as a trial not attempted.

Two independent failures, both found by reading the output, neither by a check,
and every check passed:

1. The harness certified a run with no answer. Qwen emitted its scratchpad as
   plain prose ('Here's a thinking process:', zero <think> tags), so the tag
   regex reported reasoning_present:false and recorded all 2,944 words of
   deliberation as the ANSWER; the token ceiling then cut it off mid-sentence
   before the answer began. degraded:null. The guard tested the STRING for
   emptiness while its field claimed a property of the RESULT — which is the
   previous session's open question, answered by the instrument built to audit
   instruments. Trial 02 had listed the inline-scratchpad problem as Open; the
   harness closed it assuming inline meant tagged.

2. Worse: the design forbade the region it was measuring. The self-exemption
   axis lives in Part VII; the anti-echo constraint added in trial 02 tells the
   reader to skip author-named limitations, and the scratchpad shows the model
   reaching Part VII and leaving it, citing that constraint. Silence about
   self-reference is indistinguishable from obedience. The axis was unmeasurable
   by construction, independent of the truncation. Trial 02's fix and trial 03's
   document were each sound alone; their interaction was not.

Guard now reports every degradation, not the first: empty answer, untagged
scratchpad, and token-ceiling truncation. reasoning_present renamed
think_tag_found — it was a claim about a regex wearing the name of a claim about
the model. test_degraded_guard.py is a positive control that runs against the
actual trial-03 artefact, not a synthetic one; it caught a false positive in the
first version of my own guard (a bare 'okay' matched a legitimate sentence).

The false-positive control STILL has never been run. Two attempts, two unrelated
causes — the obstacle is the instrument and the design, not the model.
This commit is contained in:
David F Glidden
2026-08-02 16:50:59 +02:00
parent b678d2f57b
commit eda11e559b
8 changed files with 608 additions and 15 deletions
+77 -10
View File
@@ -147,6 +147,27 @@ def git_revision() -> str | None:
THINK_RE = re.compile(r"<think>(.*?)</think>", re.DOTALL | re.IGNORECASE)
# Qwen3.6 does not always tag its scratchpad. In trial 03 it opened with the bare
# line "Here's a thinking process:" and never emitted a <think> tag, so the tag
# regex reported reasoning_present=false and the whole deliberation was recorded
# as the answer. These are openings of *deliberation about the task*, which no
# answer to this prompt begins with — the prompt forbids summarising the document
# and asks for named assumptions.
UNTAGGED_SCRATCHPAD_RE = re.compile(
r"^\s*(?:"
# First-person statements of intent about the task.
r"(?:here(?:'|’)s|here is|let(?:'|’)s|i(?:'|’)ll|i will|i need to|i should|"
r"first,?\s+i)\b"
# Interjections, which must actually be interjections. A bare `okay` matched
# "Okay is not a word this document uses, but its approach…" — a sentence that
# belongs in an answer. The punctuation is what distinguishes the two.
r"|(?:okay|ok|alright|right|so)\s*[,:]"
r")[^\n]{0,80}"
r"(?:thinking process|thought process|think through|reasoning|analyz|approach|"
r"plan\b|break (?:this|it) down|work through|go section by section)",
re.IGNORECASE,
)
def split_reasoning(raw: str) -> tuple[str | None, str]:
"""
@@ -205,7 +226,7 @@ def run_model(
sampler_kwargs = {"temp": sampling["temperature"], "top_p": sampling["top_p"]}
sampler = make_sampler(**sampler_kwargs)
return generate(
raw = generate(
model,
tokenizer,
prompt=text,
@@ -214,6 +235,17 @@ def run_model(
verbose=False,
)
# Re-encoding the decoded text is an ESTIMATE, not the true generated count —
# encode(decode(x)) is not guaranteed to round-trip. It is reported as an
# estimate and only used to detect the token ceiling, where being a few tokens
# out cannot change the verdict.
try:
generated_tokens = len(tokenizer.encode(raw))
except Exception:
generated_tokens = None
return raw, generated_tokens
def main() -> None:
ap = argparse.ArgumentParser(description="Run one Fool trial, reproducibly.")
@@ -264,23 +296,50 @@ def main() -> None:
sys.exit(f"FATAL: seed requested but could not be set ({exc}).")
started = datetime.now(timezone.utc)
raw = run_model(args.model, prompt_text, document_text, enable_thinking, sampling)
raw, generated_tokens = run_model(
args.model, prompt_text, document_text, enable_thinking, sampling
)
finished = datetime.now(timezone.utc)
reasoning, answer = split_reasoning(raw)
# An empty answer is NOT "the checker found nothing". It means the model
# produced only a reasoning trace, or nothing at all. Those are different
# results and must never be recorded as a finding of silence — that is the
# precise confusion trial 02 fell into. Flag it loudly and record it.
degraded: str | None = None
# A result is degraded whenever what was recorded as `answer` is not an answer.
#
# The original guard tested only for emptiness, which is a property of the
# STRING while the field claims a property of the RESULT. Trial 03 walked
# straight through it: 2,944 words of untagged deliberation, cut off at the
# token ceiling before the answer began, recorded as `degraded: null`. Trial 02
# mistook silence for restraint; that guard would have let trial 03 mistake
# deliberation for a finding. Every condition below must therefore be reported,
# not just the first.
problems: list[str] = []
if not answer.strip():
degraded = (
problems.append(
"EMPTY ANSWER: the model produced no text outside its reasoning trace. "
"This is a harness/generation failure, NOT a finding of 'nothing found'. "
"Do not grade it as restraint."
)
print(f"\n*** {degraded} ***\n", file=sys.stderr)
if reasoning is None and UNTAGGED_SCRATCHPAD_RE.match(answer):
problems.append(
"UNTAGGED SCRATCHPAD: the output opens as deliberation about the task, "
"and no <think> tag was emitted, so it was recorded as the ANSWER. "
"reasoning_present=false here means 'no tag was found', NOT 'the model "
"did not deliberate'. Do not grade this as the checker's findings."
)
ceiling = sampling["max_tokens"]
if generated_tokens is not None and generated_tokens >= ceiling - 2:
problems.append(
f"TOKEN CEILING: generation stopped at the max_tokens limit "
f"(~{generated_tokens} of {ceiling}). The output is CUT OFF, not "
f"complete. Anything absent from it may simply never have been reached."
)
degraded: str | None = "\n".join(problems) if problems else None
if degraded:
print(f"\n*** DEGRADED RUN ***\n{degraded}\n", file=sys.stderr)
stamp = started.strftime("%Y%m%dT%H%M%SZ")
slug = f"trial-{args.trial}-{stamp}"
@@ -307,8 +366,16 @@ def main() -> None:
},
"output": {
"raw_words": len(raw.split()),
"reasoning_present": reasoning is not None,
# Named for what it actually tests. The old key was `reasoning_present`,
# which read as a claim about the model and was in fact a claim about a
# regex: trial 03 deliberated for 2,944 words and this field said false.
"think_tag_found": reasoning is not None,
"answer_words": len(answer.split()),
"generated_tokens_est": generated_tokens,
"hit_token_ceiling": (
None if generated_tokens is None
else generated_tokens >= sampling["max_tokens"] - 2
),
"degraded": degraded,
},
"environment": environment(),