Public-safe authorization trace
The sentence can stay the same.
The permission can change.
Signature demonstration / 00
Matched content
One statement. Two evidence states.
Watch authorization diverge.
“The board approved Project North.”
A
Source / statusVerified summaryConfirmed
Permitted useAwaiting authorization
B
Source / statusUnverified noteProvisional
Permitted useAwaiting authorization
Same textSource and status govern permitted use
Authorization instrument / 01
Evidence → draft → release
Synthetic cases / protected method absent
Trace the release boundary.
Input field / E
Evidence / context
∅No evidence entered.
Authorization stops before model analysis.
Candidate field / C
Candidate draft
Draft not yet authorized
0 claims
Policy—
Permitted-use authorization
Output field / R
Released answer
Nothing has crossed the boundary.
ModelGPT‑5.6 contract
Policy ownerDeterministic code
Review statePending
Select a case to inspect how evidence status constrains release.
Control separation / 02Bounded public demonstration
The model describes. The policy decides.
01Structured extraction
GPT‑5.6 returns schema-constrained claims and evidence relationships. Text inside evidence is treated as untrusted data.
02Coded authorization
The release decision exists in deterministic policy. Model output cannot turn failure, conflict, or missing support into ALLOW.
03Constrained release
SAGA returns the draft, a qualified reconstruction, or no answer. Private taxonomy, prompts, thresholds, and production logic are not present.
Evidence record / 03Claims kept separate
What the evidence says. And what it does not.
01Software verification
The final Pytest, Ruff, Mypy, JavaScript, build, desktop, and mobile results describe this Build Week implementation only.
02Live integration proof
One Responses API request returned HTTP 200, requested gpt-5.6, returned gpt-5.6-sol, and matched the strict schema. Public deployment remains deterministic to prevent unrestricted API spending.
03Earlier method evidence
An earlier bounded SAGA demonstration classified 28 of 30 cases into the expected action band (93.33%): PASS 10/10, REVIEW 8/10 and HIGH RISK 10/10. This was a limited demonstration of the broader SAGA framework—not a benchmark of this Build Week implementation or a general reliability guarantee.
“Fluent chain-of-thought should not be treated as proof of validity when evidence verification is absent.”
PASS, REVIEW, and HIGH RISK are historical action bands. They are not mapped onto ALLOW, QUALIFY, and WITHHOLD.