OSL deep audit: plan, rulings, build, documents, gates, and design
Audit snapshot: 2026-08-16 10:26 PDT. Canonical source was integration/full at 4d13b8a04f48767eb9c72c2432ae9ef01caf814a. The recorded artifact was SHA-256 a219c5718fef9a87b6be03305741d73e16da00b6952ccc3c73e9768d02fd9271, attributed to commit a5b8e134b76848d81e914282e550ba12be286053.
1. Executive verdict
OSL is not release-ready. The most consequential finding is that the installed AutoScrub definition is internally impossible: TASK 6093 says the scheduled product must be Find-only, disclose that it never deletes automatically, and contain zero Find-and-delete controls, agreements, routes, jobs, or claims; TASK 7785 says the product must perform reviewed provider deletion. Both are ticked. The staged artifact independently fails the controlling TASK 6093 rule: the exact disclosure is absent and the destructive inventory is non-zero (todo/08-scrub-website-release.txt:1930-1941; todo/49-native-surfaces-release.txt:157-183; proof/autoscrub-status.md:27-43). This is a release-blocking safety contradiction.
The wider problem is measurement. The completion script can print that the artifact failed and still exit success. The design gate returns MET while the manifest contains 11 explicit failures. The universal screen gate is not invoked by the tick path. One parity checker claims zero pixel difference without reading pixels. The review animation checker watches decorative dots outside the product. Old captures can become “fresh” by being recommitted. These instruments cannot reliably distinguish broken from intentionally different.
The recorded EXE is real, hash-correct, unsigned, and a Windows GUI executable. It is not the current canonical build, not a complete installer bundle, not reproducibly tied to one clean source commit, and not observed usable. It is 23 commits behind the canonical source snapshot. It omits native-cover-writer, yet the Windows UI can expose a six-template substitute as a “verified model” and “AI Covertext.” That is a false installed feature claim, not merely an untested one (proof/canonical-artifact.json:2-14,35-43; RULING D66 at RULINGS.txt:2859-2879).
Literal plan status is 4,336 ticked and 261 open across 4,597 tasks. That is not the current-release truth. At least 31 open tasks are explicitly stale, superseded, historical, or post-V1; four are owner-deferred; TASK 1410 is already settled by D64/D64(a). No more than 225 opens are current actionable release work. Conversely, multiple ticked tasks cite proof that explicitly says their full finish line was not met.
2. Scope, reading method, and authority
This audit read the complete numbered task corpus, all 2,977 lines of RULINGS.txt, every top-level deliverable/reference document, the current and competing artifact/proof manifests, all evidence-file metadata and task-to-evidence relationships, the named gate/checker implementations, the review/parity harnesses, all 72 .dc.html design declarations, and the current route/state/dialog sources. Large generated trees were inventoried exhaustively by file, type, date, provenance field, parseability, and authority status; semantically decisive receipts, scripts, contradictions, and samples were read line by line. Binary facts were independently checked with SHA-256, PE headers, Authenticode directory inspection, strings, and Git ancestry/diff measurements.
Authority was applied in the project's own order: owner rulings, the design authority for visual screens, numbered tasks, then canonical code (deliverable/01-FULL-SPEC.md:1-3). A tick was evaluated under the README rule that it means the stated check passed and proof can be found, not merely that somebody worked on it (todo/00-README.txt:125-142). A source or component test was not treated as packaged Windows behavior; the README itself now makes that distinction (todo/00-README.txt:3-29).
The workspace was changing during the audit, so volatile counts and commit distances are explicitly timestamped. Historical documents were not rewritten to make them agree. The required correction is to label and supersede them, not erase the record.
3. Consequence rank and immediate ownership
| Rank | Finding | Consequence | Fixer |
|---|---|---|---|
| P0 | TASK 6093 and TASK 7785 define mutually exclusive AutoScrub products; the staged build fails 6093 | A fallible classifier can coexist with installed deletion authority despite the controlling find-only safety rule | Agent, unless owner explicitly reverses Round 11 |
| P0 | Windows presents six static templates as verified local AI while native writer is absent | False privacy-product capability and wrong-code-path green tests | Agent; owner only if redefining the product claim |
| P0 | plan-complete.sh discards artifact-verifier failure | Automation can declare completion with no current artifact | Agent |
| P0 | check-needs.sh reports design manifest MET with 11 failures | Downstream review can proceed without authority for live surfaces | Agent; owner decides keep/remove genuinely ambiguous surfaces |
| P0 | Screen tick gate 7054 is not wired into the tick path | Any screen task can close without parity proof | Agent |
| P1 | Recorded artifact is 23 commits behind and its “clean source” is 329 commits behind its attributed commit | “Canonical build” does not identify one reproducible source tree | Agent |
| P1 | 4,336 green tasks include explicit unfinished proofs, retired work, and proofless blocks | Progress and dependency counts do not equal product capability | Agent; owner only for genuinely missing decisions/tests |
| P1 | Owner-review freshness is not observation freshness | Owner can approve a week-old or impossible-at-commit image as current | Agent |
| P1 | Ruling IDs, supersession, and enforcement are not stable | Correct code under the latest ruling can fail an old checker; violations can pass unseen | Agent, plus one owner choice on public-name rules |
4. Release blocker: AutoScrub fails its controlling safety definition
TASK 6093 supersedes earlier deletion-oriented tasks and requires scheduled AutoScrub to be Find-only. It requires the exact sentence “AutoScrub finds possible matches for review; it does not delete them automatically,” before scheduling; zero installed Find-and-delete controls, agreements, direct routes, durable jobs, or claims; a real Windows due event; a real approved account; labelled true and false positives; zero delete attempts; and byte-exact provider after-state after restart and refresh (todo/08-scrub-website-release.txt:1930-1939).
The task is wrongly ticked. Its own proof says the Windows/provider-account evidence was unavailable, used a controlled provider double, and did not establish the installed binary, real Task Scheduler event, or refresh/re-login (evidence/6093.md:29-40). Its finish-line checklist leaves the real-account/Windows/installed/provider clause unchecked and ends: TASK 6093 FULL FINISH LINE: NOT CHECKED (evidence/6093.md:137-144). Under the README's tick rule, that heading must be open.
The fresh artifact audit confirms a product failure, not just missing proof. For exact artifact a219... at a5b8..., the disclosure has zero occurrences on the active screen/binary; the installed answer is explicitly “not zero”; deletion-agreement identifiers are in the EXE; source retains Find-and-delete authorization; and the UI makes unattended-run claims (proof/autoscrub-status.md:3-13,27-43). The run store is process-local and in-memory, so there is no durable schedule/provider due proof (proof/autoscrub-status.md:19-25). No real account, true-positive/near-miss pair, provider after-state, refresh, relogin, or quarantine has been proved (proof/autoscrub-status.md:45-55).
TASKS 7785 and 7785b are the opposite product and are also ticked. 7785 requires real provider deleters and fails if the reviewed run deletes zero (todo/49-native-surfaces-release.txt:157-183). Canonical source at the artifact commit defines find_and_delete, its authorization, agreement, and action in apps/osl-hub-ui/src/autoscrub-deletion-permission.ts:10-24,67-103; autoscrub-account-page.ts:128-131 says AutoScrub “can delete.” A zero installed deletion inventory cannot pass 7785, and a non-zero one cannot pass 6093.
Why it matters: a release gate can choose whichever green task supports the desired claim. That is precisely how a safety boundary becomes ceremonial.
Fix: agent unticks 6093, marks 7785/7785b superseded for scheduled shipping behavior, removes installed deletion controls/agreements/claims/routes, adds the exact pre-schedule disclosure, implements durable find-only discovery, rebuilds, then performs the complete real-account Windows event and after-state test. Owner action is needed only to reverse Round 11 in a new explicit ruling.
5. Release blocker: Windows “AI Covertext” is a six-template substitute
The artifact record explicitly says features are core,desktop,whatsapp-qa-shell and it was built without native-cover-writer because llama-cpp-sys-2 cannot cross-compile to x86_64-pc-windows-gnu from WSL (proof/canonical-artifact.json:35-36). D66 says no Windows test can cover that absent code and the Windows AI claim remains unproven (RULINGS.txt:2859-2879). The intended spec correctly says the Windows release must not advertise it (deliverable/01-FULL-SPEC.md:15-17,326-330).
The installed source path violates that boundary. At artifact commit a5b8e134, apps/osl-hub/src/bundled_model_pack.rs:1-6,27-32 unconditionally embeds osl-covertext-tiny-v1.oslmodel and calls it a local AI pack. Lines 162-179 load six templates; lines 199-227 choose and hash a template, requested shape, and random bytes. There is no llama inference. main.rs:10478-10501,11055-11059 unconditionally exposes/materializes it. ai_carrier.rs:55-108 calls it a local model and lets the AI selection use it. The UI says “AI Covertext uses the verified model on this device” and “AI Covertext will write the next cover on this device” (apps/osl-hub-ui/src/cover-writing-controls.ts:43-52; main.ts:10296-10306, both at a5b8e134). Direct EXE strings contain the pack name and six cover= templates.
TASKS 3518, 3522, and 3794 are ticked as shipping AI-button behavior (todo/12-final-audit.txt:488-498,572-582,7320-7330), but the 3522 proof is Linux/library and core/no-default-feature testing (evidence/3522.md:15-33,55-71). Those checks can remain valid as library proof. They cannot prove the packaged Windows feature.
Why it matters: the Windows interaction can appear green while executing different code from the feature the label promises. A signed template file and entropy do not make fixed phrases a local AI model.
Fix: agent compiles the status, commands, pack materialization, and UI behind native-cover-writer; until a native build exists, the six-template path must be labelled template/word-bank cover and the AI control must be absent or explicitly unavailable. Retick Windows claims only after the exact packaged EXE attests to the feature and the real UI causes real model inference and a decoded round trip.
6. The recorded artifact exists but is not current canonical
Independent measurement confirms the file at canonical-artifact.json:6 is 22,052,352 bytes and hashes exactly to a219c571..., matching lines 7-8 and 37. PE headers identify a 64-bit Windows GUI executable. The PE Authenticode Security Directory is zero, so it is unsigned. That matches the owner decision to ship unsigned, not the plan's signed-artifact wording.
The manifest attributes it to a5b8e134... (proof/canonical-artifact.json:14). At 10:26 PDT, integration/full was 4d13b8a...; the artifact was 23 commits behind. Those commits include Home notification-state repair, wide OSL Chats geometry repair, friend cancellation/visibility work, review-audit fixes, and capture-verdict pairing. These affect visible and behavioral review scope, not merely proof prose. Calling a219... “the current integration/full build” is false.
check-needs.sh nevertheless reports it MET because it allows any ancestor up to 400 commits behind (check-needs.sh:117-151). The comments explicitly delegate relevance to each task, but the token's message calls the bytes “the canonical Windows artifact” (:139-151). The token measures recorded bytes plus ancestry. It does not measure current behavior.
Fix: agent either freezes a named release commit and stops moving it, or rebuilds the tip. For non-release exploratory work, rename the token artifact:recorded-ancestor. For release and visual acceptance, require exact release commit or a machine-readable affected-path analysis showing that every intervening commit is irrelevant to the measured surface.
7. Artifact provenance, packaging, and usability are not established
The same manifest records clean_checkout.source_commit = e2a099a... and artifact commit = a5b8e134... (proof/canonical-artifact.json:9-14). Git shows the alleged clean source is 329 commits behind the artifact commit. A clean checkout of e2a cannot prove a clean build of a5b8.
The manifest names TASK 7751 as verifier (canonical-artifact.json:12), but project documents disagree about that task. proof/7751-build-attempt.json:2-26 records an older source and a Windows check exit 101 with no package. evidence/7751.md:1-73 later records a successful clean build and an NSIS installer at e2a..., still not the a5b8... raw EXE. This is a document-generation conflict, not a reproducible chain from one tree to the current bytes.
The record binds only the raw EXE. A loader and Tor sidecar exist beside build output, but the canonical record binds a loader only inside a patched home_qa companion and does not bind the normal runtime companion set (canonical-artifact.json:15-33). No canonical installer is recorded beside the EXE. Installation, sidecar startup, upgrade, uninstall, and clean-machine sufficiency remain unproved.
Usability is explicitly NOT OBSERVED; the 30-second manual stay-alive is explicitly not TASK 7944 (canonical-artifact.json:2-4,38-43). A living process is not a visible window, correct monitor, completed bootstrap, sign-in, or usable route. TASK 7488 separately records that the packaged artifact failed both hidden/unreachable B6 warning cases and says no tick is claimed (evidence/7488.md:3-20,120-129), yet the plan heading is ticked (todo/45-ship-the-observers.txt:514-526).
Fix: agent builds from a detached clean worktree at exactly the frozen release commit, records commands/toolchains/features/status, and emits a bundle manifest covering installer, EXE, loader, Tor sidecar, updater metadata/signatures, and hashes. Then run the exact TASK 7944 observation plus install/update/uninstall on a clean Windows VM. Owner supplies credentials only when genuinely required.
8. What is actually in the artifact
| Item | Measured state | What may honestly be claimed |
|---|---|---|
| Binary | Hash/size match; PE32+ x86-64 Windows GUI | A real Windows executable exists |
| Source identity | Attributed to a5b8..., 23 commits behind snapshot canonical | It is an ancestor build, not current canonical |
| Features | core,desktop,whatsapp-qa-shell | Those Cargo selections were recorded; behavior still needs observation |
| Native AI writer | Absent | No Windows native-model claim |
| Six-template “AI” path | Present and exposed | A template writer exists; calling it verified AI is false |
| Authenticode | Absent | Correct under unsigned decision; Windows warnings are expected |
| Companions/installer | Not bound as one canonical bundle | Raw EXE is not a distributable-product proof |
| AutoScrub | Exact disclosure absent; deletion inventory non-zero; no durable provider schedule | Release blocker |
| Runtime usability | NOT OBSERVED | No usable/sign-in claim |
| Carrier/UI strings | Numerous carrier and surface strings are present | Inventory only; strings do not prove routes, accounts, sends, or receipt behavior |
9. Plan census and what the number means
The numbered corpus contains 4,597 task headings in 56 task files; adding 00-README.txt makes the requested 57-file plan set. There is no 56-*.txt. Literal status is 4,336 ticked, 261 open. File headers and the README catalogue are not reliable inventories: several old counts differ materially from the task headings (todo/00-README.txt:35-54,144-153).
The 261 is not a release denominator. Conservative explicit-authority classification finds at least 31 stale/superseded/historical/post-V1 opens, four intentionally owner-deferred opens, and settled TASK 1410. Thus at most 225 are current actionable release work. That upper bound still includes real owner/account/physical work; it is not a completion percentage.
The same ledger uses [x] for at least three meanings: finish line passed, capability retired, and duplicate dropped. Retired GMX/Mail.com work and dropped tasks increase the same “done” count as working product. The plan needs first-class done|superseded|retired|dropped|deferred|open status.
10. Exact per-file plan audit
“Current” below means the open heading still names real work under current rulings. “Stale” is used only where explicit task/ruling/header text supersedes, retires, defers, or moves it post-V1.
| File | Coverage | Done / open (total) | Open audit |
|---|---|---|---|
| 00 | protocol, authority, progress rules | 0 / 0 (0) | current protocol; catalogue counts stale |
| 01 | foundation, build switches, machines, checks | 92 / 4 (96) | 4 current |
| 02 | allowlists, friends, block/unblock | 220 / 1 (221) | 1 current |
| 03 | onboarding, recovery, restore, name/key safety | 202 / 10 (212) | 6 current; 4 stale: 0334a, 0466a, 0463, 0463b |
| 04 | burn, timers, view-once, attachments | 181 / 11 (192) | 11 current |
| 05 | Settings, Home, tiles, friend screens | 197 / 20 (217) | 20 current |
| 06 | external chat carriers | 385 / 12 (397) | 12 current |
| 07 | browser/email carriers and OSL Chats | 246 / 9 (255) | 9 current; green retired-carrier history is misleading |
| 08 | Scrub, AutoScrub, website, installer, journeys | 142 / 16 (158) | 12 current; 1617/1618 stale; 1410 settled; 1619 has wrong signing clause; 6093 false tick |
| 09 | safety/payment/update/attack gaps | 324 / 8 (332) | 8 current |
| 10 | timed delete | 66 / 8 (74) | 8 current |
| 11 | placing text in other apps | 55 / 5 (60) | 5 current; old clipboard foundation partly superseded |
| 12 | final audit and late hardening | 611 / 56 (667) | 55 current; 5174b stale; Windows AI greens false |
| 13 | receive pipeline | 248 / 5 (253) | 5 current |
| 14 | X/Instagram/Messenger | 59 / 1 (60) | 1 current |
| 15 | mailbox reading | 82 / 13 (95) | 1 current; 12 OSL Mail opens post-V1 under D9 |
| 16 | Eye and observation groups | 26 / 4 (30) | 4 current |
| 17 | final Eye/control gap | 19 / 0 (19) | none |
| 18 | allowed places | 42 / 2 (44) | 2 current |
| 19 | allowance/top-up | 42 / 0 (42) | none |
| 20 | profiles/posts/stories | 25 / 1 (26) | 4656b stale red proof for removed rule |
| 21 | voice | 24 / 0 (24) | none open; 4712/4713 false ticks |
| 22 | discovery | 29 / 1 (30) | 1 current |
| 23 | multi-device | 40 / 0 (40) | none |
| 24 | enclaves/beacons/custom roles | 49 / 0 (49) | none; opening cut ruling obsolete |
| 25 | Tor | 42 / 0 (42) | none |
| 26 | VM/account readiness | 14 / 16 (30) | 7 live: 4959-4965; 9 explicitly historical: 4957, 4957b, 4966-4971b (26-vm-readiness.txt:1-8) |
| 27 | uncovered gaps | 50 / 0 (50) | none |
| 28 | acceptance rulings | 96 / 2 (98) | 5201/5201b stale after 5214 ladder |
| 29 | triage rulings | 14 / 2 (16) | 6582/6583 real future work, owner-deferred (29-triage-rulings.txt:1-6,156-176) |
| 30 | owner rulings | 13 / 0 (13) | none |
| 31 | redesign scope | 78 / 0 (78) | none |
| 32 | final UI spec | 65 / 0 (65) | none |
| 33 | AutoMod/bots | 21 / 0 (21) | none |
| 34 | integration | 41 / 2 (43) | 2 current; 7010/7012 appear shipped but falsely parked |
| 35 | timer modes | 8 / 0 (8) | none |
| 36 | data balance | 18 / 0 (18) | none |
| 37 | screen parity | 76 / 3 (79) | 3 current; 7049/7054 false-green dependencies |
| 38 | reverification | 34 / 0 (34) | none |
| 39 | unfailable gates | 26 / 0 (26) | none |
| 40 | cover integrity/readback | 22 / 18 (40) | 18 current |
| 41 | wiring gap | 30 / 2 (32) | 2 current |
| 42 | proof pyramid/top gate | 13 / 10 (23) | 10 current |
| 43 | lane integration | 18 / 2 (20) | 2 current |
| 44 | Windows compile/visibility | 17 / 0 (17) | none |
| 45 | observer shipping | 40 / 0 (40) | none |
| 46 | capability routing | 9 / 1 (10) | 1 current |
| 47 | product spine recovery | 16 / 0 (16) | none |
| 48 | carriers/mail | 22 / 8 (30) | 8 current |
| 49 | native surfaces/release | 26 / 6 (32) | 4 current; 7792/7793 owner-deferred identities (RULINGS.txt:2135-2139) |
| 50 | test certainty | 20 / 2 (22) | 2 current |
| 51 | owner batching | 16 / 0 (16) | none |
| 52 | parity conveyor | 28 / 0 (28) | none, but tracked evidence is stale/red |
| 53 | completion gate | 12 / 0 (12) | none |
| 54 | owner screen review | 18 / 0 (18) | none; harness contract currently disagrees with task |
| 55 | behavior audits | 9 / 0 (9) | none |
| 57 | missing design capabilities | 18 / 0 (18) | none; review-send proof is source-only |
| Total | 4,336 / 261 (4,597) | 261 literal; no more than 225 current actionable |
11. Open tasks that are done, stale, or falsely blocked
TASK 1410 is done in reality as the wording judgment it asks for. D64 calls the owner's pass “the verdict TASK 1410 asks for” and says it is settled; D64(a) confirms the four elements on the surface (RULINGS.txt:2824-2842,2957-2977). The task remains open because unrelated account scanning, deletion, and scheduling clauses were welded onto a wording review (todo/08-scrub-website-release.txt:139-148). Agent should tick the wording task and move behavioral requirements to separate tasks; the pass must not be treated as AutoScrub proof.
TASKS 7010, 7012, and 7328 appear complete in product/evidence but are falsely parked. 7010/7012 use commits created inside required disposable Git repositories; 7328's 012345... is explicitly a fixture commit. The citation matcher loses that context and treats them as fabricated. Canonical contains their task commits, and 7328's six focused tests passed. This holds a ten-task downstream chain (deliverable/04-PLAN-AUDIT.md:23-40,100-108). Agent fixes the candidate-level context parser and records the shipped work.
The dispatcher also has circular or observer-defect blocks: 1100/1130/1165 require isolated profiles before running the tasks that create them; 3512 requires the source file it is assigned to create; 3788 uses an obsolete second-physical-device token; 3514b/3669 are UNKNOWN because no checker looks in the OSL Chats store; 7241 uses an unimplemented composite token (deliverable/04-PLAN-AUDIT.md:51-81). UNKNOWN here means “the instrument did not look,” not “the account is absent.” These are agent fixes.
Real owner work remains real: unaided friend install, actual payments, real-person/provider journeys, account sign-ins/labels, owner visual judgments, and the two explicitly deferred identity provisions. It must not be auto-ticked merely because gates are inconvenient.
12. Ticked tasks that should not be ticked
The following are high-confidence false greens, not keyword guesses:
| Task(s) | Contradicting evidence | Required action |
|---|---|---|
| 6093 | Full finish explicitly NOT CHECKED; real Windows/account/provider clauses absent (evidence/6093.md:137-144) | Untick; product repair and full rerun |
| 7785/7785b | Requires provider deletion, contradicting controlling 6093 zero-inventory rule (todo/49...:157-183) | Mark superseded for scheduled shipping behavior |
| 4712 | No done:/proof: and no owner A/B/C choice (todo/21-voice-channels.txt:264-273) | Owner chooses; agent records |
| 4713 | Chose B without 4712; real call, provider/payment receipt, and Windows recording unchecked (evidence/4713.md:53-87) | Untick and run real journey after choice |
| 3518/3522/3794 | Windows feature absent under D66; proof is Linux/library (RULINGS.txt:2859-2879; evidence/3522.md:15-33) | Reclassify library proof; untick Windows claim |
| 0726a | Evidence opens “full finish line is not met” (evidence/0726a.md:5) | Untick until native Windows proof |
| 0792 | Evidence says finish line not checked because screenshot/Windows proof absent (evidence/0792.md:51) | Untick |
| 0810a | Evidence says literal finish line not checked complete (evidence/0810a.md:156) | Untick |
| 4213 | Eleven real-carrier observations absent and evidence says it remains open (evidence/4213.md:72) | Untick |
| 5402/5402b/5403 | Evidence says work remains open/must remain open (evidence/5402.md:58; 5402b.md:73; 5403.md:94) | Untick and rerun exact boundary |
| 6420 | Evidence says finish line not checked (evidence/6420.md:76) | Untick |
| 6939 | Evidence says full task finish line not checked (evidence/6939.md:96) | Untick |
| 7784 | Browser/profile/provider discovery finish line not checked (evidence/7784.md:63) | Untick |
| 7031 | Evidence says not checked as production/deployed pass (evidence/7031.md:115) | Untick or split unit/deploy tasks |
| 7488 | Exact artifact failed both B6 fault-report steps; evidence explicitly says no tick/product pass is claimed (evidence/7488.md:3-20,120-129) | Untick; repair installed warning and rerun |
| 1229,1235,1241,1247,1253,1259,1277,1290,1306,1311,1316,1324,1329,1334,1345,1350,1356,1363 | Their own visual receipts say later Liam PASS/review remains open; examples at evidence/1229.md:76, 1235.md:5, 1306.md:101, 1363.md:3 | Untick review claims; preserve rejected captures as history |
| 1262r/1269r and ordinary GMX/Mail.com rows | D35/D36 retired carriers; green describes retirement, not shipped capability (RULINGS.txt:1846-1885) | First-class retired, excluded from capability count |
A conservative scan found 756 ticked evidence files containing at least one unchecked Markdown box. That is a triage set, not a verdict: mutation tests and explicit exclusions legitimately contain unchecked examples. But the table above was manually verified against the task's own required clause. Thirty-one ticked blocks also lack the required proof: metadata: 1011, 1037c, 3082, 3095, 3118, 3119, 3129, 3130, 3131, 3132, 3133, 3134, 3203, 3609s, 3609v, 3690, 3691, 3693, 3694, 3752, 3765, 3773, 3928, 3955c, 4019, 4061, 4260a, 4601, 4602, 4712, 4916. They may be true, but the ledger cannot prove them under todo/00-README.txt:125-142.
13. Ruling inventory and effective precedence
RULINGS.txt contains 75 D-series headings plus five opening RULING: records and older named decisions. D1-D50 and D53-D68 exist; D51/D52 do not. Addenda include D17a, D18a, D21a, D33a, D45(a/b/c), D54(a), and D64(a). The latest textual entry is D64(a), after D68. Sequential prose labels are therefore not safe database keys.
Major resolved chains are clear when read in order:
- D21's performance numbers are voided by D21a; D44 is the current provisional/approved limit authority (
RULINGS.txt:1371-1440). - D31's absolute clipboard ban is narrowed through D47/D49 and finally D54/D54(a): direct user-chosen copy is allowed; automatic/background clipboard remains forbidden (
RULINGS.txt:1687-1714,2366-2424,2425-2451,2506-2547). - D35/D36 remove GMX and Mail.com (
RULINGS.txt:1846-1885). - D39's 71-page authority was overtaken by the 72-page package and later review rulings (
RULINGS.txt:1940-1960; current manifestpackageSize:72atproof/7049-route-design-manifest.json:5). - D45(a) is corrected/refined by D45(b)/(c); raw demo-pixel mismatch is not the same as structural mismatch (
RULINGS.txt:2206-2280). - D64's wording pass is corrected by D64(a), which confirms all four elements (
RULINGS.txt:2824-2842,2957-2977). - D65 exempts exactly five undesigned routes from visual-parity findings, not from behavior proof (
RULINGS.txt:2844-2857). - D66 makes the Windows native-AI gap explicit; D68 is the latest screen-review defect state (
RULINGS.txt:2859-2879,2923-2955).
The plan/checkers do not consistently apply these effective states. Old tasks and receipts remain green/current-looking after supersession, and the ruling checker still executes obsolete opening assertions.
14. Live ruling contradictions and governance defects
1. AutoScrub: TASK 6093 versus 7785/7785b is a live, mutually exclusive plan contradiction. Rulings outrank later plan drafting; agent fix unless owner explicitly reverses.
2. Public names: Decision 11 says mixed-case [a-z A-Z 0-9 _], length 1-16 (RULINGS.txt:82-83); D47 says lowercase [a-z0-9_], length 3-30 (:2372-2376). D47 is later but does not say it supersedes Decision 11. Owner must choose one; agent then updates code/tests/copy.
3. Ruling-ID collision: canonical docs/release/code-signing-decision.md:1-14 calls the 2026-07-31 unsigned decision D63; current RULINGS.txt D63 is account preconditions (RULINGS.txt:2809-2822). “D63” is not a stable identity. Agent assigns immutable/date-namespaced IDs and repairs citations.
4. Voice/beacons: opening machine-state cuts are obsolete after later “nothing left unbuilt”/D16 scope, yet the old checker still rejects voice. Agent registry fix.
5. OSL Mail: earlier hard-cut language is refined by D9 to post-V1 preservation. File 15's 12 OSL Mail opens must not gate V1 (RULINGS.txt:964-985).
6. Clipboard: prohibited automatic clipboard implementation tasks remain green dependencies even after D31/D54. Agent marks implementation clauses superseded and preserves only zero-activity regression proof.
7. Review counts: D62's prose count and named IDs disagree. Use IDs, not prose totals (RULINGS.txt:2763-2786).
8. Incomplete authority: rulings 20(b)-20(i) are said accepted but their clauses are not present; decision 17 delegates a choice without recording its outcome (RULINGS.txt:96,126-127). These cannot be reconstructed by a checker. Owner supplies missing decisions only if still operative; agent marks otherwise unresolved/historical.
15. Signed-versus-unsigned conflict
TASK 1619 requires the “exact signed artifact” (todo/08-scrub-website-release.txt:1920-1928). Related tasks 1608/1610 are ticked with signed-build wording, 1609 remains open on a signed install, and 6292/6426 retain “final signed release” language (todo/08...:1779-1818,1982-2048).
The owner decision dated 2026-07-31 is to ship Authenticode-unsigned, with an independently signed manifest; minisign is not Authenticode (integration/full:docs/release/code-signing-decision.md:1-14,37-39). The artifact's zero Security Directory matches it. deliverable/03-WHAT-LIAM-MUST-DO.md:14-32 correctly tells Liam to use the exact unsigned artifact and treats the plan text as subordinate.
Fix: agent rewrites 1619 and all release journeys to require the exact hash-bound, manifest-signed but Authenticode-unsigned release bundle and expected Windows warning. Retired signed-build clauses must not stay green as proof of an Authenticode capability. Owner action is not needed unless reversing the decision.
16. The ruling checker does not enforce the ruling corpus
ruling-check.py --plan-text reads only the five opening RULING: records, not the D-series (ruling-check.py:18-46). It checks only undone/non-held tasks (:88-105), so it cannot see false-green 6093/7785. Two opening assertions it does enforce—voice absence and beacon/custom-role cut—are obsolete. Current execution therefore punishes the intended shipping scope.
Coverage mode parses ordinary D headings but its ID regex cannot parse parenthetical D45(a/b/c), D54(a), or D64(a) (ruling-check.py:242). Runnable checks exist only for the five opening assertions and D31/D35/D36/D37 (:161-195). D31 is source-only despite its packaged-behavior language. The saved report claims 44 holes and stops at D39 (proof/7327-ruling-coverage.md:1-56); current execution rejects it as stale against the current RULINGS hash (ruling-check.py:382-385). plan-complete.sh then scrapes a prose holes: N phrase using a regex that does not even match the report's “Hole count” wording (plan-complete.sh:77-84).
Fix: agent creates one versioned ruling registry with immutable IDs, supersedes, effective state, affected task/product surfaces, and executable packaged checks. Check open and green tasks separately. Regenerate against the exact current RULINGS digest and fail if an operative ruling has neither a check nor an explicit non-machine-review disposition.
17. Document-corpus authority is undefined
OSL-AUDITS is a historical warehouse, not a single current document set. It contains roughly 19,500 files across root documents, deliverables, reference exports, proof, evidence, lane logs, and unintegrated lane material. Within the specifically requested document/proof families there are thousands of generated files: 4,355 evidence Markdown receipts, 529 top-level proof JSON files, 352 proof Markdown files, hundreds of captures, several full design generations, backups, and staging trees.
proof/README.md:1-12 says where outputs go but defines no current-authority index, immutability rule, source/artifact binding, observation timestamp, schema, or supersession field. Consequently, “canonical,” “latest,” “live,” and “current” are prose adjectives used by mutually inconsistent generations.
Fix: agent adds an authority index for root/reference/proof/evidence with current|historical|superseded|staging, created/observed time, source commit, complete artifact digest set, producer/checker digest, and supersedes. Move backups/staging out of the searchable live namespace. A real file with a real PASS must not be current merely because nothing points elsewhere.
18. Deliverable documents
| Document | Status | Finding and evidence |
|---|---|---|
01-FULL-SPEC.md | Current intended specification, not build certification | It explicitly says that at line 3; correctly records Windows AI absent, find-only AutoScrub, unsigned release, and non-release state (:17,256-262,326-344,377-384). It is bound to source 2b971..., already behind snapshot canonical. Regenerate code-discrepancy citations when source moves. |
03-WHAT-LIAM-MUST-DO.md | Same-day queue snapshot already stale | Line 3 says 261 open/105 ready/25 agent, while the later measured plan audit says 104 ready/24 agent (04-PLAN-AUDIT.md:11-21). It correctly resolves 1619 to unsigned (03...:14-32) but still asks Liam to review 1410 at :463 after D64/D64(a) settled it. Stamp and regenerate from semantic task status. |
04-PLAN-AUDIT.md | Current dated operational audit | Snapshot is explicit (:1-4). Its findings—false parks, circular needs, stale account observers, and work agents can do—are current as of 10:13 (:23-81,100-108). It is not a release/build verdict and will age with dispatcher state. |
.02-rulings-inventory.tmp.md, .02-spec-inventory-notes.md, .02-task-proof-notes.md | Hidden work products, not final deliverables | They are valuable audit inputs but must live in a work/staging directory or be finalized. Their presence makes the delivery directory look complete while 02 itself is absent. |
19. Root documents
The Aug 5/6 area-audit family 00-INDEX.md and 01-onboarding.md through 31-osl-homepage.md describes old owner build b3ce...; the index itself says the audit is current only through Aug 5 and the integration route was stale. Individual files use current-tense labels such as “RIGHT NOW,” “Works today,” and “Missing.” Every member needs a historical banner, not merely the index.
OSL-EVERYTHING.txt:3580-3677, DECISIONS.txt:69-80, and SPEC.md:108-110 promise AutoScrub deletion. TASK 6093 and the current intended spec require scheduled find-only behavior. SPEC.md:1-7 says it folds in rulings and “wins,” making its stale deletion claim particularly dangerous. Mark these sources historical/superseded; preserve owner quotations separately from the effective product contract.
RESUME-HERE.md:1-18 calls itself the single entry point but describes 26 files/2,986 tasks. The current plan has 57 files including README and 4,597 tasks. PRODUCT.txt, SCOPE-CHANGES-2026-08-06.txt, OSL-TODO.txt, PLAN-48H.md, PLAN-DESIGN.txt, FINAL-AUDIT.txt, CLEANUP-PLAN.txt, STORAGE-RULING.txt, and UI-FEEDBACK.txt are dated intent/history, not reconciled current authority.
RECONCILIATION-LOG.md and the _BLUEPRINT/_AUDIT/_RESEARCH/_STYLE/_TASK/_TEST files remain useful process/history if labelled. MODEL-TIERS.md is routing, not product truth. deferred.txt and its backups need one current pointer.
20. Reference documents and design exports
Current/useful with stated limits: 006-007-BEHAVIOR.md when overlaid with D53/D54; AGENT-WINDOW-PLACEMENT.md as an operational rule; the three COVER research files and DISCORD-AUTOMOD-AND-BOTS.md as research, not build proof; LOCAL-CARRIER-IDENTITIES.md as a dated inventory; and PROJECT-CAPABILITY-ROUTES.md as a routing map that delegates live state to regenerated machine state.
Stale or hazardous: MACHINE-STATE.md says “right now” but was generated Aug 15 and instructs readers to regenerate if older than the task (reference/MACHINE-STATE.md:1-10). It names an older artifact/head (:12-24), 71/70 design declarations (:83), and older account state (:55-77). MACHINE-AND-CAPACITY.md is a dated snapshot. GATE-FIX-DISPATCH.md, GATE-REMAINING.txt, and REPASS-QUEUE.md are Aug 6 imperative queues and are superseded. FORMAT-CONTRACT.md, INTEGRATION-LESSONS.md, OSLCHATS-POLISH.md, and PLAN-WRITER-BRIEF.md are historical inputs unless re-ratified. Tool/patch/build-note files are not product evidence and need a bound source/toolchain before reuse.
The asset tree contains multiple competing exports: Aug 6/8/9 designs, an Aug 13 zip, Aug 11 redesign scope, canon-1280x800, dated canon-1280x809, and staging trees. Filename and dimensions do not establish authority. Agent adds reference/INDEX.json naming one current design root, retired roots, ruling overlays, and hashes.
21. Proof receipts: current, stale, and contradictory
Current and useful within limits:
canonical-artifact.jsonis the current byte locator, but not a clean-source/package/usability proof (:2-14,35-43).autoscrub-status.mdis current, exact-digest-bound, and appropriately negative (:3-13,27-55).task-7060-20260816T170343Z-blocked.jsonis a fresh negative preflight: its artifact does not match checkout head, onboarding routes remain unreferenced, the expected command is absent, and no capture set exists (:22-45,58-71).7049-route-design-manifest.jsonis the current structured map, but contains 11 failures (:638-649).
Stale/misleading if read as current:
verify7944-artifact.json,7317/canonical-windows-artifact.provenance.json,build-inventory-run.md, andpreflight-rerun.mddescribe predecessor hashes/commits. The latter two called an older executable canonical and reported exit 78/no window. Their negative history is useful; their “current” label is not.canonical-artifact.json.bak-*leaves old “canonical” records beside the current name.spec-vs-latest-gui.md:1-9,28-30,60-61calls itself latest and recommends reviewed AutoScrub provider deletion, contradicting 6093.build-inventory-run.mdtreats D65's five ruled-undesigned routes as design failures; D65 requires them classified unavailable-by-ruling, while preserving behavioral proof (RULINGS.txt:2844-2857).7900-owner-console.*and7950-live-review-harness.jsonbind older/no source identity; “live” is not provenance.
Review generations contradict one another: evidence/7997.md:35-51 says ready 0/still 69; proof/7997-owner-review-site.json:1-7 says ready 31/still 38; later 8003 evidence says 35 ready. These can all be honest historical generations. Without a supersession index, none is the current count.
Machine integrity is also weak: one proof JSON (task-7060-20260815T124647Z-blocked.json:51) is malformed; hundreds of zero-byte outputs exist, including an empty required-looking receipt; 456 paths explicitly contain staging/backup/legacy markers. Structured ingestion must parse and schema-check required receipts, not count path existence.
22. Evidence receipts: historical work is being mistaken for current product proof
There are 4,355 evidence Markdown files; 4,256 start with a task heading. Only 2,040 mention a commit and 762 mention a SHA/digest. 2,544 point into /home/liamw/osl-exec-* lane worktrees; none points to /home/liamw/osl-integrate. Age distribution is overwhelmingly historical: 22 files date Aug 5, 845 Aug 6, 438 Aug 7, 557 Aug 8, 182 Aug 9, 261 Aug 10, 470 Aug 11, 141 Aug 12, 722 Aug 13, 421 Aug 14, 230 Aug 15, and 56 Aug 16.
Old evidence is not worthless. It proves the scope it actually ran. The defect is flattening component, mutation, fixture, lane, live provider, Windows package, and owner-review evidence into one PASS. Examples responsibly say what did not run: evidence/0334b.md:132, 0328.md:55-56,75, 0135.md:5,54,65, and 0310b.md:15,152,155. The consuming ledger ignores that scope.
Of 4,336 ticked tasks, 4,310 have an exact same-ID Markdown evidence file; 26 do not. Of 261 open tasks, only 24 have exact same-ID evidence, usually blocked or rejected history. A filename therefore cannot determine current status.
Fix: agent requires a receipt schema with task ID, result, each scope proved/unproved, fixture/live authority, source commit, reached-canonical commit or patch identity, complete artifact digest, commands/exits, checker digest, observation time, and supersession. Only an explicit exact-artifact outcome supports a current-build claim.
23. Owner-review freshness remains unmeasured
The known 74-of-76 week-old incident is structurally reproducible. Current owner-materials contain 85 capture references. They generally bind subject JSON to commit 9f916... but provide path/hash without captured_at, renderer/source digest, or a real freshness verdict. Git-history comparison found 37 references last changed Aug 16; 6 Aug 10; 13 Aug 9; 15 Aug 7; 1 Aug 6; and 13 paths absent at their asserted commit. At least 35 references are days old and 13 are impossible under their declared source. proof/owner-materials/0028.json:12-22 calls an Aug 6 image a shipping-product capture under the later commit.
The newer TASK 7997 checker still dates the Git commit that last changed an image, not when the image was observed (integration/full:apps/osl-hub-ui/scripts/task-7997-owner-reviews-lib.mjs:161-189). Recommitting an old image makes it fresh; an identical new recapture remains old. Failure to obtain a timestamp returns without failing (:180-184). Gate receipts are read as self-reported PASS JSON rather than rerun (:278-298).
Fix: capture process emits a bound receipt containing observation time, exact artifact SHA, exact source commit, route/state/seed, renderer/checker digest, image/design SHA, dimensions/DPR, and Windows session identity. Review refuses any image predating its source, absent at its asserted commit, over maximum age, or from a non-current artifact. Git file age is not capture freshness.
24. check-needs.sh: two false MET definitions
The artifact token measures file existence, recorded SHA, ancestry, and lag up to 400 commits (check-needs.sh:117-151). It does not check PE type, architecture, startup, feature inventory, native modules, package companions, signing policy, route behavior, or usability. It can report MET over a renamed non-PE file if the record/hash/ancestry agree. Current USABILITY = NOT OBSERVED demonstrates the semantic gap.
The design token validates that all 72 declarations and accounting entries are readable, but it explicitly accepts UNREFERENCED rows when a reason exists and never requires failures.length == 0 (check-needs.sh:154-221, especially :168,205-221). The current manifest has 11 failures at proof/7049-route-design-manifest.json:638-649; the checker returns MET. TASK 7049's own finish line says one unreferenced route must exit 1 (todo/37-screen-parity.txt:157-171). This is the next clear example of a measurement that cannot distinguish “accounted for as broken” from “satisfied.”
Account needs also conflate cached/self-authored state with live observation. The record is accepted for seven days (check-needs.sh:224-250); future dates are not rejected. The separate verifier uses recent store writes/markers for sign-in and has hard-coded UNKNOWN tokens. deliverable/04-PLAN-AUDIT.md:51-71 shows circular, impossible, and observer-missing requirements.
Fix: split tokens by what they actually establish; fail design on every live unpaired route and non-empty failure list; derive the live route set in the same invocation; add PE/feature/startup/package tokens; use service-specific non-secret account/session observation; reject future/stale timestamps.
25. plan-complete.sh: the top answer can fail open
Clause A correctly counts 261 unticked headings but cannot distinguish current from retired/post-V1 (plan-complete.sh:48-51). Clause B claims every live requirement maps to a check that passes, yet it only verifies acceptance is non-empty; it does not execute or validate the acceptance (:53-75). Clause C scrapes the first prose holes: N phrase from stale Markdown (:77-85).
The fatal defect is clause D. It runs the exact-current artifact verifier followed by || true, then reads PIPESTATUS[0] (plan-complete.sh:87-102). || true turns the shell result into success. The Python can print D artifact FAIL and exit 1 without setting global fail. If other clauses pass, the script can print FINISHED and exit 0 (:156-163).
Clause E accepts non-empty transcript/digest strings without hashing/opening them or binding to artifact (:103-119). Clause G chooses the first worktree containing a checker, not necessarily canonical checker bytes (:129-145). Clause F passes if stale Markdown contains verdict: green anywhere (:148-153).
Fix: agent captures real exit codes, makes every parser/checker absence nonzero, uses versioned JSON receipts bound to inputs/checker source, executes acceptance checks, verifies transcript digests, and adds mutants for missing/wrong/stale artifact, stale ruling report, false green prose, and dirty noncanonical checker.
26. record-done.sh: prose length and vocabulary decide truth
The tick boundary selects the lexicographically first of TASK.md and TASK-*.md, not the newest or invocation-bound receipt (record-done.sh:19-22). It requires lane exit 0, evidence existence, and 400 bytes (:61-75). Its central truth test is a prose regex for known admission phrases with layers of exclusions (:78-160). A required clause can pass by using a different honest phrase; an unrelated mutation phrase can falsely block; a table/quote/task-ID exclusion can hide the admission.
Commit validation is conditional on helper executability (:163-190,225-259), so missing checkers fail open. Untick freshness relies on mutable file mtime (:192-223). It never parses or executes the task's done when, validates dependencies at record time, binds source/artifact/capture, or requires task-specific evidence. It then edits the heading and inserts done/proof (:268-294). Already ticked returns success without revalidation (:273-275).
This explains the false-green table: long honest “not done” reports were historically accepted until their exact phrasing was added. The gate is learning vocabulary, not measuring completion.
Fix: agent requires an invocation-specific machine receipt and executable acceptance or explicit human-review schema per task class. Parse every required clause; bind evidence path/digest, source, artifact, checker, and observation. Missing mandatory helper is UNKNOWN/nonzero. Owner defines only which task classes may close by judgment rather than execution.
27. Product reach, cited commits, and parked-task release
check-code-reached-product.sh fails open if the repository/canonical ref is absent (:46-48). It discovers work by task-ID commit subject (:50-72). With no match it passes unless evidence fits one narrow source-path-plus-verb regex (:74-92). If any one historical task-named commit is an ancestor, it passes even when the final repair is stranded (:95-104). It explicitly does not inspect content (:32-39). The earlier matcher also re-offered shipped work 85 times; its comments document that failure (:54-59).
check-cited-commits.sh:39-52 skips whole lines containing broad words such as hash, fixture, or mutant, and accepts any Git object type rather than ${sha}^{commit}. That produces both false accusations (7010/7012/7328 disposable/fixture commits) and false acceptance (a blob or skipped fabricated “commit hash”).
release-externally-parked.sh reads only evidence/$id.md, while the recorder accepts suffixed files (:70-75). It checks only the first 20 hex strings against one repository (:90-97), uses a much narrower admission regex (:98-102), selects candidates outside the lock and resets them later without revalidating the row (:50-63,107-114). A task can gain a real own-fault or new blocker in between.
Fix: lane receipt records exact implementation commit(s), patch IDs, affected paths, task result, and blocker provenance. Verify the exact final patch is in frozen canonical and touches the claims. Candidate-context parsing must understand disposable repositories without skipping real citations. Revalidate typed blocker/evidence generation inside the lock.
28. Other top-level gates and account/deploy checks
plan-gate.sh ignores producer exits, suppresses parser errors, defaults absent metrics to zero, and exits success unless callers remember --strict (plan-gate.sh:31-35,90-97). A parser crash becomes the passing value. Strict must be default; report-only should be explicit.
verify-accounts.sh treats marker directories plus any file modified within 14 days as live (verify-accounts.sh:70-85). Future attestation dates pass; Signal requires two sessions but receives only one store argument (:55-66,128). Gmail and nine composite tokens are hard-coded UNKNOWN (:147-189). Use service-specific bounded session/count probes and distinguish inaccessible from absent.
ship-review-when-green.sh counts error TS text through a pipeline (:24). If TypeScript exits 78, times out, or crashes before rendering with no matching line, grep -c returns zero: the known false MET. It does not check npm ci exit (:63), does not run 7950/7997/8001 acceptance before deploy, ignores Wrangler exit in favor of text, and logs any HTTP response—including 404/500—as PUBLISHED (:64-80). Capture pipeline statuses; run all gates; deploy an immutable commit/digest manifest; fetch it with HTTP 200 and exact byte verification.
29. Parity checkers that do not observe parity
TASK 7054 says no screen work can pass around its gate (todo/37-screen-parity.txt:288), but record-done.sh, plan-gate.sh, guard.sh, run-task.sh, dispatch.sh, and todo/next.sh never invoke it. Its standalone checker validates caller-supplied JSON shapes, not artifact/design/capture bytes (integration/full:apps/osl-hub-ui/scripts/task-7054-screen-parity-tick-gate.mjs:9-28). Any screen task can bypass it; a crafted receipt can satisfy it.
The legacy task-7923-strip-conveyor.mjs creates a parity register from source hashes/hooks and hard-coded assertions, then writes structural_differences:0, pixel_difference_percent:0, and GREEN without loading an image, DOM, browser, EXE, or design geometry (integration/full:apps/osl-hub-ui/scripts/task-7923-strip-conveyor.mjs:58-81,118-177). CSS can move the strip off-screen while the checker claims pixel-perfect parity.
The newer canonical 7923 conveyor is materially better but tracked proof is bound to old a83d... and all 38 routes are red/unpaired. Its capture is Vite/Linux headless Chromium with a synthetic Tauri state seam, not the Windows EXE; the manifest lacks artifact SHA/commit and forces reduced motion. TASK 7923 is nonetheless ticked (todo/52-parity-is-a-conveyor.txt:123).
TASK 7055 authenticates coordinated tool echoes, not canonical tool implementations. TASK 7965 hard-codes D43 defects and substitutes them for observations when none are supplied (task-7965-structural-parity.mjs:33-109,149,187,230-238). TASK 7902 maps eight distinct TASK 4022 states to the same image and validates labels/hashes, not pixels/state. These are useful contract/ledger checks only if named honestly.
Fix: one packaged-artifact parity command enumerates normalized reachable routes/states/dialogs, drives exact Windows artifact, captures with provenance, compares DOM/geometry/pixels against exact design/accepted differences, and proves sensitivity by mutating a real control. Source-contract and disposition ledgers must not publish visual verdicts.
30. Review harnesses can present static, stale, or unsent work as live
build-flow-review.py classifies every verdict except strings containing fail, mismatch, or ungradable as green. ?, UNKNOWN, BLOCKED, CRASHED, and NOT OBSERVED become ok (build-flow-review.py:63-69). It copies any existing capture without artifact SHA, commit, capture time, or content revalidation (:72-132) and labels the result captured from canonical through the real app (:187) without establishing that.
TASK 7950's “both animating” checker observes .motion-dot elements manufactured in the parent review shell, not motion inside either iframe (integration/full:apps/osl-hub-ui/scripts/task-7950-review-harness.mjs:91-119; review-harness/review.js:14-35). Two static pages pass. The current builder emits seven unsettled screens while the checker requires 23; the original task requires all 43. This is no coherent current acceptance contract.
TASK 8001 scans source strings and reports hard-coded green; the 7950 builder never invokes the inbox-config generator. One review page only saves local radio state and has no send path. The static generator would also put a shared write secret in browser-readable JSON (task-8001-review-controls-check.mjs:12-19; task-7950-build-review-harness.mjs:125-135; owner-reviews.js:61-93; task-8001-review-inbox-config.mjs:5-17). A public UI can therefore appear to send while nothing reaches an inbox.
Legacy uidiff.py and ui-review/build.py are hard-coded to an old worktree. uidiff.py lets stray PNGs fill missing manifest routes, resizes dimension mismatches, has no failure threshold, exits zero after differences, and labels images “Fresh” without checking (uidiff.py:22,119-144,235). Quarantine them from the live workflow.
31. UI authority: the 72-page count is structurally complete and semantically wrong
The 72 .dc.html declarations account for 41 routed pages, 30 states, and one retired page. The design checker validates route rows but requires only that the states array exists; it never validates state parent, drive recipe, identity, or reachability (check-needs.sh:168-221). It never compares manifest routes with current product routes.
Current top-level routes at snapshot canonical are onboarding, home, arrange-tiles, scrub, service, settings, mullvad, osl-chat, osl-mail, osl-mail-status, osl-notes-status, osl-servers, and signal-qa (integration/full:apps/osl-hub-ui/src/main.ts:365-366). Onboarding values and actual normalization/reachability are defined at main.ts:473-474 and first-run-spine.ts:19-65; settings choices at settings-home.ts:41-97.
The 7049 manifest claims top-level Activity, Connections, Inbox, People, and Privacy. Content functions remain, but the live render switch does not call them (main.ts:6028-6037, with content around :6264,6341,6407,6825,7648). Dead functions are not surfaces. Simultaneously, current routes are listed as failures. Counting every page exactly once proves bookkeeping, not correspondence.
D65 allows five specific undesigned routes—timed-delete-scheduler, settings/schedule, auto-whitelist-rules, friend, behaviour—to ship and removes only their visual-parity finding; behavior obligations remain (RULINGS.txt:2844-2857). They should be explicit ruled exceptions, not silently paired or repeatedly reported missing.
32. Design pages with no current reachable surface
| Design authority | Evidence/status | Fixer |
|---|---|---|
Activity + Activity Empty | Declared activity; only dead content remains, absent from current Route/render switch | Owner restore/delete; agent classifies |
Connections + Connections Empty | Same: dead function, no destination | Owner/agent |
Inbox + Inbox Empty | Same: dead destination | Owner/agent |
People + People Empty | Page route absent; current people/friend dialogs are different surfaces | Owner chooses destination vs state; agent maps |
Privacy | Declared top-level route absent; Settings/onboarding privacy are distinct | Owner/agent |
Onboarding Visibility | Current union has no visibility; old silent-visible normalizes to browser | Agent marks retired under rulings |
Onboarding Install | Declaration is retired; old install normalizes to detected, yet manifest still reports it as a live failure | Agent fixes reachable enumeration |
Website Home | Declares settings/appearance; an external website is not that desktop surface | Owner classifies web authority; agent removes fake pair |
WhatsApp Overlay Draft/Empty/Error | Product catalogue marks WhatsApp coming-soon/no overlay | Owner decides scope; agent marks not-built/excluded |
Device Storage | Declares settings/cleanup; no current setting matches exactly | Owner/agent reclassify |
Native Protect Picker states | Built around a retired forward-secrecy placement; current dialog needs exact mapping | Agent rebinds or retires |
Several pages use invented route labels for real dialogs/states: Friends Dialog Populated as friend-pictures, Owned Confirmation Verify as onboarding/recovery-check, People In Chat Dialog Populated as auto-whitelist-rules, and People Key Change/Send Mode as settings-like routes. Reclassify them as exact driven states/regions; do not count fake route names as current navigation.
33. Current product surfaces with no exact design authority
The current 7049 manifest itself names 11 unreferenced routes: arrange-tiles, mullvad, onboarding/apps, onboarding/decoy, onboarding/install, onboarding/mullvad, onboarding/silent-visible, onboarding/tutorial, osl-mail-status, osl-notes-status, and scrub (proof/7049-route-design-manifest.json:638-649). The onboarding legacy entries should be removed/normalized under D1/D4/D9/D12 rather than designed. The remaining current surfaces require a real authority or explicit scope decision.
Additional exact gaps found by joining current source to declarations:
- Top-level
scrubis not the same navigation context asScrub.dc.htmldeclared forsettings/scrub. - Reachable onboarding
cleanuphas no onboarding-cleanup page;Device Storageclaims settings cleanup. - Reachable onboarding privacy has no declaration naming that actual route.
- Settings Security exists as content/choice but lacks an exact route/drive record.
- Whitelist roster, safety-number, OSL Chat settings, and burn result/error dialogs lack exact authority or an explicit region/drive mapping.
- Home notification popover, toast, people dialog, native-protect friend dialog, scrub review dialog, update dialog, and owned-confirmation dialog need validated parent/state/drive records even where a similarly named state page exists.
Fix: agent builds one source-derived normalized surface inventory covering routes, onboarding states, settings sections, dialogs, and substates. Join it to all 72 declarations and require each live surface to have an exact page/region plus tested drive recipe; each design page must be live, ruled-retired, external-web, or explicitly not-built. Owner decides only ambiguous restore/delete/scope/design questions.
34. Five beliefs most likely to be true-looking and false now
| Believed true | What is actually known | One measurement that settles it |
|---|---|---|
| “AutoScrub is safely Find-only.” | Exact disclosure is absent; destructive/unattended inventory is non-zero; real due/provider after-state never ran. | On the exact release artifact, inventory zero destructive controls/agreements/routes/jobs/claims, observe the exact disclosure before save, then run one real Windows due event with independently labelled near-miss and true positive and prove zero delete attempts plus byte-identical provider rows after restart/re-login. |
| “The canonical artifact is the current product.” | The raw EXE is hash-correct but 23 commits behind, companion/installer-unbound, and NOT OBSERVED. | Build a complete bundle from a detached clean worktree at the frozen current release commit; require manifest source commit equality and hashes for installer/EXE/loader/sidecar, then install-launch-sign-in-update-uninstall on a clean VM. |
| “AI Covertext on Windows uses the verified local model.” | Native feature is absent; the visible path loads six templates and calls them AI. | From the exact Windows EXE, attest the compiled native-cover-writer feature and observe the real llama inference call caused by the real button; decode the result. If absent, the AI control must be absent. |
| “A green task means its full finish line passed and reached the product.” | Explicitly unfinished receipts, retired work, proofless blocks, and old task-named commits all count green. | Parse every green task into required clauses and require an invocation-bound structured verdict for each, plus exact final patch containment in canonical and exact artifact binding where behavior is claimed. Any required unchecked/unavailable clause fails. |
| “The owner is reviewing current, correctly designed screens.” | The design gate says MET with 11 failures; current/dead routes are mixed; captures lack observation provenance and include old/impossible-at-commit images. | Drive every normalized route/state/dialog in the exact packaged Windows artifact, issue per-capture artifact/source/time/image/design receipts, join to all 72 declarations, and require zero unpaired live surfaces, zero dead-page claims, and zero stale captures. |
The single priority order is therefore: remove the AutoScrub contradiction and installed deletion inventory; stop the false Windows AI claim; make the completion/design/tick gates fail on their known red cases; freeze and rebuild one reproducible complete artifact; then recapture and re-review the source-derived live UI set. Until those measurements are in place, adding more PASS prose increases confidence faster than truth.