Task-ledger and proof-artifact audit notes
Scope: OSL-AUDITS/todo/[0-9][0-9]-*.txt and OSL-AUDITS/proof/**, read on
2026-08-16. These notes are an input to the spec audit, not a claim that every
task was individually re-executed. A checked task is plan bookkeeping. It is not
runtime proof unless the cited artifact actually satisfies the task's `done
when` clause.
Ledger inventory
OSL-AUDITS/todo/status.sh reports **4,336 checked / 4,597 specified plan tasks;
261 remain open**. The useful family-level view is:
| Plan files | Capability family | Checked / total | Open |
|---|---|---|---|
| 01-05 | foundations, identities, crypto, storage, cleanup | 892 / 938 | 46 |
| 06-11 | chat/email carriers, browser, network, payment | 1,218 / 1,276 | 58 |
| 12-16 | final audit, receive path, settings/mailbox/attachments | 1,026 / 1,105 | 79 |
| 17-25 | deployment, site, policy, documentation, community | 312 / 316 | 4 |
| 26-30 | VM/account readiness, copy/content and studies | 187 / 207 | 20 |
| 31-36 | integration, security, implementation reconciliation | 231 / 233 | 2 |
| 37-43 | screen parity, runtime observations, proof pyramid | 219 / 254 | 35 |
| 44-50 | release observers, carriers, final release | 150 / 167 | 17 |
| 51-57 | completion controls and final reconciliation | 101 / 101 | 0 |
The apparent closure in files 51-57 does not establish product closure. Those
files mostly build/refine gates. Earlier product finish lines remain open, and
several checked gate tasks truthfully record that the product observation is
blocked or failed.
Build-kind totals from status.sh are also revealing: backend 1,102/1,123,
decision 64/71, deploy 22/28, test 2,556/2,704, and UI 592/671. Most ledger rows
are tests or test infrastructure, not independently observed capabilities.
Ledger integrity warnings
- 31 checked tasks have no
proof:line: 1011, 1037c, 3082, 3095, 3118,
3119, 3129-3134, 3203, 3609s, 3609v, 3690, 3691, 3693, 3694, 3752, 3765,
3773, 3928, 3955c, 4019, 4061, 4260a, 4601, 4602, 4712, 4916. Many are owner
decisions, so this is not automatically a product defect, but the tick alone
carries no reproducible evidence.
- **773 evidence Markdown files contain at least one unchecked finish-line
box. Parsing checked tasks against their cited evidence flags 756 checked
tasks** whose evidence still contains - [ ]. This is a review queue, not a
claim that all 756 tasks are false: some unchecked boxes can be historical or
secondary. It does prove that a checked ledger row cannot be accepted without
reading the evidence.
- Concrete examples of misleading green-looking ticks are TASK 3506b, 7761,
7488, and 7944 below. Each is checked, while its own evidence says the live
product finish line is blocked, failed, or not observed.
Proof-strength rubric used here
| Strength | What it establishes | What it does not establish |
|---|---|---|
| Strong | Shipping Windows artifact, exact digest, real user action/real accounts or network, independent observation of outcome, negative mutation, and restoration | Nothing outside the exercised route/account/environment |
| Moderate | Shipping Windows artifact observed through UIA/capture, but no external carrier effect or no complete admission chain | End-to-end send/receive, durability, parity, or all routes |
| Weak | Unit/component test, isolated source compile, fixture, synthetic harness, source search, screenshot, manifest, checksum, or refusal-gate self-test | That the packaged product performs the user capability |
| None/red | Planned, blocked, refused, not credited, missing observation, failed packaged run, or unchecked finish line | Any positive product claim |
File count is not proof strength. OSL-AUDITS/proof contains 2,552 files:
907 PNG, 528 JSON, 351 Markdown, 154 stdout, 147 stderr, 135 HTML, 87 text,
plus scripts/fonts/CSS/logs. Large fixture/staging directories dominate (for
example parity-audit-images 373 files, task-1548-website-candidate 306, and
task3677-fixture-1h 253). A screenshot, fixture, or generated manifest must not
be counted as a successfully exercised capability.
Hard findings
TASK 3506: two-party Discord carrier receive has never passed
Specified: OSL-AUDITS/todo/12-final-audit.txt:188-200. The finish line
requires two real Discord release processes and distinct signed-in accounts on
Windows; three fresh protected messages through the shipping carrier/network
must create exactly three receiver records with exact private words; an ordinary
message must create none; sender close, receiver network block, and removal of
the shipping receive job must each fail.
Built: code/checking infrastructure exists, but the task itself is open.
The ledger contains an explicit 2026-08-14 untick because its evidence says it
was blocked. OSL-AUDITS/evidence/3506.md records zero protected sends, zero
ordinary sends, no receiver count, no receive/decode, and no negative cases. The
isolated attempt reached WebView/create-account/recovery before PID/profile
collision killed it.
Proven: no. This is not weak proof; it is an explicit red/blocked result.
OSL-AUDITS/proof/7805-observed-red-batch-status.json classifies 3506 as
pending-observed-red.
The adjacent checked TASK 3506b is not a pass either. Its own
OSL-AUDITS/evidence/3506b.md says “Status: blocked; finish line not met. This
is not a simulated pass,” that no 3506 checker exists, and that the only native
Discord receive target is a loopback fixture explicitly ineligible under the
done-when clause. The Windows build did not complete (exit 143), so neither the
mutation nor restoration ran. **No carrier may be described as end-to-end
proven by borrowing the 3506b tick.**
TASK 7761 supplies another direct contradiction to a release claim.
OSL-AUDITS/todo/48-carriers-and-mail.txt:44 is checked, but
OSL-AUDITS/evidence/7761.md begins “Not promotable as packaged Discord
functionality.” Its only completed result is a 19-condition local refusal-gate
self-test. Both owner identities and every live send/read/reveal/timer/burn/
attachment/reference/restart leg remain unobserved; the focused hub test did not
complete. This proves a fail-closed gate was built, not Discord send/receive.
B6 Discord product-observation monitor: shipped, packaged behavior failed
Specified: TASK 7488 in OSL-AUDITS/todo/45-ship-the-observers.txt:514 and
its cited B6 observation matrix.
Built: yes in the current canonical binary. `OSL-AUDITS/proof/canonical-
artifact.json` records SHA-256
a219c5718fef9a87b6be03305741d73e16da00b6952ccc3c73e9768d02fd9271 and
launch_observation.contains_b6_monitor: true.
Proven: no; the live packaged test failed.
OSL-AUDITS/proof/7488/b6-discord-visibility.json binds that exact binary to
actual Discord PID 8240 and UIA HWND 9175624. The required sentence was “The
Discord window is no longer available.” Exact-warning count was zero while
healthy, zero after minimizing Discord, zero after hiding Discord, and zero
after both restorations. Verdict: fail; failing steps say the packaged UIA
report was absent for both hidden and unreachable cases.
OSL-AUDITS/proof/7488/b6-mutations.json is only weak component proof: an
isolated exact-source compile had two baseline tests pass, both injected source
mutations went red with exit 101, and restoration passed. The full hub target
was explicitly “not credited” because of 62 unrelated compile errors. That
isolated sensitivity test cannot override the failed shipping-binary result.
Canonical artifact gate can be MET without observing a usable app
The historical failure is preserved in
OSL-AUDITS/proof/7757/task-7757-windows-launch-receipt.json: the packaged exe
exited 78 before a top-level HWND because the strict English catalogue was
missing dialog.account_delete.confirm. OSL-AUDITS/proof/preflight-rerun.md
records the independent repeated failure: 12 runners yielded 5 REFUSED, 7
UNAVAILABLE, and 0 PASS.
The cause of the false-looking green prerequisite is visible in code.
OSL-AUDITS/check-needs.sh:117-151 implements artifact:canonical by checking
record readability, file existence, SHA-256 equality, ancestor relationship,
and an at-most-400-commit age. It performs **no launch, survival, HWND, UIA,
sign-in, or journey check**, then emits MET at line 151. Therefore “canonical
artifact MET” means only recorded bytes/ancestry are acceptable; it never meant
the executable rendered or worked.
The replacement artifact is better, but still not formally proven usable.
OSL-AUDITS/proof/canonical-artifact.json says USABILITY.state: NOT OBSERVED.
Its manual note says the process stayed alive past 30 seconds and contains B6,
but explicitly says it is not TASK 7944 OBSERVED RUNNING. TASK 7944 is checked
because its verifier/hardening exists; OSL-AUDITS/evidence/7944.md records that
the verifier still did not promote the artifact. The later identifier-isolated
Windows run failed password validation before UIA and produced no transcript.
There is moderate evidence that a related packaged executable can render:
OSL-AUDITS/proof/task-7060-windows-e2a099ac-bbdc0200/capture-provenance.json
binds SHA bbdc0200... to 13 nonblank, live interactive Windows captures. But
admission.admitted is false because the exact executable lacks a clean TASK
7059 ledger paired to complete TASK 7049 route/page pairings. This is rendering
evidence, not admitted design parity or working carrier behavior.
AI cover writer: absent from the Windows shipping binary
Specified/ruled: owner ruling D66 in OSL-AUDITS/RULINGS.txt:2859-2879.
It states that native-cover-writer requires llama-cpp-sys-2, which does not
cross-compile from WSL to windows-gnu; only the Linux library real-model proof
may be credited, and Windows behavior remains unproven.
Built: not in the canonical Windows artifact. `OSL-AUDITS/proof/canonical-
artifact.json lists features core,desktop,whatsapp-qa-shell` and explicitly
says “Built WITHOUT native-cover-writer.”
Proven: **no Windows AI cover-writer coverage is possible against this
artifact.** Any claim that the Windows product locally generates AI covers is
false. The separate owner naturalness task 3524 remains open as well.
Owner-review capture freshness: today’s refresh is not packaged proof
The exact historical sentence “74 of 76 were a week stale” could not be
reconstructed from the current post-refresh tree and should not be repeated as
independently verified. The underlying problem is confirmed:
- current
proof/owner-materials/[0-9][0-9][0-9][0-9].jsonrecords reference
85 captures (81 unique paths);
- all of those manifests now bind the 2026-08-16 canonical commit
9f916..., but Git last-touch dates show only 47/81 referenced assets changed
on 2026-08-16; 34 were last changed on August 6, 7, 9, or 10;
- commit
5a1dd67a0f80adf48d636061ede555458283034e(“refresh owner review captures
and guidance”) changed 42 files today, including 40 PNGs;
OSL-AUDITS/evidence/8003-legacy-draft.mdsays 31/69 reviewed items were
ready and 38 were not. Its “ready” material is live review-app URLs, not native
packaged-Windows evidence; missing items include provider, transient, and
native artifacts.
Thus freshness/manifests can support owner review preparation, but **do not
prove the packaged Windows product matches the design or works**. Manifest
commit binding is not capture-time provenance by itself.
Cross-cutting proof gaps
1. No B1-B7 shipping exercise is currently green. TASK 7479 is checked
because it was asked to close honestly or enumerate the gap. Its outputs are
unambiguous: OSL-AUDITS/proof/7479-passing-exercise-list.md says Count: 0;
OSL-AUDITS/proof/7479-missing-exercise-list.md says Count: 237. The tick
proves the missing-proof inventory was produced, not that any row passed.
2. Canonical parity conveyor is red. `OSL-AUDITS/proof/task-7923-canonical-
verdict/summary.json` reports failures for every listed route (38 failures),
including hundreds of structural differences on home, service, OSL Chat,
OSL Mail, OSL Servers, Signal QA, and onboarding routes, plus unpaired pages.
This is source/capture comparison evidence, not a live Windows journey, but
it directly refutes a blanket parity claim.
3. Independent product-wide review is blocked. `OSL-AUDITS/proof/6324-
product-wide-review-status.json is blocked_fail_closed`: no signed
candidate, no independent 3204/3205 inventory, and the outside review on
file used synthetic “Example Security / A. Reviewer” material.
4. TASK 7060 has useful but unadmitted Windows evidence. Thirteen live
interactive captures demonstrate that one digest rendered onboarding, but
exact parity admission is false and it exercises no real carrier effect.
5. A checked refusal gate is often the artifact, not the user outcome. This
is explicit in 3506b, 7761, 7944, and 7479. Count BUILT for the gate; count
PROVEN only when the underlying product task's positive and negative live
legs passed.
Major open capability clusters to carry into the final gap list
- Chat/social carriers (
todo/06): 12 open, including two-person Signal send
(1040), Signal attachment/story timer, X profile/timer/burn, Instagram and
Messenger profiles/timers/view-once/story burn, carrier plaintext, and privacy
transforms.
- Email/chat carrier (
todo/07): 9 open, including Outlook desktop send/read
(1287r) and several live/privacy finish lines.
- Final audit (
todo/12): 56 open, including 3506, real large-file handling,
mixed groups, clue tests, live burns/purchases, atomicity/destruction/payment
release prerequisites, and AI-cover naturalness.
- Receive path (
todo/13): 5 open, including non-Liam real send/read (4153). - Mailbox (
todo/15): 13 open, including cryptographic recipient-device
guarantees, retained-thread reachability, nonce safety, device recovery
disclosure, and owner review.
- VM/account readiness (
todo/26): 16 open, including second Signal, iCloud,
second X, Instagram, Messenger, signed-in snapshots, and window shapes.
- Screen parity (
todo/37): 7060, 7060b, and 7061 remain open. - Cover integrity (
todo/40): 18 of 40 open, including paired positive/red rows
7233, 7241-7243, and 7245-7249.
- Proof pyramid (
todo/42): 10 open, including real Windows journeys (7324)
and the top/release gates 7330-7337.
- Carrier release (
todo/48): 8 open: paired Signal, Instagram, Messenger, and
iCloud live/red tasks.
- Final release (
todo/49): 6 open, including every-carrier lifecycle (7786),
the 61/61 release claim (7791), and second iCloud/social identities.
- Completion clauses (
todo/50): 7810 and 7810b remain open.
Bottom line for the deliverable
The most important specified, believed-done, but not proven capability is **a
real protected message sent through a shipping carrier and received/decrypted by
a second real account**. TASK 3506 is the canonical test of that claim. It is
open, its last attempt sent zero messages, and its checked negative companion is
explicitly blocked. The B6 packaged monitor failure and the zero-of-237 shipping
exercise inventory reinforce the same conclusion: source, gates, and manifests
exist, but the central two-party product transaction has not been observed
working end to end.