← all audits

.02 task proof notes

Task-ledger and proof-artifact audit notes

Scope: OSL-AUDITS/todo/[0-9][0-9]-*.txt and OSL-AUDITS/proof/**, read on

2026-08-16. These notes are an input to the spec audit, not a claim that every

task was individually re-executed. A checked task is plan bookkeeping. It is not

runtime proof unless the cited artifact actually satisfies the task's `done

when` clause.

Ledger inventory

OSL-AUDITS/todo/status.sh reports **4,336 checked / 4,597 specified plan tasks;

261 remain open**. The useful family-level view is:

Plan filesCapability familyChecked / totalOpen
01-05foundations, identities, crypto, storage, cleanup892 / 93846
06-11chat/email carriers, browser, network, payment1,218 / 1,27658
12-16final audit, receive path, settings/mailbox/attachments1,026 / 1,10579
17-25deployment, site, policy, documentation, community312 / 3164
26-30VM/account readiness, copy/content and studies187 / 20720
31-36integration, security, implementation reconciliation231 / 2332
37-43screen parity, runtime observations, proof pyramid219 / 25435
44-50release observers, carriers, final release150 / 16717
51-57completion controls and final reconciliation101 / 1010

The apparent closure in files 51-57 does not establish product closure. Those

files mostly build/refine gates. Earlier product finish lines remain open, and

several checked gate tasks truthfully record that the product observation is

blocked or failed.

Build-kind totals from status.sh are also revealing: backend 1,102/1,123,

decision 64/71, deploy 22/28, test 2,556/2,704, and UI 592/671. Most ledger rows

are tests or test infrastructure, not independently observed capabilities.

Ledger integrity warnings

3119, 3129-3134, 3203, 3609s, 3609v, 3690, 3691, 3693, 3694, 3752, 3765,

3773, 3928, 3955c, 4019, 4061, 4260a, 4601, 4602, 4712, 4916. Many are owner

decisions, so this is not automatically a product defect, but the tick alone

carries no reproducible evidence.

box. Parsing checked tasks against their cited evidence flags 756 checked

tasks** whose evidence still contains - [ ]. This is a review queue, not a

claim that all 756 tasks are false: some unchecked boxes can be historical or

secondary. It does prove that a checked ledger row cannot be accepted without

reading the evidence.

7488, and 7944 below. Each is checked, while its own evidence says the live

product finish line is blocked, failed, or not observed.

Proof-strength rubric used here

StrengthWhat it establishesWhat it does not establish
StrongShipping Windows artifact, exact digest, real user action/real accounts or network, independent observation of outcome, negative mutation, and restorationNothing outside the exercised route/account/environment
ModerateShipping Windows artifact observed through UIA/capture, but no external carrier effect or no complete admission chainEnd-to-end send/receive, durability, parity, or all routes
WeakUnit/component test, isolated source compile, fixture, synthetic harness, source search, screenshot, manifest, checksum, or refusal-gate self-testThat the packaged product performs the user capability
None/redPlanned, blocked, refused, not credited, missing observation, failed packaged run, or unchecked finish lineAny positive product claim

File count is not proof strength. OSL-AUDITS/proof contains 2,552 files:

907 PNG, 528 JSON, 351 Markdown, 154 stdout, 147 stderr, 135 HTML, 87 text,

plus scripts/fonts/CSS/logs. Large fixture/staging directories dominate (for

example parity-audit-images 373 files, task-1548-website-candidate 306, and

task3677-fixture-1h 253). A screenshot, fixture, or generated manifest must not

be counted as a successfully exercised capability.

Hard findings

TASK 3506: two-party Discord carrier receive has never passed

Specified: OSL-AUDITS/todo/12-final-audit.txt:188-200. The finish line

requires two real Discord release processes and distinct signed-in accounts on

Windows; three fresh protected messages through the shipping carrier/network

must create exactly three receiver records with exact private words; an ordinary

message must create none; sender close, receiver network block, and removal of

the shipping receive job must each fail.

Built: code/checking infrastructure exists, but the task itself is open.

The ledger contains an explicit 2026-08-14 untick because its evidence says it

was blocked. OSL-AUDITS/evidence/3506.md records zero protected sends, zero

ordinary sends, no receiver count, no receive/decode, and no negative cases. The

isolated attempt reached WebView/create-account/recovery before PID/profile

collision killed it.

Proven: no. This is not weak proof; it is an explicit red/blocked result.

OSL-AUDITS/proof/7805-observed-red-batch-status.json classifies 3506 as

pending-observed-red.

The adjacent checked TASK 3506b is not a pass either. Its own

OSL-AUDITS/evidence/3506b.md says “Status: blocked; finish line not met. This

is not a simulated pass,” that no 3506 checker exists, and that the only native

Discord receive target is a loopback fixture explicitly ineligible under the

done-when clause. The Windows build did not complete (exit 143), so neither the

mutation nor restoration ran. **No carrier may be described as end-to-end

proven by borrowing the 3506b tick.**

TASK 7761 supplies another direct contradiction to a release claim.

OSL-AUDITS/todo/48-carriers-and-mail.txt:44 is checked, but

OSL-AUDITS/evidence/7761.md begins “Not promotable as packaged Discord

functionality.” Its only completed result is a 19-condition local refusal-gate

self-test. Both owner identities and every live send/read/reveal/timer/burn/

attachment/reference/restart leg remain unobserved; the focused hub test did not

complete. This proves a fail-closed gate was built, not Discord send/receive.

B6 Discord product-observation monitor: shipped, packaged behavior failed

Specified: TASK 7488 in OSL-AUDITS/todo/45-ship-the-observers.txt:514 and

its cited B6 observation matrix.

Built: yes in the current canonical binary. `OSL-AUDITS/proof/canonical-

artifact.json` records SHA-256

a219c5718fef9a87b6be03305741d73e16da00b6952ccc3c73e9768d02fd9271 and

launch_observation.contains_b6_monitor: true.

Proven: no; the live packaged test failed.

OSL-AUDITS/proof/7488/b6-discord-visibility.json binds that exact binary to

actual Discord PID 8240 and UIA HWND 9175624. The required sentence was “The

Discord window is no longer available.” Exact-warning count was zero while

healthy, zero after minimizing Discord, zero after hiding Discord, and zero

after both restorations. Verdict: fail; failing steps say the packaged UIA

report was absent for both hidden and unreachable cases.

OSL-AUDITS/proof/7488/b6-mutations.json is only weak component proof: an

isolated exact-source compile had two baseline tests pass, both injected source

mutations went red with exit 101, and restoration passed. The full hub target

was explicitly “not credited” because of 62 unrelated compile errors. That

isolated sensitivity test cannot override the failed shipping-binary result.

Canonical artifact gate can be MET without observing a usable app

The historical failure is preserved in

OSL-AUDITS/proof/7757/task-7757-windows-launch-receipt.json: the packaged exe

exited 78 before a top-level HWND because the strict English catalogue was

missing dialog.account_delete.confirm. OSL-AUDITS/proof/preflight-rerun.md

records the independent repeated failure: 12 runners yielded 5 REFUSED, 7

UNAVAILABLE, and 0 PASS.

The cause of the false-looking green prerequisite is visible in code.

OSL-AUDITS/check-needs.sh:117-151 implements artifact:canonical by checking

record readability, file existence, SHA-256 equality, ancestor relationship,

and an at-most-400-commit age. It performs **no launch, survival, HWND, UIA,

sign-in, or journey check**, then emits MET at line 151. Therefore “canonical

artifact MET” means only recorded bytes/ancestry are acceptable; it never meant

the executable rendered or worked.

The replacement artifact is better, but still not formally proven usable.

OSL-AUDITS/proof/canonical-artifact.json says USABILITY.state: NOT OBSERVED.

Its manual note says the process stayed alive past 30 seconds and contains B6,

but explicitly says it is not TASK 7944 OBSERVED RUNNING. TASK 7944 is checked

because its verifier/hardening exists; OSL-AUDITS/evidence/7944.md records that

the verifier still did not promote the artifact. The later identifier-isolated

Windows run failed password validation before UIA and produced no transcript.

There is moderate evidence that a related packaged executable can render:

OSL-AUDITS/proof/task-7060-windows-e2a099ac-bbdc0200/capture-provenance.json

binds SHA bbdc0200... to 13 nonblank, live interactive Windows captures. But

admission.admitted is false because the exact executable lacks a clean TASK

7059 ledger paired to complete TASK 7049 route/page pairings. This is rendering

evidence, not admitted design parity or working carrier behavior.

AI cover writer: absent from the Windows shipping binary

Specified/ruled: owner ruling D66 in OSL-AUDITS/RULINGS.txt:2859-2879.

It states that native-cover-writer requires llama-cpp-sys-2, which does not

cross-compile from WSL to windows-gnu; only the Linux library real-model proof

may be credited, and Windows behavior remains unproven.

Built: not in the canonical Windows artifact. `OSL-AUDITS/proof/canonical-

artifact.json lists features core,desktop,whatsapp-qa-shell` and explicitly

says “Built WITHOUT native-cover-writer.”

Proven: **no Windows AI cover-writer coverage is possible against this

artifact.** Any claim that the Windows product locally generates AI covers is

false. The separate owner naturalness task 3524 remains open as well.

Owner-review capture freshness: today’s refresh is not packaged proof

The exact historical sentence “74 of 76 were a week stale” could not be

reconstructed from the current post-refresh tree and should not be repeated as

independently verified. The underlying problem is confirmed:

85 captures (81 unique paths);

9f916..., but Git last-touch dates show only 47/81 referenced assets changed

on 2026-08-16; 34 were last changed on August 6, 7, 9, or 10;

and guidance”) changed 42 files today, including 40 PNGs;

ready and 38 were not. Its “ready” material is live review-app URLs, not native

packaged-Windows evidence; missing items include provider, transient, and

native artifacts.

Thus freshness/manifests can support owner review preparation, but **do not

prove the packaged Windows product matches the design or works**. Manifest

commit binding is not capture-time provenance by itself.

Cross-cutting proof gaps

1. No B1-B7 shipping exercise is currently green. TASK 7479 is checked

because it was asked to close honestly or enumerate the gap. Its outputs are

unambiguous: OSL-AUDITS/proof/7479-passing-exercise-list.md says Count: 0;

OSL-AUDITS/proof/7479-missing-exercise-list.md says Count: 237. The tick

proves the missing-proof inventory was produced, not that any row passed.

2. Canonical parity conveyor is red. `OSL-AUDITS/proof/task-7923-canonical-

verdict/summary.json` reports failures for every listed route (38 failures),

including hundreds of structural differences on home, service, OSL Chat,

OSL Mail, OSL Servers, Signal QA, and onboarding routes, plus unpaired pages.

This is source/capture comparison evidence, not a live Windows journey, but

it directly refutes a blanket parity claim.

3. Independent product-wide review is blocked. `OSL-AUDITS/proof/6324-

product-wide-review-status.json is blocked_fail_closed`: no signed

candidate, no independent 3204/3205 inventory, and the outside review on

file used synthetic “Example Security / A. Reviewer” material.

4. TASK 7060 has useful but unadmitted Windows evidence. Thirteen live

interactive captures demonstrate that one digest rendered onboarding, but

exact parity admission is false and it exercises no real carrier effect.

5. A checked refusal gate is often the artifact, not the user outcome. This

is explicit in 3506b, 7761, 7944, and 7479. Count BUILT for the gate; count

PROVEN only when the underlying product task's positive and negative live

legs passed.

Major open capability clusters to carry into the final gap list

(1040), Signal attachment/story timer, X profile/timer/burn, Instagram and

Messenger profiles/timers/view-once/story burn, carrier plaintext, and privacy

transforms.

(1287r) and several live/privacy finish lines.

mixed groups, clue tests, live burns/purchases, atomicity/destruction/payment

release prerequisites, and AI-cover naturalness.

guarantees, retained-thread reachability, nonce safety, device recovery

disclosure, and owner review.

second X, Instagram, Messenger, signed-in snapshots, and window shapes.

7233, 7241-7243, and 7245-7249.

and the top/release gates 7330-7337.

iCloud live/red tasks.

the 61/61 release claim (7791), and second iCloud/social identities.

Bottom line for the deliverable

The most important specified, believed-done, but not proven capability is **a

real protected message sent through a shipping carrier and received/decrypted by

a second real account**. TASK 3506 is the canonical test of that claim. It is

open, its last attempt sent zero messages, and its checked negative companion is

explicitly blocked. The B6 packaged monitor failure and the zero-of-237 shipping

exercise inventory reinforce the same conclusion: source, gates, and manifests

exist, but the central two-party product transaction has not been observed

working end to end.