← all audits

What only you can do

Your remaining work, with the steps, ordered by what it unblocks.

What Liam must do

The fleet is not short of lanes. It is short of work it is allowed to do without Liam. There are 261 open tasks. Of the 105 measured gate-ready tasks, only 25 are agent-runnable, while 80 belong to Liam: 56 screen reviews, 14 real tests, 7 decisions, and 3 deploys. The pack ladder is already approved and ticked, so it is outside those 80. TASK 4402 is inside the 80 but needs no further owner answer. That leaves 79 items with physical actions below.

Read this before starting

The eight highest-leverage actions

1. TASK 1619 — let a real friend install OSL with absolutely no help

This proves that a new person can install and set up the release without Liam coaching them, and it unblocks six tasks.

Why only Liam: it needs Liam's real friend, a fresh real Windows computer, and an honest record of whether Liam intervened.

Do this:

1. Ask an agent to identify the exact release installer under test, record its hash, and put that exact file where the friend can obtain it. Use the unsigned installer. The task's words say “exact signed artifact,” but the 2026-07-31 owner ruling says ship unsigned and accept the warning; owner rulings outrank plan text.

2. Before the attempt, confirm the friend's Windows machine has never had OSL installed and contains no old OSL state.

3. Start a continuous observation record. From this point on, do not prompt, answer, point, touch the machine, type, scan a code, solve a CAPTCHA, help with two-factor authentication, or give even one hint.

4. Send the installer to the friend and stay silent while they get past whichever Windows unsigned-publisher or SmartScreen warning appears. Because the installer is unsigned, a Windows warning is expected, and the friend must work out how to continue unaided.

5. Stay silent while the friend installs OSL, launches it, creates a password, records or handles recovery, creates an identity, and completes first setup.

6. When the friend says they are finished, have the independent checker read the installed state. It must show a successful setup and exactly one new, non-empty identity.

7. Record either PASS or the exact first point of failure. A failed attempt is useful evidence but does not close the task.

It fails if: Liam gives one hint or performs one action; the machine is not fresh; the file is not the exact unsigned release artifact selected for this run; the friend stops early; any production setup action is disconnected; or the final independent read does not show exactly one new non-empty identity and successful setup.

Time and needs: 30–60 minutes. Needs one real friend, that friend's fresh Windows PC, the exact unsigned release installer and hash, and an independent final-state checker.

2. TASK 1287r — send and read one protected message through real Outlook desktop

This proves the installed Outlook desktop route with two real accounts and unblocks four tasks.

Why only Liam: only Liam may enter the two real mailbox passwords and handle their CAPTCHA or two-factor prompts.

Do this:

1. Have an agent open two isolated Outlook desktop profiles on MON-2 and drive each profile to its real sign-in page.

2. Type the password for the first mailbox yourself and complete any human check. Do the same for the second mailbox. Type nothing else; return control to the agent after sign-in.

3. Have the agent record that mailbox one contains zero messages marked OSL-OUTLOOK-DESKTOP-1287R.

4. Have the agent use OSL to create and send one protected message with that marker from mailbox one.

5. In mailbox two, open the arriving message and have the agent independently read back the exact private words.

6. Let the agent deliberately break the OSL Send path and try again. Confirm OSL refuses safely, sends nothing in the clear, and the marked-message counts in both mailboxes do not change.

It fails if: either profile is a fixture or signed out; the starting count is not zero; the message does not arrive exactly once; the private words do not match; the broken Send path sends anything; or either count changes during the deliberate failure.

Time and needs: 15–25 minutes. Needs two real mailboxes, their credentials and second factors, installed Outlook desktop, and two isolated Outlook/OSL profiles.

3. TASK 7793 — make the missing Signal, Instagram, and Messenger test identities available

This supplies the isolated social accounts the fleet cannot create or sign into for Liam and unblocks three tasks.

Why only Liam: account registration, passwords, QR linking, CAPTCHA, and two-factor authentication must be performed by the account owner, not an agent.

Do this:

1. Have an agent open the second isolated Signal profile and the isolated Instagram and Messenger browser profiles on MON-2.

2. For Signal, use the phone holding the second registered Signal identity to scan Link a New Device. Do not link the same Signal identity twice.

3. For Instagram, sign the second test identity into its isolated profile and complete any code, CAPTCHA, or “was this you?” check yourself.

4. For Messenger, sign a separate Facebook/Messenger test identity into its isolated profile and complete the human checks yourself. An Instagram-only session does not count as Messenger.

5. Close and reopen each profile. Have the agent read the account control and confirm the expected distinct identity is still signed in.

6. Record READY or WAIT separately for Signal, Instagram, and Messenger. For every ready service, record only the safe profile label and account label; never record a password, token, code, cookie, or session ID. For every wait, record the concrete missing thing and when it can be ready.

It fails if: a directory exists but the account is still at a QR or login page; two profiles resolve to the same identity; Messenger is inferred from an Instagram session; an agent handles a password/CAPTCHA/2FA step; or credentials appear in the record.

Time and needs: 10–30 minutes per service. Needs a second registered Signal identity and phone, a second Instagram identity, a Facebook/Messenger test identity, and access to their recovery mailboxes or second factors.

4. TASK 3227 — perform the 40-message blind cover review

This measures whether Liam can spot hidden messages from provider-side cover data without seeing the answer key and unblocks three tasks.

Why only Liam: this is an independent human judgment; an agent knows or can derive the generated answer key.

Do this:

1. Ask an agent for the numbered set of 40 saved messages with the answer key hidden. Do not open the key or any generation log.

2. Read all 40 in order. Beside every row, write HIDDEN or ORDINARY based only on what the provider shows.

3. For every guess, write any clue you actually used: repeated wording, odd word choice, similar length, clustered time, unnatural phrasing, or NO CLUE.

4. Confirm all 40 rows have a guess and a clue entry before the key is revealed.

5. Give the completed blind sheet back to the agent. Let the agent reveal the precommitted answer key, calculate the exact correct count, and append the key without changing Liam's guesses.

6. Confirm the saved record starts with the exact sentence research only - this does not prove cover messages are unspottable.

It fails if: Liam sees the key before finishing; one of the 40 rows is absent or unanswered; clues are not recorded; the key was not fixed before the run; the correct count is missing or wrong; or the required research-only sentence is absent.

Time and needs: 25–45 minutes. Needs the 40-message blind packet, a separately sealed/precommitted answer key, and somewhere to write one guess and clue per row.

5. TASK 1040 — complete the real two-person Signal send

This proves a protected Signal message can travel between two distinct real accounts and unblocks three tasks.

Why only Liam: Liam must perform any QR, password, CAPTCHA, or two-factor step for the real accounts; the agent runs everything else.

Do this:

1. Have the agent open two isolated Signal/OSL instances on MON-2 and show their distinct profile paths and account labels.

2. If either instance asks for linking or authentication, perform only that QR/password/CAPTCHA/2FA step, then hand control back.

3. Have the agent confirm the receiver holds zero messages matching the new random marker.

4. Let the agent send one protected message from account one and open it on account two.

5. Confirm the receiver now holds exactly one matching message and that its private words read back exactly.

6. Let the agent save sender and receiver captures, both Signal versions, both processes, both account labels, and both OSL data paths, then run the missing-message and blank-capture negative controls.

It fails if: the accounts, processes, or OSL data paths are not distinct; the count is not zero then one; the words differ; either capture is blank/flat/transparent or lacks its controls and bounds; versions are missing; or the unchanged checker does not fail when the arriving message or a valid capture is removed.

Time and needs: 15–25 minutes after both identities are linked. Needs two real Signal identities, the phone(s) used to link them, and two isolated Signal/OSL profiles.

6. TASK 3201 — make exactly one real production card purchase

This proves the live hosted card checkout, live webhook, and production voucher issuance in one paid action and unblocks two tasks.

Why only Liam: only Liam can authorize a real charge on his card and inspect the live merchant result.

Do this:

1. Have an agent stage the public production purchase page on MON-2 and show the intended live product and price. Confirm the URL is production, not QA or test mode.

2. Click the purchase control and confirm the processor-hosted page shows that same product and price.

3. Enter the real card details yourself and authorize one charge.

4. Stop. Do not retry or make a second charge, even if evidence collection is awkward.

5. Have the agent record the resulting live Checkout Session and live webhook event, including livemode=true and the exact product/price, without exposing card details.

6. Have the agent independently confirm production delivery minted exactly one new voucher from that event. Do not redeem it or rerun decline, dispute, privacy-wire, or duplicate-redemption tests.

It fails if: the page, key, session, or webhook is test/QA; the charged product or price is wrong; the webhook is not verified; no voucher or more than one voucher is minted; evidence comes from test-mode coverage; or more than one real charge is made.

Time and needs: 10–20 minutes. Needs one real card, access to its authentication step, the public production URL, the intended live price, live processor access, and voucher-delivery observation.

7. TASK 0569 — enable recurring timer deletion in production

This turns on the reviewed live timer worker and proves the scheduler keeps owning it across restarts and missed intervals, unblocking two tasks.

Why only Liam: this changes the live service and requires owner authorization for production configuration and restarts.

Do this:

1. Have an agent stage the reviewed deployment change, live scheduler view, storage observer, lease/run-ID recorder, and freshness alarm on MON-2. Confirm the target is production.

2. Authenticate and approve enabling the timer worker. Do not invoke the deletion job manually.

3. Before the first scheduled interval, have the agent seed one fresh non-empty expired target, one unexpired sentinel, and one no-expiry sentinel. Observe the scheduler remove only the expired target at its first eligible interval.

4. Repeat with a fresh target for at least four independently observed scheduled intervals total. Record the authoritative lease and run ID plus an independent storage read after each interval.

5. During those intervals, approve one worker restart and one service restart. Confirm recurring execution resumes without a person calling the job.

6. Stop the worker, make one target become eligible while it is stopped, then restart it. Confirm the scheduler removes that target within the documented catch-up interval and preserves both sentinels byte-for-byte.

7. Disable or starve one scheduled run long enough for the freshness alarm to become visibly unhealthy. Restore scheduling and confirm the alarm becomes healthy only after a new independently observed successful run.

It fails if: evidence is only an enabled flag or a manual/direct job call; fewer than four intervals are observed; a target or sentinel is missing; either restart or the catch-up case is absent; a sentinel changes; lease/run IDs or independent storage reads are missing; or the freshness alarm does not go unhealthy and recover for the right reasons.

Time and needs: 45–120 minutes plus the configured interval length. Needs production deploy access, authority to restart worker and service, live scheduler/lease access, independent storage reads, test records, and the freshness alarm.

8. TASK 7792 — sign in the second isolated iCloud mailbox

This provides the second real iCloud identity needed for the two-mailbox test and unblocks one task.

Why only Liam: only Liam may enter the mailbox credentials and complete Apple's human or two-factor checks.

Do this:

1. Have an agent open the two isolated iCloud mail profiles on MON-2, with the already-ready profile and the second profile clearly labelled.

2. Sign the second mailbox in yourself and complete any Apple approval or two-factor prompt.

3. Close and reopen both profiles.

4. Have the agent read the account control in each and confirm two different mailboxes remain signed in.

5. Record READY NOW and the two safe profile labels, or record WAIT with the missing requirement and ready date. Do not record addresses if they are considered secret, and never record credentials, codes, cookies, or tokens.

It fails if: there is only one signed-in mailbox; the same mailbox is in both profiles; a login page is counted as ready; reopen loses the session; or the record contains secrets rather than safe profile labels.

Time and needs: 10–20 minutes. Needs the second iCloud identity, its password, trusted device or phone number, and two isolated browser/mail profiles.

The other real-world tests

9. TASK 1051 — send one real protected Signal attachment

This checks that one protected file sent from one real Signal account arrives intact on the other.

Why only Liam: he owns the real accounts and must handle any authentication prompt.

Steps: 1. Have the agent open the two signed-in Signal profiles on MON-2 and record zero matching saved files. 2. If prompted, Liam completes only authentication. 3. Let the agent attach and send one uniquely named protected file from account one. 4. On account two, click Open or Save and independently compare the file name, byte size, and contents. 5. Save both account-labelled captures and run the missing-file and blank-capture negative controls.

Fail: anything other than a zero-to-one saved-file count, exact byte match, distinct accounts, usable captures, and negative controls that go red. Time/needs: 15–25 minutes; two linked Signal accounts and one harmless test file.

10. TASK 1058 — watch a protected Signal story expire on both profiles

This checks one real protected story appears in both Signal copies and disappears at its earlier ending.

Why only Liam: real signed-in Signal identities and any linking prompts belong to him.

Steps: 1. Open both real profiles on MON-2. 2. Let the agent publish one random marked protected story to both and record exact text and count 1 in each. 3. Wait once for the earlier configured ending; do not delete it manually. 4. Refresh both and record count 0 and no marked text. 5. Run the disabled-receiving-job negative control.

Fail: either starting count is not one, either ending count is not zero, text survives, the story is manually removed, or disabling the real receive job does not fail the check. Time/needs: 10–20 minutes hands-on plus timer wait; two Signal profiles.

11. TASK 1117 — run X timer, view-once, post-burn, and reply-burn on two profiles

This checks all four destructive X behaviours against real carrier state while preserving an unrelated marker.

Why only Liam: the two real X identities, login challenges, and destructive approval belong to him.

Steps: 1. Sign two isolated X browser profiles in on MON-2. 2. Let a separate installed sender deliver the five declared markers: timer, view once, post burn, reply burn, and keep. 3. Before each action, confirm both-profile target counts and the preserve set. 4. Exercise the shipping timer, open view-once twice, burn the post, and burn the reply. 5. Refresh both profiles and confirm each target is gone or unreadable exactly as declared while KEEP is byte-identical. 6. Run one mutation that retains a target, one that removes KEEP, and one with the X receiving job stubbed to do nothing; each unchanged check must go red, then restoration must pass.

Fail: fixture/seeded messages, a missing marker/profile/action, a second view that succeeds, an undeleted target, deleted KEEP, blank capture, or any retain/delete/disabled-receiver negative control staying green. Time/needs: 35–60 minutes plus timer wait; two distinct X identities and isolated profiles.

12. TASK 1146 — run Instagram timer, view-once, and every burn side

This checks the full Instagram destructive set against two real signed-in profiles.

Why only Liam: account login, human checks, and destructive actions require the account owner.

Steps: 1. Sign in two isolated Instagram profiles on MON-2. 2. Let a separate installed sender deliver the declared timer, view-once, burn-yours, burn-theirs, burn-both, and KEEP markers. 3. Record nonzero before-counts on both profiles. 4. Exercise timer, open view-once once and then attempt an identical second open, and run each burn side separately. 5. Refresh both profiles and compare exact declared target and preserve sets. 6. Run retain-target, delete-KEEP, and stubbed-Instagram-receiving-job negative controls; each unchanged check must go red, then restoration must pass.

Fail: a missing side/marker/profile, seeded or fixture state, target survival, KEEP deletion/change, second view success, blank capture, or any retain/delete/disabled-receiver negative control staying green. Time/needs: 40–70 minutes plus timer wait; two Instagram identities and isolated profiles.

13. TASK 1160 — burn one real Instagram story without touching the control post

This proves a selected-audience story is removed while an unrelated post for another audience remains.

Why only Liam: he controls the real Instagram accounts, audiences, and any login challenges.

Steps: 1. Have a separate installed sender publish INSTAGRAM-STORY-1160 to the selected audience and INSTAGRAM-KEEP-1160 to another. 2. On MON-2, read both through the real carrier and record count 1 for each. 3. Use the shipping burn action once on STORY. 4. Refresh and record STORY 0, KEEP 1, with KEEP byte-identical. 5. Run one copy that wrongly retains STORY and another that deletes KEEP.

Fail: either marker was seeded locally, either is missing before burn, STORY survives, KEEP changes/disappears, or either bad copy stays green. Time/needs: 20–35 minutes; two Instagram audiences and real accounts.

14. TASK 3202 — make one real Bitcoin purchase and one real Monero purchase

This proves both live crypto payment paths from mainnet invoice to production voucher.

Why only Liam: only Liam can authorize spending real BTC and XMR.

Steps: 1. On the public production checkout, choose the intended product and create one live BTC invoice. 2. Send exactly the requested BTC amount to that mainnet address. 3. Record the invoice and chain transaction, then wait for the production watcher to confirm and mint exactly one voucher. 4. Repeat once for XMR using a new live mainnet invoice/address. 5. Record both production settlements without redeeming the vouchers or rerunning testnet/stagenet tests.

Fail: testnet/stagenet, mock watchers, wrong product/address/amount, missing live transaction or invoice, no voucher, extra voucher, or any extra payment. Time/needs: 15–25 minutes hands-on plus chain confirmation; funded BTC and XMR wallets and live merchant access.

15. TASK 3624 — judge English cover text inside every supported conversation language

This checks whether English-bank cover text looks conspicuous in non-English conversations.

Why only Liam: the finish line is a reasoned language-by-language human judgment.

Steps: 1. Have the agent stage fresh real-carrier conversations for every supported language on MON-2, beginning with Japanese private text in Japanese. 2. Open each same-run Windows capture and read the English cover in context. 3. Mark NOT CONSPICUOUS or CONSPICUOUS and give one reason for every language. 4. Confirm a deliberately literal English canary in Japanese is judged conspicuous. 5. Return the completed set for the missing-language and removed-handler negative checks.

Fail: one shipping language lacks a judgment, any shipping language is conspicuous, the canary is not, the carrier/capture is not live and nonblank, or removing language handling does not fail. Time/needs: 20–40 minutes; complete supported-language packet.

16. TASK 3718 — send one protected Discord message to a mixed four-person group

This proves only the OSL member can recover private words while two real non-OSL members see the cover.

Why only Liam: the real Discord accounts and fourth browser identity require his access.

Steps: 1. Have the agent open Stable and PTB as sender and OSL member, Canary as non-OSL member one, and a separate signed-in browser profile as non-OSL member two on MON-2. 2. Liam handles any login/human checks. 3. Confirm all three receiving accounts have zero matching messages. 4. Send one marked protected group message. 5. Inspect all three receivers: each must have exactly one matching message; only the OSL member may read the private words. 6. Record private-word count 0 for both no-OSL members.

Fail: fewer than four distinct accounts, missing browser fourth seat, counts other than zero-to-one, either no-OSL account sees private words, or the OSL member cannot. Time/needs: 20–35 minutes; four Discord test accounts across Stable, PTB, Canary, and browser.

17. TASK 3721 — inspect visible clues on every shipping chat carrier

This checks every carrier in the versioned manifest from the viewpoint of an account without OSL.

Why only Liam: real carrier accounts and human login checks belong to him.

Steps: 1. Have the agent load the current shipping-carrier manifest and open a no-OSL observer account for each carrier. 2. Liam completes any login/2FA/CAPTCHA. 3. Send one fresh marked cover through each carrier and confirm its count changes from zero to one. 4. On the no-OSL account, open the visible message and details view while the agent saves raw emitted/stored bytes and the carrier message ID. 5. Confirm each row reports zero OSL names, zero private-sender clues, and the exact visible-detail count. 6. Run mutations inserting an OSL name and private identity into the raw bytes.

Fail: manifest rows are missing/extra/hard-coded, an account or byte capture is absent, IDs do not join, a clue exists, or either raw-byte mutation stays green because the rendered view hid it. Time/needs: 10–15 minutes per shipping carrier; all real carrier identities.

18. TASK 4153 — have a non-Liam person send and read a private message

This proves another person can install OSL, add Liam, send, and read a reply using only the app.

Why only Liam: he must supply the other real person and be the real friend at the far end.

Steps: 1. Use the successful TASK 1619 friend's installed copy, or give the current installer to another person who is not Liam if 1619 cannot continue. 2. Do not explain the app; observe them add Liam as a friend and send a private message. 3. Reply once from Liam's account. 4. Have them open the reply and describe in their own words what they saw. 5. Record opened-private-message count zero before and exactly one after, plus the number of explanations Liam gave. 6. After the successful human run, let the agent stub the receiving job to do nothing and confirm the unchanged check goes red, then restore it.

Fail: the person is coached beyond what the app says, the count is not zero-to-one, they cannot send/read, their own-words reaction is absent, or disabling the real receiving job does not fail. Time/needs: 25–45 minutes; another person, another Windows machine, and Liam's OSL identity.

The remaining owner decisions

19. TASK 0475 — approve or reject the independently measured live revision

This decides whether the exact reviewed image and configuration are what current users reach.

Why only Liam: only the owner can accept the production target as the intended current-user server.

Steps: 1. Have the agent put 0442a's fresh outside-trust-path measurement, signed reviewed manifest, image digest, configuration digest, target name, and old-binary/self-report substitution results on MON-2. 2. Compare the target to the current-user server and both digests byte-for-byte to the manifest. 3. Confirm every hostile substitution result is green. 4. Write dated PASS, or REJECTED with one numbered repair.

Fail: relying on the server's own revision string, a stale/replayed measurement, any digest/target mismatch, unreachable target, or missing substitution result. Time/needs: 5–10 minutes; prepared 0442a packet.

20. TASK 3952 — choose a remedy only for carriers that fail the fresh cover rerun

This chooses “change the cover” or “turn the carrier off” only where a genuine live rerun proves carrier damage.

Why only Liam: product behavior for a failing carrier is an owner choice.

Steps: 1. Do not answer until the agent supplies fresh 3950/3951 results for every currently signed-in carrier, each with live sent cover, independent read-back, and ordinary control. 2. If failures are zero, record NO DECISION NEEDED. 3. For each actual failing carrier, write exactly one choice: CHANGE COVER or TURN OFF. 4. If changing cover, wait for a new real send that returns the exact pointer. If turning off, read the promised plain refusal and exercise it in the release build.

Fail: deciding from a zero-app/fixture/replay run, omitting a carrier, leaving a failure undecided, or accepting an unexercised remedy. Time/needs: 5 minutes plus reruns; complete fresh carrier packet.

21. TASK 4763 — judge the Anyone consent gate

This decides whether the three-sentence gate honestly explains the retroactive and permanent trade-off.

Why only Liam: this wording carries an owner decision made against advice and needs his explicit acceptance.

Steps: 1. In the shipping Windows build on MON-2, open Settings and click Anyone. 2. Read all three sentences at shipped size. 3. Confirm Turn on starts disabled, tick the consent box, and confirm it becomes enabled. 4. Click Cancel and verify the setting is unchanged. 5. Reopen, tick, confirm, and verify the consent behavior occurs. 6. Write dated PASS or REJECTED with a repair.

Fail: unreadable sentence, unacceptable retroactive/permanent wording, enabled-before-tick, Cancel changes the setting, Confirm is inert/bypassed, missing bounded capture/UIA record, blank capture, or design-rule breach. Time/needs: 8–12 minutes; shipping Windows gate and prepared recorder.

22. TASK 7508 — safely retire credential-bearing scratch residue

This decides what secret scratch entries are still needed, migrates only those, and removes the volatile copies without leaking values.

Why only Liam: only the owner may decide which credentials remain needed and authorize their destruction.

Steps: 1. With all recording/log output disabled for values, open the volatile credential-named CSV locally; do not copy or print it. 2. Mark each entry needed or not needed. 3. Move every needed entry into the approved credential store and verify it opens its required test identity. 4. Record only the number migrated and destination class, never a value. 5. After verification, securely retire the scratch CSV and browser-session-copy residue. 6. Have the agent verify the files are absent, scan plan/evidence/logs for leaked values without displaying them, and reopen required identity surfaces.

Fail: deleting before review, printing/copying a value, leaving residue, losing a required identity, or skipping store verification, absence check, or leak scan. Time/needs: 20–45 minutes; approved credential store and local access to the named residue.

The other production changes

23. TASK 0570 — enable view-once enforcement live

This turns on the reviewed production single-use claim and proves only one simultaneous claimant receives content.

Why only Liam: it changes a live production setting.

Steps: 1. Have the agent stage the exact reviewed setting and live claim observer on MON-2. 2. Authenticate and change the setting from disabled to enabled. 3. Create one non-empty live record. 4. Launch two byte-identical claims simultaneously. 5. Confirm exactly one returns content and one is refused. 6. Attempt a later replay and confirm it is refused. 7. Let the agent run separate red proofs that change the setting back, remove the atomic guard, and starve the live claim service; each unchanged check must fail, then restoration must pass.

Fail: an unchanged/default flag, sequential rather than simultaneous proof, two winners, zero winners, replay success, fixture service, missing live record, or any setting/atomic-guard/starved-service red proof staying green. Time/needs: 15–25 minutes; production configuration and live claim access.

24. TASK 7409 — preserve the exact canonical integration history off-machine

This chooses the durable remote/ref for the exact integrated history and publishes it without rewriting anything.

Why only Liam: updating or choosing the shared canonical remote is an owner-only operation.

Steps: 1. Have the agent fetch without changing local history and show the exact final integration/full SHA, origin/integration/full SHA, ancestry, and divergence. 2. Decide whether origin/integration/full is canonical or name a different protected private remote/ref. 3. If origin is still an ancestor, approve only a normal fast-forward push. If the remote has any new commit, stop and create a reconciliation task. 4. If using an alternate ref, approve creation of that new durable off-machine ref. 5. Never approve force-push, rewrite, or deletion. 6. From a clean clone/fetch, verify the chosen ref resolves byte-for-byte to the final 7407 SHA and all source-to-canonical mappings are reachable; record the decision and ref in proof/7409-canonical-remote.md.

Fail: remote-only divergence is ignored, tip differs, history is only patch-equivalent, copy is on the same disk, choice is unrecorded, reachability is incomplete, or any force/rewrite/delete occurs. Time/needs: 15–30 minutes; authenticated remote access and a clean verification location.

No action: two owner choices are already settled

Do not ask Liam about the pack ladder again. D57 is approved and TASK 7794 is ticked.

Do not ask Liam for the ship-or-not list again. TASK 4402's seven open groups B1–B7 were answered YES on 2026-08-14 and confirmed twice; the 411 forced-no rows and 48 not-a-decision rows are unchanged. TASK 4402 remains open for a different reason: every yes row must have its own shipping-build exercise from TASK 4401 and zero prohibited-policy failures. The current evidence says shipping-build=not-run, so there are zero such exercises. Its finish line explicitly says, “Decision text cannot count as behavioural evidence or override an unsafe result.” Tasks 7450 onward, not another Liam decision, must supply that evidence.

The live screen journeys

25. TASK 1510 — confirm the public production purchase route without buying again

This checks that Download leads to the same live hosted product and price already proved by the paid runs. Why only Liam: he owns the production purchase judgment. Steps: 1. After TASKS 3201 and 3202, open a clean external browser on MON-2 at the public production Download page. 2. Click the purchase control. 3. Confirm the hosted checkout shows the exact live product/price in the saved paid record and is not a QA/test route. 4. Save the route and destination capture, then leave checkout without paying. Fail: local/QA/test page, wrong product/price, static screenshot instead of a click, or any additional charge. Time/needs: 5–10 minutes; public URL and 3201/3202 records.

26. TASK 1609 — walk from the clean installer to Welcome

This checks whether the installer and first screen give a new person a clear next step. Why only Liam: the finish line requires his product judgment. Steps: 1. Have the exact unsigned release installer staged on a clean Windows VM on MON-2. 2. Run it and use only the visible installer controls. 3. Continue through first launch until Welcome appears. 4. Read the installer and Welcome as a new user and write PASS only if the next action is clear; otherwise name the confusing point and create a repair. Fail: wrong artifact, non-clean machine, fixture/render/Linux evidence, missing before/after OS/app state, inert action, or unclear next step. Time/needs: 10–20 minutes; clean VM and exact release installer.

27. TASK 0595 — judge the Free/Pro view-once split

This checks that creation is Pro-only while a Free recipient can still open a received view-once message. Why only Liam: “is the split obvious?” is an owner judgment. Steps: 1. On MON-2, open the create control as a signed-in Free account and record disabled plus its reason. 2. Open it as Pro, create one marked view-once message, and retain its provider ID. 3. As a separate Free account, open that exact message once. 4. Check all three states individually and write PASS or REJECTED. Fail: fixture/signed-out run, inert control, missing provider ID, blank capture, or any state not reviewed. Time/needs: 10–15 minutes; signed-in Free and Pro accounts plus a separate Free recipient.

Prepared visual-review sitting

For TASKS 28–71 below, have the agent prepare one packet per task before Liam sits down. Each packet must use the current shipping Windows build, the current canonical commit, a live nonblank capture and screen tree, the task's control-activation transcript, and—where the plan requires it—the 7055 result showing zero structural differences and pixel difference below 1%. Liam opens the packet on MON-2, checks the named criterion, physically activates each named control if the packet calls for it, and writes PASS or REJECTED. Any rejection gets one numbered repair and a repeat. A stale, flat, blank, legacy, fixture, or mismatched capture; absent transcript; inert control; unchecked criterion; or red/missing machine comparison fails every item in this sitting.

28. TASK 0028 — review the fixed-screen Linux Welcome

This checks the Welcome layout and every route from its buttons. Why only Liam: final readability and spacing are visual judgments. Steps: 1. Open the dated Welcome packet. 2. Read every word and inspect clipping and spacing. 3. Match every visible required button to the same-run transcript showing its named shipping route/state. 4. Mark the checklist and verdict. Fail: any common review failure, clipped word, hidden/inert button, wrong route, or empty checklist. Time/needs: 4–6 minutes; prepared 0028 packet.

29. TASK 0381 — review onboarding batch one

This judges the first seven onboarding pages for truth, completeness, and design conformance. Why only Liam: final product judgment cannot be automated. Steps: 1. Open all seven dated live packets. 2. For each page, read the title and locate every named control. 3. Check dark ground, named accent palette only, 0–3 px radii, outline buttons, no shadows/gradients, and uppercase Consolas status. 4. Write a verdict beside every page; repeat repairs until all seven pass. Fail: any common review failure, omitted page, missing control/title, filled cyan button, shadow/gradient, or one page left unchecked. Time/needs: 15–25 minutes; seven current packets.

30. TASK 0382 — review manifest-derived onboarding batch two

This reviews every retained 0360–0366 onboarding/recovery page and excludes only manifest-deleted pages. Why only Liam: readability, truth, and finish need his judgment. Steps: 1. Open the current 7049 manifest and derive the batch from it. 2. For every retained page, compare its route, 7055 verdict, live capture, and control transcript. 3. For every excluded page, confirm DELETED and its ruling. 4. Apply the onboarding design rules and write one verdict per retained page. Fail: any common review failure, hard-coded count, omitted retained/deleted entry, absent ruling, or design breach. Time/needs: 15–25 minutes; current manifest-derived packet.

31. TASK 0383 — review onboarding batch three

This judges the third seven onboarding pages for truth, completeness, and design conformance. Why only Liam: final product judgment is his. Steps: 1. Open all seven current packets. 2. Read every title and locate every named control. 3. Check the required dark ground, palette, small radii, outline buttons, no shadows/gradients, and uppercase Consolas status. 4. Record seven verdicts and repeat rejected pages after repair. Fail: any common review failure, omitted page, unreadable title, missing/inert control, or design breach. Time/needs: 15–25 minutes; seven current packets.

32. TASK 0384 — review manifest-derived onboarding batch four

This reviews every retained 0374–0380 onboarding/recovery page and the rulings for deletions. Why only Liam: readability, truth, and finish need his judgment. Steps: 1. Derive the batch from the current 7049 manifest. 2. Inspect route, 7055 verdict, live capture, and transcript for every retained entry. 3. Confirm every exclusion is recorded DELETED with its ruling. 4. Apply the onboarding design rules and record one verdict per retained entry. Fail: any common review failure, hard-coded count, omitted entry/ruling, or nonconforming page. Time/needs: 15–25 minutes; current manifest-derived packet.

33. TASK 0056 — review the Free attachment screen

This checks that the Free limit, exact file size, and upgrade offer are readable and true. Why only Liam: he must judge clarity and honesty. Steps: 1. Open the dated Free attachment packet. 2. Read the displayed limit and exact size against the evidence. 3. Locate and activate the upgrade action. 4. Record the three checks and verdict. Fail: any common review failure, missing/wrong limit or size, hidden/inert upgrade, or unsupported claim. Time/needs: 3–5 minutes; 0056 packet.

34. TASK 0057 — review the Pro attachment screen

This checks that the Pro limit and no-upgrade message are readable and true. Why only Liam: he must judge clarity and honesty. Steps: 1. Open the dated Pro packet. 2. Compare its displayed limit with evidence. 3. Confirm the no-upgrade message is visible and there is no upgrade action. 4. Record both checks and verdict. Fail: any common review failure, wrong/missing limit or message, or an upgrade action. Time/needs: 3–5 minutes; 0057 packet.

35. TASK 0717 — judge Settings home

This checks that Settings is calm, scannable, and does not make safe use feel like setup. Why only Liam: this is a product-tone judgment. Steps: 1. Open the Settings home packet. 2. Scan it once without instructions, then read every choice. 3. Activate the named controls and confirm their routes. 4. Mark calm, clear, easy to scan, and safe-use-not-setup separately. Fail: any common review failure or one of those four judgments not passing. Time/needs: 3–5 minutes; 0717 packet.

36. TASK 0721 — judge privacy-level settings

This checks that all three privacy choices state their consequences without pressure. Why only Liam: neutrality and plainness need his judgment. Steps: 1. Open the privacy-level packet. 2. Read all three choices and consequences. 3. Activate each choice and inspect the resulting state. 4. Mark visibility, plain consequences, and no pressure. Fail: any common review failure, missing choice/consequence, pressure language, or inert choice. Time/needs: 4–6 minutes; 0721 packet.

37. TASK 0725 — judge Notifications

This checks that people can tell what will interrupt them and what every app tick controls. Why only Liam: this is a comprehension judgment. Steps: 1. Open the Notifications packet. 2. Read the interruption rules and every app tick label. 3. Toggle each named control and check its persisted result. 4. Record whether both questions are unambiguous. Fail: any common review failure, unclear interruption, ambiguous app tick, or inert toggle. Time/needs: 4–6 minutes; 0725 packet.

38. TASK 0741 — judge auto-whitelist rules

This checks that direct messages, groups, servers, channels, threads, email, and posts are unambiguous. Why only Liam: scope clarity is an owner judgment. Steps: 1. Open the auto-whitelist packet. 2. Read and check each of the seven named scopes individually. 3. Activate each scope and inspect its persisted effect. 4. Record a verdict only after all seven are checked. Fail: any common review failure, omitted/ambiguous scope, or inert control. Time/needs: 5–7 minutes; 0741 packet.

39. TASK 0749 — judge the verification warning

This checks that the safety trade-off is clear without frightening technical words. Why only Liam: tone and honesty belong to the owner. Steps: 1. Open the warning at shipped size. 2. Read it once as a new user. 3. State in plain words what trade-off it communicated. 4. Mark clear trade-off and nontechnical wording separately. Fail: any common review failure, unclear consequence, or scary/technical wording. Time/needs: 3–5 minutes; 0749 packet.

40. TASK 0757 — judge message defaults

This checks that timers, burning, view once, and writing choice cannot be confused. Why only Liam: he must judge whether the choices read distinctly. Steps: 1. Open the defaults packet. 2. Identify the four choices without help. 3. Activate each and inspect its persisted effect. 4. Mark each distinction separately. Fail: any common review failure, two concepts blending together, or inert choice. Time/needs: 4–6 minutes; 0757 packet.

41. TASK 0761 — judge Apps and sending

This checks that account, sending style, and protected-message choices are clear at a glance. Why only Liam: at-a-glance comprehension is his product call. Steps: 1. Open the packet and identify all three choices without instructions. 2. Read their labels. 3. Activate each named control and inspect the result. 4. Mark the three criteria separately. Fail: any common review failure, unclear/missing choice, or inert control. Time/needs: 4–6 minutes; 0761 packet.

42. TASK 0765 — judge Whitelisting

This checks that OSL's allowed scope is obvious and Save cannot be mistaken for Reset. Why only Liam: he owns the safety presentation judgment. Steps: 1. Open the Whitelisting packet. 2. State what OSL may touch from the screen alone. 3. Exercise Save and Reset separately and inspect persisted state. 4. Mark scope, Save, and Reset clarity. Fail: any common review failure, unclear scope, confused buttons, or wrong/inert result. Time/needs: 4–6 minutes; 0765 packet.

43. TASK 0773 — judge Look

This checks that maximum appearance choice still feels simple and calm. Why only Liam: “simple, calm, not technical” is subjective product judgment. Steps: 1. Open the Look packet. 2. Scan choices without instructions. 3. Change each named choice once and inspect the live result. 4. Record simple, calm, and not-technical separately. Fail: any common review failure or a control-panel feel that needs explanation. Time/needs: 4–6 minutes; 0773 packet.

44. TASK 0781 — judge Window and sounds

This checks that position, movement, tray pictures, sounds, and mute are findable and understandable. Why only Liam: usability judgment is his. Steps: 1. Open the packet. 2. Locate each of the five named areas without help. 3. Activate each control and inspect the physical/persisted effect. 4. Record one check per area. Fail: any common review failure, missing/unclear area, or inert control. Time/needs: 5–7 minutes; 0781 packet.

45. TASK 0789 — judge friend pictures

This checks recognisable pictures, useful fallback initials, and a plain privacy explanation. Why only Liam: recognition and wording need human judgment. Steps: 1. Open picture, fallback, and privacy states. 2. Identify the friend from each picture/initial state. 3. Read the privacy sentence in plain language. 4. Mark all three criteria. Fail: any common review failure, unrecognisable picture/initials, or unclear privacy wording. Time/needs: 3–5 minutes; 0789 packet.

46. TASK 0793 — judge Account

This checks that dangerous account actions are clearly separated from ordinary changes. Why only Liam: safety hierarchy is an owner judgment. Steps: 1. Open the Account packet. 2. Point out ordinary and dangerous areas without instructions. 3. Read each action and exercise only the prepared safe transcript/negative control, not a real destructive action. 4. Mark separation and comprehension. Fail: any common review failure, dangerous action blending with ordinary changes, or unclear wording. Time/needs: 4–6 minutes; 0793 packet.

47. TASK 0817 — judge the Home top bar

This checks that every top-bar control is recognisable, calm, and uncrowded. Why only Liam: visual hierarchy is his call. Steps: 1. Open the Home packet. 2. Name each control from its appearance. 3. Activate each and compare its route/state transcript. 4. Mark recognisable, calm, and uncrowded. Fail: any common review failure, unidentified/inert control, or crowding. Time/needs: 3–5 minutes; 0817 packet.

48. TASK 0825 — judge the Home protection panel

This checks that the summary is reassuring, honest, and gives a new person the next useful step. Why only Liam: trust and tone require owner judgment. Steps: 1. Open the panel packet. 2. Read the summary and state what it promises. 3. Identify and activate the next-step control. 4. Mark reassuring, honest, and useful-next-step separately. Fail: any common review failure, overclaim, unclear next step, or inert control. Time/needs: 3–5 minutes; 0825 packet.

49. TASK 0833 — judge the OSL Friends panel

This checks that friends are recognisable without exposing pictures too widely. Why only Liam: recognition/privacy balance is his call. Steps: 1. Open the friends-panel packet. 2. Identify each friend from the allowed display. 3. Inspect states where pictures must not be exposed. 4. Record recognition and exposure judgments. Fail: any common review failure, unrecognisable friend, or picture visible outside its allowed scope. Time/needs: 4–6 minutes; 0833 packet.

50. TASK 0841 — judge an OSL friend page

This checks that what a friend may see is obvious and Remove differs clearly from Save. Why only Liam: safety comprehension needs his judgment. Steps: 1. Open the friend-page packet. 2. State the visible-to-friend scope from the page alone. 3. Exercise Save and the prepared safe Remove transcript separately. 4. Mark scope and button distinction. Fail: any common review failure, unclear scope, confused buttons, or wrong/inert result. Time/needs: 4–6 minutes; 0841 packet.

51. TASK 0845 — judge OSL tool and service tiles against custody proof

This checks that no tile claims more safety than current continuous-custody evidence proves. Why only Liam: only the owner can accept the product claim. Steps: 1. Open the current tile packet and its fingerprint-joined 0810a receipts. 2. For every protected or Ready tile, verify a complete current interval/surface/trace receipt and green 0810b mutation results. 3. Inspect the two bad candidates: plaintext round-trip only and transient service decryption; reject both if labelled protected/Ready. 4. Activate each tile and record a verdict. Fail: any common review failure, missing/mismatched receipt, starved observer, overclaim, or acceptance of either bad candidate. Time/needs: 8–12 minutes; tiles plus current 0810a/0810b packet.

52. TASK 0849 — judge Arrange tiles

This checks that arrange, hide, restore, and finish are obvious without instructions. Why only Liam: discoverability is his judgment. Steps: 1. Open Arrange tiles. 2. Without help, arrange one tile, hide one, restore it, and finish. 3. Confirm each result persists. 4. Mark all four actions. Fail: any common review failure, action needing explanation, inert control, or lost state. Time/needs: 4–6 minutes; 0849 packet.

53. TASK 0857 — judge honest tile status against custody proof

This checks that the status page is candid, useful, and no safer-sounding than its continuous proof. Why only Liam: final claim honesty is his call. Steps: 1. Open each current status page with its fingerprint-joined 0810a receipt and 0810b mutation record. 2. Check candid limits and useful next information. 3. Reject the plaintext-only and transient-decrypt candidates if either says protected/Ready. 4. Activate named controls and record the verdict. Fail: any common review failure, missing/mismatched receipt, overclaim, unchecked criterion, or accepted bad candidate. Time/needs: 8–12 minutes; status pages and custody packet.

For TASKS 54–61, the review packet is not complete until the same shipping Windows Scrub path uses a real signed-in supported account with a unique sacrificial marker: the scan must discover that marker, an approved delete run must make it absent after provider refresh or re-login, and a real Windows due or Run now event must invoke the shipping scan. A fixture, empty account, canned connection, or disconnected reader/deleter/scheduler fails every one of these eight reviews.

54. TASK 1410 — confirm Discovery consent wording

This checks that the consent panel plainly states real reading, service-rule risk, ban risk, and stopping. Why only Liam: risk acceptance and wording belong to the owner. Steps: 1. Open the shipping Windows Scrub path against a real signed-in supported account containing the unique sacrificial marker. 2. Open Discovery consent and read every risk sentence. 3. Exercise the named consent controls and independently inspect their real effect/refusal. 4. Check off real reading, service-rule risk, ban risk, and stopping, then record the verdict. Fail: any common review failure, fixture/empty account, missing risk, inert control, or absent real marker path. Time/needs: 5–8 minutes; real supported account and 1410 packet.

55. TASK 1415 — confirm the What to find screen

This checks that a new user understands the filters produce possible matches for review, not facts or automatic deletion. Why only Liam: that comprehension judgment is his. Steps: 1. In the same real Scrub run, open Discovery filters with every choice visible. 2. Read and activate every filter. 3. State what the results mean from the screen alone. 4. Write PASS only if “possible matches for review” is clear and the real marker path remains connected. Fail: any common review failure, hidden/inert filter, automatic/factual implication, fixture, or missing real account marker. Time/needs: 5–8 minutes; live Scrub account and packet.

56. TASK 1420 — confirm the ready-to-run Scrub workspace

This checks that Discovery and AutoScrub look different, Run discovery is clear, and file scanning is optional. Why only Liam: mode clarity is his judgment. Steps: 1. Open the shipping Scrub workspace ready to run against the real supported account. 2. Identify Discovery and AutoScrub without help. 3. Locate Run discovery and the optional file-scan choice. 4. Activate the prepared safe controls and inspect the real marker result, then record the three criteria. Fail: any common review failure, confused modes, unclear run action, file scan appearing mandatory, fixture, or inert production handler. Time/needs: 5–8 minutes; live Scrub packet.

57. TASK 1436 — confirm the running/stop screen

This checks that Stop now cannot be confused with Keep scanning. Why only Liam: destructive workflow clarity needs his judgment. Steps: 1. Start the prepared shipping Discovery run against the real signed-in account. 2. Open the stop confirmation while it is active. 3. Identify the two actions without help. 4. Exercise each in its prepared run and independently inspect whether scanning stopped or continued. 5. Record the verdict. Fail: any common review failure, confused buttons, fixture-only run, or either action not producing its named effect. Time/needs: 6–10 minutes; active real Discovery run.

58. TASK 1448 — confirm the Discovery result screen

This checks that a Discord DM result clearly shows person, date, and time but offers no jump to the message. Why only Liam: clarity and intentional absence of a route need his judgment. Steps: 1. Open the real discovered sacrificial Discord DM result. 2. Locate DM, person, date, and time. 3. Inspect every focusable/announced control and try the prepared activation transcript. 4. Confirm there is no jump route, then record the verdict. Fail: any common review failure, missing location field, jump route/ghost control, or fixture result. Time/needs: 4–7 minutes; real result packet.

59. TASK 1453 — confirm the Pro deletion review

This checks that the exact final deletion count is plain and Cancel really cancels. Why only Liam: destructive consent wording needs owner acceptance. Steps: 1. Open the shipping deletion review for the prepared real sacrificial target. 2. Read the exact count and what will be deleted. 3. Click Cancel and independently refresh the provider to confirm nothing was deleted. 4. Reopen the review in the prepared approved-delete run, confirm the same target/count, perform the deletion, and refresh or re-login to prove the real marker is absent. 5. Record wording, cancellation, and deletion-result verdicts. Fail: any common review failure, unclear/wrong count, inert Cancel, deletion after Cancel, target surviving the approved delete, or fixture-only state. Time/needs: 6–10 minutes; real sacrificial account and provider refresh.

60. TASK 1467 — confirm the AutoScrub schedule

This checks that Only when I choose does not masquerade as an automatic schedule. Why only Liam: schedule meaning is a product judgment. Steps: 1. Open the shipping schedule pane against the real account. 2. Read every schedule choice. 3. Select Only when I choose, save, and inspect the persisted schedule. 4. State whether a new person would mistake it for automation and record the verdict. Fail: any common review failure, option that looks automatic, wrong persisted result, or inert control. Time/needs: 4–7 minutes; schedule packet.

61. TASK 1479 — confirm AutoScrub activity after one failed deletion

This checks that the required next action and the no-jump-link rule are clear. Why only Liam: failure recovery clarity is his judgment. Steps: 1. Use the prepared real account run to produce one genuine failed deletion. 2. Open the AutoScrub activity pane. 3. Identify what the person must do next. 4. Inspect every focusable/announced control and confirm there is no jump link. 5. Record the verdict. Fail: any common review failure, fixture failure, unclear next action, hidden jump/ghost control, or absent provider state. Time/needs: 5–10 minutes; one real failed deletion packet.

62. TASK 1379 — confirm enclave member permissions

This checks whether three members at different permission levels visibly have the right different actions. Why only Liam: role comprehension requires his judgment. Steps: 1. Have two independently running installed OSL clients open the same deployed-service enclave on MON-2. 2. Open the three-member list and read every row. 3. For each member, identify the permission level and activate every allowed button through the shipping screen. 4. Independently read the persisted/remote effect on the other client. 5. Record the verdict. Fail: fixture/in-memory transport, one process with two folders, missing marker/client, unclear permissions, inert control, or blank capture. Time/needs: 8–15 minutes; two real installed clients and a three-member enclave.

63. TASK 1605 — confirm the unsigned-publisher note

This checks that the real-size install note honestly prepares a person for Windows's unsigned/Unknown Publisher warnings. Why only Liam: honesty and readability before install are his judgment. Steps: 1. Open the current release note at its actual display size on a clean Windows machine on MON-2. 2. Read the unsigned and Unknown Publisher warnings, the sentence saying neither proves the installer unsafe, and the checksum instructions. 3. Follow the prepared real action to the release installer state. 4. Write PASS only if it is honest and easy to understand. Fail: any common review failure, stale signed-build claim, missing warning/checksum, promise that Windows trusts it, wrong release artifact, or render-only evidence. Time/needs: 4–7 minutes; exact unsigned release note and clean Windows state.

64. TASK 3108 — review the two tick labels

This checks whether an ordinary person can distinguish the words beside the two marks. Why only Liam: wording judgment belongs to him. Steps: 1. Open the two live mark states together. 2. Read both labels aloud and state what each means. 3. Activate each mark and inspect its persisted result. 4. Write PASS, or REJECTED with the exact replacement words wanted. Fail: any common review failure, indistinguishable meaning, inert mark, vague rejection, or no exact replacement text. Time/needs: 3–5 minutes; 3108 packet.

65. TASK 3117 — review messaging-risk wording

This checks whether the messaging risk page is as honest as Scrub consent. Why only Liam: comparative risk acceptance is his call. Steps: 1. Open the messaging-risk and Scrub-consent pages side by side. 2. Read every risk and consequence. 3. Exercise the named controls and inspect their results. 4. Write PASS, or REJECTED with the exact replacement words wanted. Fail: any common review failure, omitted/softened risk, inert control, or rejection without exact wording. Time/needs: 4–7 minutes; both current packets.

66. TASK 3167 — judge the Behaviour screen

This checks that all six settings, all six Reset buttons, and Back read clearly and work. Why only Liam: final clarity is his judgment. Steps: 1. Open the unresized Behaviour capture and live screen. 2. Locate all six setting names, six Reset buttons, heading, and Back in the UIA tree. 3. Change and reset each prepared setting, then click Back. 4. Record what should change if rejected. Fail: any common review failure, missing control, wrong ROI/mode/alpha, low-colour capture, inert reset/back, or vague verdict. Time/needs: 6–10 minutes; 3167 packet.

67. TASK 3305 — confirm carrier timer limits

This checks honest timer availability and limits across shipping and contract-only carriers. Why only Liam: release honesty is his judgment. Steps: 1. Open the shipping timer screen for Signal, Discord, and protected email. 2. Read and compare Signal/Discord limits and email pointer-only wording. 3. Inspect X, Instagram, and Messenger rows and confirm they are visibly unavailable with no timer control. 4. Record every carrier row. Fail: any common review failure, missing row, contract-only timer, hidden email limit, or starved shipping control. Time/needs: 5–8 minutes; complete 3305 packet.

68. TASK 3317 — confirm the failed-delete list

This checks whether a person can tell which messages failed, in which apps, and why. Why only Liam: recovery comprehension is his judgment. Steps: 1. Open the list containing both prepared real failure entries. 2. For each row, identify message, app, and reason. 3. Exercise every named recovery control and inspect real remote/persisted state. 4. Record one verdict per row. Fail: any common review failure, missing field/row, fixture timer/deletion, inert control, or no independent receiver/provider refresh. Time/needs: 5–10 minutes; two real failure rows and live accounts.

69. TASK 3330 — confirm the waiting-list screen

This checks that three running timers clearly say what will be deleted and when. Why only Liam: countdown comprehension is his judgment. Steps: 1. Open the live waiting list with three nonzero real timer targets. 2. Read all four required fields on every row. 3. Watch each countdown visibly change. 4. Activate named controls and inspect the real state. 5. Record all rows. Fail: any common review failure, zero-of-zero fixture, missing field/row, frozen countdown, inert control, or simulated carrier deletion. Time/needs: 6–10 minutes; three live timer targets.

70. TASK 3332 — confirm the plain-words timer warning

This checks every warning sentence against installed evidence so it does not oversell. Why only Liam: only he decides claim honesty. Steps: 1. Open the real shipping warning and its sentence-by-sentence 3331 evidence. 2. Compare every sentence with the installed release scope. 3. Confirm no X/Instagram/Messenger timer claim and exact email pointer-only limits. 4. Record a verdict per sentence. Fail: any common review failure, omitted sentence, absent evidence, contract-only control/claim, or overstated carrier effect. Time/needs: 5–8 minutes; warning plus 3331 evidence.

71. TASK 3427 — confirm the placing-failure message

This checks whether a person understands which app failed and that nothing was sent. Why only Liam: failure-message clarity is his judgment. Steps: 1. Force the real prepared placing failure in the shipping build. 2. Open the resulting screen. 3. Locate the app name and the words saying nothing was sent. 4. Check the OSL visual rules and record the verdict. Fail: any common review failure, missing app/nothing-sent words, gradient, filled button, shadow, wrong palette/type, or failure branch not real. Time/needs: 3–6 minutes; forced shipping failure packet.

The remaining human and interactive reviews

72. TASK 3524 — blind-review 20 AI cover messages

This checks whether the shipping AI cover writer produces natural messages without using another detector as the judge. Why only Liam: every result needs his own message-specific judgment. Steps: 1. Have the agent invoke AI Covertext 20 times through the shipping Windows control with the local model pack present and save the messages in shuffled order. 2. Review them blind to generation order. 3. Write PASS or REJECTED plus a message-specific reason for every one. 4. Confirm the deliberately repeated/non-sentence canary is rejected. 5. Return the frozen judgments for the missing-pack/control negative checks. Fail: missing message/reason, visible order, accepting the broken canary, fixture generation, blank capture, or missing model/control. Time/needs: 20–30 minutes; prepared blind 20-message set.

73. TASK 3720 — record three unprompted no-OSL reactions

This observes how real recipients react to cover messages when nobody primes them. Why only Liam: he must recruit and honestly avoid briefing three real people. Steps: 1. Arrange three independently identified recipients who do not have OSL; give no advance explanation. 2. Send each a different fresh cover through the real shipping carrier/network. 3. For each, record carrier count zero before and one after. 4. Write their exact first unprompted words, first action, whether they asked the sender, and the exact visible cover. Fail: staged/primed/fixture reaction, delivery not zero-to-one, reused cover, missing recipient, or missing reaction field. Time/needs: 30–60 minutes plus scheduling; three real recipients and live carriers.

74. TASK 3912 — judge the shipping eye open, closed, and unavailable

This checks that the eye's real behavior is clear without claiming unproved carrier delivery. Why only Liam: the final meaning/claim judgment is his. Steps: 1. Open a real signed-in carrier conversation with a marked protected message on MON-2. 2. With the eye closed, confirm words are hidden. 3. Click once and confirm exact words appear; click again and confirm they hide. 4. Open an unavailable-carrier state, click the disabled/refusing eye, and read its reason. 5. Check OSL visual rules and record the verdict. Fail: fixture conversation, blank capture, visible words while closed, wrong words, inert toggle/refusal, missing reason, or design breach. Time/needs: 10–15 minutes; real carrier conversation and unavailable state.

75. TASK 4022 — judge all eight live failure sentences and their controls

This checks that an ordinary person knows what happened and what to do after every shipping failure. Why only Liam: plainness and next-action comprehension need his judgment. Steps: 1. Have the agent force each of the eight live failures one at a time. 2. For each, read the sentence and say what happened/what to do next. 3. Click every retry, unlock, eye, or protection control shown and record its resulting state or refusal. 4. Check OSL visual rules and write one verdict per case. Fail: missing failure, fixture/blank capture, unclear sentence, inert/missing handler, absent click result, or design breach. Time/needs: 25–45 minutes; eight prepared live failures.

76. TASK 4068 — judge Messages waiting from three real arrivals

This checks that the reopened app tells a person something is waiting without revealing what it is. Why only Liam: this is a product-comprehension judgment. Steps: 1. Close the shipping app. 2. From a real sender, deliver three new messages. 3. Reopen OSL and confirm the waiting control shows 3 without private content. 4. Click it, confirm the conversation opens and count becomes 0, then return to the previous screen. 5. Record clarity and state transitions. Fail: fixture arrivals, wrong count, content leak, inert control, count not clearing, blank capture, or design breach. Time/needs: 10–15 minutes; real sender and three live messages.

77. TASK 4262 — judge the honest absence of X, Instagram, and Messenger

This checks that the release omits the three contract-only carriers cleanly while keeping a usable proven-carrier route. Why only Liam: release honesty and layout are his judgment. Steps: 1. Open the shipping no-sidebar Home launcher and service-selection surface on MON-2. 2. Inspect visually, by keyboard focus, and with the announced-control inventory. 3. Confirm zero X, Instagram, or Messenger controls, gaps, ghost targets, or claims. 4. Follow one usable proven-carrier path. 5. Record the verdict. Fail: any common review failure, any visible/focusable/announced contract-only carrier, blank gap/ghost control, or no usable supported path. Time/needs: 5–8 minutes; launcher, selection surface, and 4261 inventory.

78. TASK 4353 — judge mailbox reading consent and undo

This checks whether a person understands permission to read a mailbox and how to turn it off. Why only Liam: consent comprehension requires owner judgment. Steps: 1. Against a real mailbox on MON-2, open the shipping permission screen. 2. Click Refuse and inspect provider/app state. 3. Reopen, click Agree, and inspect the connected state. 4. Click Turn off and inspect the disconnected state. 5. Reopen and confirm the screen accurately reflects that state. 6. Check OSL visual rules and record the verdict. Fail: missing/broken action, fixture account, no provider-open evidence, unclear consent/undo, blank capture, or design breach. Time/needs: 15–20 minutes; real mailbox account.

79. TASK 4414 — watch real arrivals across every privacy lifecycle boundary

This proves unattended arrivals repaint the row but keep private words out of pixels and accessibility surfaces until a fresh reveal. Why only Liam: it requires sustained human observation and his judgment of the real arrival/recovery experience.

Steps:

1. Open a focused real carrier conversation on MON-2. Turn the eye on, and within 60 seconds have a distinct sender deliver unique private words through the real carrier and wake-up service. Record the active arrival.

2. Repeat with separate fresh arrivals after each boundary: default start, Windows lock, user switch, suspend, hibernate, app restart, focus loss, and 60 seconds of inactivity. Leave the eye shown before the boundary where possible.

3. In every boundary run, confirm the row/unread state repaints but zero private words appear before a new reveal.

4. For every hidden run, have the independent observer save the Windows UIA/accessibility tree and cache scan; confirm the exact private words, substrings, and recoverable values occur zero times outside encrypted storage.

5. With the app focused again, reveal and hide the independently saved words using the pointer, keyboard-only Enter/Space, and a screen reader. Confirm the screen reader announces reveal/hide and the changed state.

6. Compare frames for eye state/position, focus, and window bounds; confirm no unexpected jump.

7. Review all saved negative controls: gradient, filled button, starved arrival, inert/inaccessible eye, stale state, focus jump, words surviving in pixels, and words surviving only in UIA. All must be rejected.

Fail: fixture/injected notice, omitted lifecycle boundary, missing UIA/cache scan, premature private words in pixels or accessibility state, pointer-only operation, unlabeled/inert eye, missing screen-reader announcement, focus jump, blank capture, missing negative control, or any negative control accepted. Time/needs: 90–180 minutes; real sender/carrier, lifecycle control, continuous capture, independent UIA/cache scanner, keyboard, and screen reader.

Finish order after the eight unblockers

Keep the relevant accounts and people in place: run 1051 and 1058 immediately after 1040; run 4153 immediately after a successful 1619; run 0570 after 0569; run 3202 and 1510 after 3201; then complete the X/Instagram, Discord/carrier, prepared review, and long lifecycle sittings. Leave credential cleanup and canonical-remote publication until their prepared evidence is ready, but do not let them turn into another open-ended owner decision.

Steps are provided for 79 current task IDs. TASK 4402 requires no new owner action; its missing work is behavioural evidence owned by the agent tasks.