Status: local candidate verification passed; native results are tracked on PR #111. This is not a package release.
The starting tree was c92f7e86757edbd98e4b71532d26d5aea45f8858. Its existing
suite passed 1,664 tests with eight slow tests deselected and 92.261478% combined
statement-and-branch coverage. Newly authored process tests reproduced failures
despite that green baseline.
Three agents received separate assignments and fresh contexts. They authored
and froze their acceptance before the corresponding repairs; the implementing
agent did not change those tests. These are separate investigations within one
agent system, not three vendors or human reviewers. An additional local Ollama
qwen2.5-coder:7b review restated launcher facts but supplied no new executable
defect evidence. No finding is attributed to Lou or Jarvis: their actual fix and
reproducer were not available.
| Failure | Why prior verification missed it | Before → after |
|---|---|---|
| Rejected MCP tools reported success; an update banner corrupted JSON | Helpers extracted text without checking the SDK result flag; banner coverage did not parse the first real response | Six new process cases failed → six passed |
| Invalid UTF-8 update cache blocked startup; a Unicode version value reported failure after a committed write | Ordinary valid-cache cases did not exercise startup and post-commit notification failures | Four council failures and one control → five passed, including timeline below |
| Timeline returned no match when a newer unrelated drawer consumed the limit | Filtering was tested without a smaller limit than the candidate set | Included in the five MCP council cases |
| JSON initialization dropped the grant environment, widening a restarted connection to owner access | Setup assertions did not launch a restricted connection before and after reconfiguration | Twenty-two council failures and one control → twenty-three passed, including store identity below |
| Relative store overrides selected different databases from different host directories; tilde was literal | Existing tests supplied absolute fixture paths | Included in the twenty-three continuity cases |
| Generated launchers depended on shell PATH; automatic Claude registration omitted host identity | Setup mocked subprocess success and asserted bare command strings; it never launched the resulting configuration | Four council failures and two controls → six passed |
| Installed-package semantic recall could pass with zero hits | A substring assertion matched the query echoed by the no-results message | Old oracle accepted a real zero-hit response; replacement has five positive/negative outcome cases |
Frozen definitions are plans/council-mcp.freeze.json,
plans/council-store-identity.freeze.json and
plans/council-registration.freeze.json. All 34 council acceptance cases passed
unchanged in independent repair round one. The preservation gate pins those
manifests and their four test files alongside the existing 24 manifests and
115 protected artifacts. Unit expectation updates replace obsolete bare-launcher
assertions; no existing frozen acceptance or coverage exclusions were changed.
The published 1.12.2 wheel also fails the new real-MCP failed-write receipt check.
The injected late SQLite write failure rolled back all tables and a corrected
retry succeeded, but the rejected tool was marked successful. Public wheel SHA-256:
6b79adc1b92db524ad49765cdb1275ef9530b25c37bde6f67e161ba923f35d39.
The bounded final review independently reproduced a new serializer regression:
non-BMP characters in the installing interpreter path were escaped as surrogate
pairs, producing invalid Codex TOML despite a successful configuration receipt.
A fourth frozen manifest, plans/council-codex-unicode.freeze.json, preserves
three cases: the non-BMP case failed before repair while ASCII/accented controls
passed; all three now pass. The second repair round changed only Unicode
serialization. All 32 continuity, registration and Unicode cases passed unchanged
on independent rerun. A fifteenth mutation restores the broken serializer and
is caught by the frozen assertion, with its clean control passing.
Before the final Unicode serializer repair, all 1,745 collected cases passed across the normal, permission-rerun, new boundary, public evidence-view, model and slow selections. The initial full run had four sandbox permission failures; those four passed with network/process permissions. The combined coverage was 92.262837% with unchanged exclusions, including additional transport files measured by the new subprocess scenarios.
All fourteen original deliberate mutations were caught with passing clean controls;
all five historical broken/fixed comparisons passed. Seven production-model
evaluation cases and one slow acceptance case passed. The newly installed
candidate verified 74 package files and all 21 MCP tools, abrupt resumption,
all-table restoration, real semantic retrieval, failed-write rollback/retry and
unmodified JSON with update notices. The machine-readable receipt is
plans/council-verification-evidence.json.
The first hosted run passed regression and mutation evidence on all three platforms. macOS and Linux Python 3.12 each passed 1,734 source cases and skipped one Windows-only case, but failed the unchanged coverage gate at 91.827676% and 91.879896%. Two independently authored process tests then preserved missing public outcomes: startup grants survive host refresh, and a completed learning episode remains connected across session context, discovery review, calibration, retention and idempotent replay. Both pass locally; native reruns determine whether each platform clears the gate. No runtime or coverage configuration was changed to address this coverage failure. The next Linux Python 3.12 run reached 92.097476%, but two policy checks caught encoding damage from the documentation update. The text was repaired, and all 46 targeted policy/host checks passed. The final tree requires fresh full-suite and installed-package gates on every platform; the PR retains failed runs as well as the final results.
Local process evidence is Windows evidence. The CI changes add macOS behavioral and installed-wheel runs to the existing three-platform source matrix; those hosted results must pass before claiming native macOS verification. A hosted runner cannot prove behavior inside Lou’s particular GUI applications.
Council scenarios used disposable stores and hash embeddings where a production model was unnecessary. Installed-package verification separately uses the production embedding model. No fault injection touched the production store; intentional progress checkpoints are separate. No user application was restarted and no production package was upgraded.
Acknowledged interruption/resumption, exact checkpoint replay, stale revision rejection, concurrent revisions, uncommitted transaction rollback, ambiguity and terminal task behavior held in the continuity investigation. This is bounded evidence, not certification of every possible interaction. Missing-walkthrough retention reporting remains a minor observation outside the frozen repair scope.
Custom Codex TOML and existing CLI-owned registrations may require a manual command/argument update. The repair preserves their access settings and reports the conflict rather than deleting the registration. Pinned virtual environments must remain at their configured paths; rerun setup after relocating an install.
Local raw reports and before/after receipts are retained under
.cairntir/integration-audit-20260910/, including council-mcp,
council-continuity, council-verification and the Ollama input/output. They are
private workspace evidence; CI artifacts provide the hosted verification record.