cairntir

Council integration audit — 2026-09-10

Status: local candidate verification passed; native results are tracked on PR #111. This is not a package release.

The starting tree was c92f7e86757edbd98e4b71532d26d5aea45f8858. Its existing suite passed 1,664 tests with eight slow tests deselected and 92.261478% combined statement-and-branch coverage. Newly authored process tests reproduced failures despite that green baseline.

Three agents received separate assignments and fresh contexts. They authored and froze their acceptance before the corresponding repairs; the implementing agent did not change those tests. These are separate investigations within one agent system, not three vendors or human reviewers. An additional local Ollama qwen2.5-coder:7b review restated launcher facts but supplied no new executable defect evidence. No finding is attributed to Lou or Jarvis: their actual fix and reproducer were not available.

Reproduced failures and missed checks

Failure Why prior verification missed it Before → after
Rejected MCP tools reported success; an update banner corrupted JSON Helpers extracted text without checking the SDK result flag; banner coverage did not parse the first real response Six new process cases failed → six passed
Invalid UTF-8 update cache blocked startup; a Unicode version value reported failure after a committed write Ordinary valid-cache cases did not exercise startup and post-commit notification failures Four council failures and one control → five passed, including timeline below
Timeline returned no match when a newer unrelated drawer consumed the limit Filtering was tested without a smaller limit than the candidate set Included in the five MCP council cases
JSON initialization dropped the grant environment, widening a restarted connection to owner access Setup assertions did not launch a restricted connection before and after reconfiguration Twenty-two council failures and one control → twenty-three passed, including store identity below
Relative store overrides selected different databases from different host directories; tilde was literal Existing tests supplied absolute fixture paths Included in the twenty-three continuity cases
Generated launchers depended on shell PATH; automatic Claude registration omitted host identity Setup mocked subprocess success and asserted bare command strings; it never launched the resulting configuration Four council failures and two controls → six passed
Installed-package semantic recall could pass with zero hits A substring assertion matched the query echoed by the no-results message Old oracle accepted a real zero-hit response; replacement has five positive/negative outcome cases

Frozen definitions are plans/council-mcp.freeze.json, plans/council-store-identity.freeze.json and plans/council-registration.freeze.json. All 34 council acceptance cases passed unchanged in independent repair round one. The preservation gate pins those manifests and their four test files alongside the existing 24 manifests and 115 protected artifacts. Unit expectation updates replace obsolete bare-launcher assertions; no existing frozen acceptance or coverage exclusions were changed.

The published 1.12.2 wheel also fails the new real-MCP failed-write receipt check. The injected late SQLite write failure rolled back all tables and a corrected retry succeeded, but the rejected tool was marked successful. Public wheel SHA-256: 6b79adc1b92db524ad49765cdb1275ef9530b25c37bde6f67e161ba923f35d39.

The bounded final review independently reproduced a new serializer regression: non-BMP characters in the installing interpreter path were escaped as surrogate pairs, producing invalid Codex TOML despite a successful configuration receipt. A fourth frozen manifest, plans/council-codex-unicode.freeze.json, preserves three cases: the non-BMP case failed before repair while ASCII/accented controls passed; all three now pass. The second repair round changed only Unicode serialization. All 32 continuity, registration and Unicode cases passed unchanged on independent rerun. A fifteenth mutation restores the broken serializer and is caught by the frozen assertion, with its clean control passing.

Candidate verification

Before the final Unicode serializer repair, all 1,745 collected cases passed across the normal, permission-rerun, new boundary, public evidence-view, model and slow selections. The initial full run had four sandbox permission failures; those four passed with network/process permissions. The combined coverage was 92.262837% with unchanged exclusions, including additional transport files measured by the new subprocess scenarios.

All fourteen original deliberate mutations were caught with passing clean controls; all five historical broken/fixed comparisons passed. Seven production-model evaluation cases and one slow acceptance case passed. The newly installed candidate verified 74 package files and all 21 MCP tools, abrupt resumption, all-table restoration, real semantic retrieval, failed-write rollback/retry and unmodified JSON with update notices. The machine-readable receipt is plans/council-verification-evidence.json.

The first hosted run passed regression and mutation evidence on all three platforms. macOS and Linux Python 3.12 each passed 1,734 source cases and skipped one Windows-only case, but failed the unchanged coverage gate at 91.827676% and 91.879896%. Two independently authored process tests then preserved missing public outcomes: startup grants survive host refresh, and a completed learning episode remains connected across session context, discovery review, calibration, retention and idempotent replay. Both pass locally; native reruns determine whether each platform clears the gate. No runtime or coverage configuration was changed to address this coverage failure. The next Linux Python 3.12 run reached 92.097476%, but two policy checks caught encoding damage from the documentation update. The text was repaired, and all 46 targeted policy/host checks passed. The final tree requires fresh full-suite and installed-package gates on every platform; the PR retains failed runs as well as the final results.

Boundaries

Local process evidence is Windows evidence. The CI changes add macOS behavioral and installed-wheel runs to the existing three-platform source matrix; those hosted results must pass before claiming native macOS verification. A hosted runner cannot prove behavior inside Lou’s particular GUI applications.

Council scenarios used disposable stores and hash embeddings where a production model was unnecessary. Installed-package verification separately uses the production embedding model. No fault injection touched the production store; intentional progress checkpoints are separate. No user application was restarted and no production package was upgraded.

Acknowledged interruption/resumption, exact checkpoint replay, stale revision rejection, concurrent revisions, uncommitted transaction rollback, ambiguity and terminal task behavior held in the continuity investigation. This is bounded evidence, not certification of every possible interaction. Missing-walkthrough retention reporting remains a minor observation outside the frozen repair scope.

Custom Codex TOML and existing CLI-owned registrations may require a manual command/argument update. The repair preserves their access settings and reports the conflict rather than deleting the registration. Pinned virtual environments must remain at their configured paths; rerun setup after relocating an install.

Local raw reports and before/after receipts are retained under .cairntir/integration-audit-20260910/, including council-mcp, council-continuity, council-verification and the Ollama input/output. They are private workspace evidence; CI artifacts provide the hosted verification record.