1
0
Fork 0
Code Issues Pull requests Projects Releases 2 Packages Wiki Activity Actions Pages
DocForge2/toreview.md

4.5 KiB
Raw Permalink Blame History

Incremental Adapter Indexing — Implemented Design Record

Status: implemented on Dev-Rewrite. The maintained contract and adapter guide now live in docs/INCREMENTAL_INDEXING.md.

Current behavior

DocForge adapters currently return a complete project projection. A build:

  1. Enumerates and captures every source selected by the adapter.
  2. Re-extracts every node and relationship.
  3. Calculates source, node, edge, and content hashes.
  4. Writes a complete temporary SQLite index, including full-text search data.
  5. Runs an integrity check.
  6. Verifies that canonical source did not change during the build.
  7. Atomically replaces the previous index.

The hashes currently provide identity, stale-state detection, proposal safety, and publication consistency. They are not used as a per-source extraction cache. DocForge changesets record proposed manual-node operations, but changesets are separate from adapter source extraction and index construction.

Ani-web's current adapter build produces roughly 3,700 nodes and 6,200 relationships in about 10 seconds. Its generated adapter index is approximately 9.4 MB. This is inexpensive at development-slice or commit cadence. It would be wasteful, though still initially tolerable for most SSDs, to rebuild continuously on every file-save event.

Proposed direction

Incremental indexing should be a backward-compatible, opt-in DocForge capability, not a private Ani-web optimization.

DocForge core should provide:

  • Persistent per-source fingerprints.
  • Versioned cached extraction records.
  • A contract for added, changed, deleted, and affected sources.
  • Cache invalidation when the adapter version, graph schema, extraction rules, or project descriptor changes.
  • Whole-candidate graph validation before publication.
  • Atomic publication so readers see only a complete old or complete new graph.
  • A deterministic full-rebuild fallback.
  • Reporting that distinguishes cache hits, reparsed sources, invalidated sources, rebuilt relationships, and total build time.

Adapters opting into incremental extraction should:

  1. Compare current source fingerprints with the last validated build.
  2. Reparse only added and changed sources.
  3. Remove cached facts belonging to deleted sources.
  4. Identify reverse dependencies and other globally affected sources.
  5. Recompute all relationships whose evidence may have changed.
  6. Assemble a complete candidate projection from cached and refreshed facts.
  7. Validate that candidate as strictly as a full rebuild.
  8. Publish only after source-stability and integrity checks pass.

The SQLite index may still be rewritten completely because it is small and atomic replacement is simple and safe. The primary optimization is avoiding repeated parsing and static relationship discovery, not avoiding every database page write.

Correctness requirements

  • Incremental and full builds from the same revision must produce identical canonical projections, hashes, node sets, edge sets, and query results.
  • Deleted and renamed sources must leave no orphaned nodes, edges, backlinks, or full-text-search rows.
  • A changed symbol, import, route, plugin declaration, SQL reference, test, or operational unit must invalidate every relationship that can depend on it.
  • Uncertain dependency impact must trigger broader invalidation or a full rebuild. It must never retain a possibly stale edge.
  • Interrupted or failed builds must leave the previous validated index intact.
  • Existing adapters must continue using the complete-projection contract without modification until they explicitly opt in.

Suggested implementation sequence

  1. Add generic cache and invalidation primitives to DocForge without changing the existing adapter behavior.
  2. Add contract tests for cache identity, invalidation, deletion, interruption, and atomic publication.
  3. Implement incremental extraction in the Ani-web adapter as the first adopter.
  4. Run full and incremental builds against identical revisions and compare their complete outputs.
  5. Measure wall time, parsed-source counts, cache hit rate, database writes, and index size before enabling automatic rebuilds.
  6. Add a debounced watcher only after equivalence and failure-recovery tests pass.
  • Rebuild after a coherent development slice.
  • Rebuild before graph-dependent documentation work when the index is stale.
  • Rebuild before validation and rendering.
  • Do not rebuild on every keystroke or un-debounced save event.
  • If automatic rebuilding is added first, debounce and coalesce changes for at least 3060 seconds.