94 lines
4.4 KiB
Markdown
94 lines
4.4 KiB
Markdown
|
|
# Incremental Adapter Indexing — Design Review Notes
|
|||
|
|
|
|||
|
|
## Current behavior
|
|||
|
|
|
|||
|
|
DocForge adapters currently return a complete project projection. A build:
|
|||
|
|
|
|||
|
|
1. Enumerates and captures every source selected by the adapter.
|
|||
|
|
2. Re-extracts every node and relationship.
|
|||
|
|
3. Calculates source, node, edge, and content hashes.
|
|||
|
|
4. Writes a complete temporary SQLite index, including full-text search data.
|
|||
|
|
5. Runs an integrity check.
|
|||
|
|
6. Verifies that canonical source did not change during the build.
|
|||
|
|
7. Atomically replaces the previous index.
|
|||
|
|
|
|||
|
|
The hashes currently provide identity, stale-state detection, proposal safety, and
|
|||
|
|
publication consistency. They are not used as a per-source extraction cache.
|
|||
|
|
DocForge changesets record proposed manual-node operations, but changesets are
|
|||
|
|
separate from adapter source extraction and index construction.
|
|||
|
|
|
|||
|
|
Ani-web's current adapter build produces roughly 3,700 nodes and 6,200
|
|||
|
|
relationships in about 10 seconds. Its generated adapter index is approximately
|
|||
|
|
9.4 MB. This is inexpensive at development-slice or commit cadence. It would be
|
|||
|
|
wasteful, though still initially tolerable for most SSDs, to rebuild continuously
|
|||
|
|
on every file-save event.
|
|||
|
|
|
|||
|
|
## Proposed direction
|
|||
|
|
|
|||
|
|
Incremental indexing should be a backward-compatible, opt-in DocForge capability,
|
|||
|
|
not a private Ani-web optimization.
|
|||
|
|
|
|||
|
|
DocForge core should provide:
|
|||
|
|
|
|||
|
|
- Persistent per-source fingerprints.
|
|||
|
|
- Versioned cached extraction records.
|
|||
|
|
- A contract for added, changed, deleted, and affected sources.
|
|||
|
|
- Cache invalidation when the adapter version, graph schema, extraction rules, or
|
|||
|
|
project descriptor changes.
|
|||
|
|
- Whole-candidate graph validation before publication.
|
|||
|
|
- Atomic publication so readers see only a complete old or complete new graph.
|
|||
|
|
- A deterministic full-rebuild fallback.
|
|||
|
|
- Reporting that distinguishes cache hits, reparsed sources, invalidated sources,
|
|||
|
|
rebuilt relationships, and total build time.
|
|||
|
|
|
|||
|
|
Adapters opting into incremental extraction should:
|
|||
|
|
|
|||
|
|
1. Compare current source fingerprints with the last validated build.
|
|||
|
|
2. Reparse only added and changed sources.
|
|||
|
|
3. Remove cached facts belonging to deleted sources.
|
|||
|
|
4. Identify reverse dependencies and other globally affected sources.
|
|||
|
|
5. Recompute all relationships whose evidence may have changed.
|
|||
|
|
6. Assemble a complete candidate projection from cached and refreshed facts.
|
|||
|
|
7. Validate that candidate as strictly as a full rebuild.
|
|||
|
|
8. Publish only after source-stability and integrity checks pass.
|
|||
|
|
|
|||
|
|
The SQLite index may still be rewritten completely because it is small and atomic
|
|||
|
|
replacement is simple and safe. The primary optimization is avoiding repeated
|
|||
|
|
parsing and static relationship discovery, not avoiding every database page write.
|
|||
|
|
|
|||
|
|
## Correctness requirements
|
|||
|
|
|
|||
|
|
- Incremental and full builds from the same revision must produce identical
|
|||
|
|
canonical projections, hashes, node sets, edge sets, and query results.
|
|||
|
|
- Deleted and renamed sources must leave no orphaned nodes, edges, backlinks, or
|
|||
|
|
full-text-search rows.
|
|||
|
|
- A changed symbol, import, route, plugin declaration, SQL reference, test, or
|
|||
|
|
operational unit must invalidate every relationship that can depend on it.
|
|||
|
|
- Uncertain dependency impact must trigger broader invalidation or a full rebuild.
|
|||
|
|
It must never retain a possibly stale edge.
|
|||
|
|
- Interrupted or failed builds must leave the previous validated index intact.
|
|||
|
|
- Existing adapters must continue using the complete-projection contract without
|
|||
|
|
modification until they explicitly opt in.
|
|||
|
|
|
|||
|
|
## Suggested implementation sequence
|
|||
|
|
|
|||
|
|
1. Add generic cache and invalidation primitives to DocForge without changing the
|
|||
|
|
existing adapter behavior.
|
|||
|
|
2. Add contract tests for cache identity, invalidation, deletion, interruption,
|
|||
|
|
and atomic publication.
|
|||
|
|
3. Implement incremental extraction in the Ani-web adapter as the first adopter.
|
|||
|
|
4. Run full and incremental builds against identical revisions and compare their
|
|||
|
|
complete outputs.
|
|||
|
|
5. Measure wall time, parsed-source counts, cache hit rate, database writes, and
|
|||
|
|
index size before enabling automatic rebuilds.
|
|||
|
|
6. Add a debounced watcher only after equivalence and failure-recovery tests pass.
|
|||
|
|
|
|||
|
|
## Recommended rebuild policy before this exists
|
|||
|
|
|
|||
|
|
- Rebuild after a coherent development slice.
|
|||
|
|
- Rebuild before graph-dependent documentation work when the index is stale.
|
|||
|
|
- Rebuild before validation and rendering.
|
|||
|
|
- Do not rebuild on every keystroke or un-debounced save event.
|
|||
|
|
- If automatic rebuilding is added first, debounce and coalesce changes for at
|
|||
|
|
least 30–60 seconds.
|