17 KiB
DocForge2 development notes
This is the running implementation record for DocForge2. It records what is active, what was measured, what changed, what failed, why architectural decisions were made, and which ideas were deferred. Stable user and compatibility contracts still belong in dedicated documentation.
Working rules
- Only one milestone is active at a time.
mainremains the last fully verified milestone.- Active implementation occurs on
dev. - Every milestone begins from direct repository evidence and ends with focused tests, the complete repository gate, updated measurements, documentation closeout, and a clean pushed state.
- WorldForge, ScrapeStation, legacy DocForge, and production MCP bindings remain out of scope.
- DocForge2 does not self-host during this program.
- Release tags and Forgejo releases require Rob's explicit approval.
Milestone 0 — complete
Milestone 0 established the public successor, preserved the complete lineage and v1 tag, integrated the no-AST and adapter-lifecycle work, froze compatibility guarantees, added repository-native quality and contract gates, and recorded cold/warm performance, memory, rendering, and response sizes.
The central measurement was decisive: a 1,000-node warm exact lookup took about 286 ms while the generation-pinned SQLite query path took about 0.4–1.4 ms. Repeated whole-source loading and validation, not SQLite, is the first optimization target.
Milestone 1 — active: fast, observable core
Outcome
Warm retrieval should disappear into normal tool overhead. Routine reads must not parse project sources. Status must not render or rebuild hidden work. Results must remain bounded independently of project size.
Starting evidence
- Generic
Project.load()walks, captures, parses, rereads, validates, hashes, and checks Git for the complete source set. - Exact retrieval validates twice around one bounded SQLite query.
- Context compilation performs three full project loads.
- Render status recompiles the complete manual.
- Incremental adapters already prove that manifest attestation can make no-change synchronization and exact retrieval sub-millisecond on a tiny fixture.
- Pinned viewer queries prove the current SQLite schema can serve bounded reads quickly.
Current work
- Audit request-scoped immutable snapshot and persistent-generation options.
- Audit result receipts, pagination, bounded response contracts, and side-effect-free status.
- Audit graph validation complexity, indexed traversal, profiling, and zero-source-parse proofs.
- Reconcile the audits into the smallest additive design that preserves v1 behavior.
- Implement and measure coherent slices, committing only after their gates pass.
Work log
Linear dependency validation
The inherited dependency-cycle preparation scanned every edge once for every node. The graph validator now constructs dependency adjacency in one edge pass and sorts each adjacency list before an iterative deterministic depth-first cycle check. The iterative stack also removes recursion depth as a failure mode on large valid graphs. The same edge pass now rejects missing sources as well as missing targets.
A 10,000-node regression test counts complete edge-collection iteration passes and caps them at four. The focused correctness and bounded-pass tests pass, and the configured strict source type gate is clean.
One validation command initially included tests/test_core.py in a direct Pyright invocation.
Repository Pyright intentionally covers src and tools, so that command reported existing
untyped test-result indexing rather than a source defect. Rerunning the repository-configured type
gate produced zero diagnostics.
Persistent generic source generations
Generic projects now persist a version-1 source-generation receipt only after complete source loading and fully verified index publication. The receipt binds the explicit generic source contract, project and root identity, adapter, revision, source hash, every canonical/authority/ descriptor regular-file identity, and every source-membership directory identity.
A normal warm check reads no canonical source bytes. It validates the known directories and files
directly using device, inode, mode, size, nanosecond modification time, and nanosecond change time.
Directory identities detect add, delete, and rename operations without an rglob. Any missing,
malformed, incompatible, foreign, or dirty receipt becomes a cache miss and falls back to the full
canonical load and row-verification oracle. Successful fallback verification repairs the disposable
receipt.
The receipt is deliberately generic-project behavior. Incremental adapter manifests retain authority over generated or specialist source identities. A one-method legacy adapter continues to work even when it cannot provide a cheap generation.
Request-scoped immutable reads
Index reads now use one read-only SQLite transaction pinned to one verified file signature and one source identity. Existence checks and queries share that connection. Before returning, the request rechecks the index signature and current cheap source generation. A concurrent source or index change fails closed.
Context compilation hydrates nodes and edges from the pinned derived snapshot while retaining
profiles from the immutable descriptor. It no longer loads or parses canonical sources. The public
full Project.load() and deep ProjectIndex.check() behavior remains the recovery and equivalence
oracle.
Focused tests prove that fresh-process-style generic reads can run exact, search, filter,
backlinks, dependency, impact, context, and no-change synchronization operations while
Project.load() is forbidden. They also prove a final source-generation change is rejected before
return and missing/corrupt receipts fall back and repair.
On the maintained 1,000-file fixture, the current work-in-progress measurements are:
| Operation | Milestone 0 median | Milestone 1 WIP median |
|---|---|---|
| Warm no-change synchronize | 142.479 ms | 20.007 ms |
| Exact node | 286.306 ms | 40.277 ms |
| Search, limit 20 | 288.793 ms | 41.455 ms |
| Dependencies, depth 8 | 287.791 ms | 41.551 ms |
| Context, 32k | 436.897 ms | 46.245 ms |
| MCP exact node | 287.094 ms | 40.051 ms |
| MCP context, 32k | 434.853 ms | 46.454 ms |
The three-sample WIP run is directional, not the final Milestone 1 baseline. The final evidence run will use the maintained sample counts and committed clean-tree revision.
Read-only audit reconciliation
The three Milestone 1 audits agreed on the main architecture:
- Keep complete loading and deep checking as independent truth oracles.
- Trust only versioned, identity-bound disposable generation receipts.
- Use one pinned read transaction and retain a final dirty check.
- Hydrate context from the current index.
- Replace full-edge traversal scans with bounded indexed frontier reads.
- Add compact success receipts before allowing large mutations to report post-write size errors.
- Replace hidden render-status rendering with a receipt comparison.
- Add algorithmic counters and parse-count gates alongside wall-clock thresholds.
One audit identified a correctness risk beyond latency: a large mutating MCP operation can commit
successfully and then be replaced by result_too_large. This must be fixed in Milestone 1 so
exactly-once operations never report a false failure after mutation.
Mutation success receipts
Proposal, preview, and canonical-application MCP mutations now declare an internal response policy.
Before runtime validation or mutation, the service proves that a minimum receipt containing the
actual input identity and fixed-length hash fields fits the configured output limit. If it cannot,
the operation returns a preflight size error with mutation_committed = false and does not call the
mutation.
Small results retain the existing full payload. Oversized successful results become a version-1 compact receipt that preserves exact changeset identity, hash, workflow scalars, and lifecycle state while omitting full operations. Application receipts also preserve changed-source counts and derived-refresh status/counts. If the compact form is still too large, the service returns the minimum receipt proven by preflight. It never converts committed success into a post-write size failure.
End-to-end MCP tests exercise two large hash-chained appends followed by canonical application. Each response stays within 1,600 compact JSON characters, exposes the new exact hash, and reports committed success. A separate 700-character preflight test proves the callback and changeset file are never created. Changeset lifecycle receipts now obey the configured changeset byte limit on both write and read.
Receipt-based render status
Successful declared renders now publish a bounded, atomic version-1 receipt below the disposable cache. It binds project/root/adapter/source identity, the normalized view configuration, renderer identity, template and output hashes, byte size, and safe regular-file identities. Generic renders also publish the verified source generation used by cheap status.
Normal status compares only source-generation, descriptor/view, template-file, output-file, and
receipt identities. It does not call Project.load(), prepare the renderer, construct HTML, read
the full output, rebuild the index, or repair missing state. Missing and corrupt receipts are
unverified; source, template, or output changes are stale. An explicit deep option on the
Python, CLI, and MCP status surfaces preserves the old side-effect-free full-render equivalence
oracle.
Receipt failure after atomic output replacement is reported as degraded publication success, not a false render failure. Canonical application converts the same condition into a degraded derived-refresh report while retaining canonical success. Focused tests forbid source loading and renderer preparation during warm status and cover output, template, missing-receipt, corrupt- receipt, and post-publication receipt-failure behavior.
After race hardening, a 50-sample three-node receipt-status check measured a 3.202 ms median and 3.509 ms p95, compared with the 1.941 ms Milestone 0 three-node full-render status. The small fixture does not show the scaling benefit; the 1,000-node Milestone 0 status baseline was 150.591 ms and will be rerun in the final Milestone 1 evidence pass.
Bounded indexed retrieval
Search, metadata filtering, backlinks, dependency traversal, and impact traversal now query one
extra row beyond the requested bound and report limit plus truncated. Backlinks, dependency,
and impact APIs accept the same additive limit option through Python, CLI, and MCP surfaces.
Omitted limits are capped by the project max_results policy.
Traversal no longer loads the complete edge table and repeatedly scans it. It performs
deterministically ordered frontier queries through the existing source primary key or target index.
Each request also has a deterministic edge-examination budget derived from its result limit. The
response includes candidate_edges_consumed, candidate_edges_limit, and truncation_reason
counters so algorithmic work can be asserted independently of machine timing. truncated is true
when either another unique result exists or the work budget prevents proving completeness.
The read-only query-plan audit found that source-ordered unfiltered incoming traversal required a
temporary SQLite sort with the version-2 (target_id, relation, source_id) index. Direct
EXPLAIN QUERY PLAN evidence showed USE TEMP B-TREE FOR ORDER BY. A measured additive
(target_id, source_id, relation) index removes that sort. The disposable index schema is now
version 3, so existing version-2 indexes rebuild without changing canonical source or proposals.
Frontier cursors are streamed and stop immediately on the first omitted unique result. A focused
core, CLI, MCP, Ruff, and Pyright gate passes for this work-in-progress slice.
Structured profiling and zero-work gates
DocForge now has an opt-in, request-local diagnostics collector backed by ContextVar. It emits
one bounded version-1 aggregate with a fixed operation name, outcome, total elapsed nanoseconds,
fixed stage timing keys, and fixed integer counters. It never records paths, node IDs, queries,
source text, or SQL. Disabled mode reads no clock and adds no response field, preserving the
existing CLI and MCP payloads.
The generic loader, adapter projection and extraction paths, source-generation checks, index checks/synchronization/build/read transactions, render status/preparation/output hashing, MCP runtime validation, and viewer-manager requests now expose direct proof counters. A warm incremental adapter cache hit still counts the enclosing project load, so the counters cannot hide full adapter assembly merely because extraction was reused.
MCP servers and the CLI accept the additive --diagnostics startup option. Diagnostics are
attached to structured successes and errors only when the complete MCP response still fits its
configured output budget; they are discarded before any primary result or compact mutation
receipt. Warm generic error decoration now reads the persisted source generation before falling
back to complete loading. Render- and visualization-status error paths explicitly disable both
recovery synchronization and complete identity loading.
Context isolation tests cover threads, concurrent async tasks, repeated stages, nested collectors, exceptions, and disabled collection. Repository tests assert that warm success and error reads, render status, and visualization status perform zero project loads, source parses, adapter projection/extraction, index builds, render preparation, output construction, and output hashing. The result JSON schema contains the same closed operation, stage, and counter sets as the implementation.
The maintained tools/milestone1_benchmark.py harness adds hard counter and p95 latency gates to a
disposable generic project. The smoke target is part of make gate; the 1,000-node evidence run
will be recorded only from a clean committed revision. The historical Milestone 0 harness remains
behaviorally unchanged as comparison evidence; it only exposes shared fixture and measurement
helpers to the Milestone 1 harness.
Visualization snapshot freshness
Visualization workers now receive a version-1 snapshot specification containing the exact validated index publication signature: device, inode, size, modification time, and change time. Both the manager and worker reject a launch if that publication changes before startup. The transmitted project root, root fingerprint, source identity, adapter, counts, limits, and confined index path are strictly validated before the worker may serve source or graph data.
Worker health reports index freshness through stat-only comparison. It does not open SQLite and
does not renew the browser activity lease. The version-2 viewer-manager protocol validates the
complete worker identity and distinguishes an unreachable worker from a live stale worker. A stale
worker stays lifecycle running for accurate diagnosis, but the next visualize request stops it
and launches a newly validated snapshot instead of reusing it.
Client status separately compares the worker's pinned source identity with
IncrementalStateProject.incremental_state(). The composite snapshot is stale if either proof is
stale, current only when both proofs are current, and unknown otherwise. A stopped worker has
unknown snapshot identity. MCP preserves this state at the top-level staleness field and disables
recovery synchronization and full-load error decoration.
Tests cover signature mutation before worker startup, malformed identity, missing and symlinked indexes, stat-only health, unchanged activity, current/unknown/stale source states, live stale workers, non-reuse, and zero-load status. The Milestone 1 benchmark now measures current, stale, not-running, and unavailable visualization status separately with the same zero-work and 50 ms p95 gates as other receipt status operations.
Initial design constraints
- Full rebuild remains the recovery and equivalence oracle.
- Canonical content remains authoritative.
- Existing one-method
load_projection()adapters remain unchanged. - No-AST adapters remain first-class.
- Indexes, source-generation receipts, and caches remain disposable.
- Cheap reads may trust only identity-bound, versioned, corruption-checked receipts.
- Any optimization must fail closed on source mutation and must preserve stale-read refusal.
Future ideas and suggestions
These are notes, not commitments:
- A stable source-generation provider may deserve a public adapter capability only after both the generic project and one incremental adapter prove the same boundary.
- Profiling receipts could eventually feed the human-facing project control panel, but Milestone 1 should expose structured data before adding UI.
- A durable telemetry exporter remains deliberately deferred. Request-local bounded aggregates are enough to prove compiler work in Milestone 1 without adding persistence, cardinality, or privacy risks.
- The stat identity is a cheap publication proof, not a cryptographic integrity scan. Full index validation remains the launch and query oracle.
- Large context and changeset payloads may need cursor pagination or compact immutable receipts. The choice should follow actual client workflows rather than generic pagination machinery.