1
0
Fork 0
Code Issues Pull requests Projects Releases 2 Packages Wiki Activity Actions Pages
DocForge2/README.md

10 KiB

DocForge

DocForge is a project-scoped documentation graph for people and AI agents. It validates canonical documentation, builds a disposable search and relationship index, compiles bounded context, renders declared manuals, visualizes project structure, and manages reviewable documentation changesets.

What it does

  • Validates stable Markdown/TOML nodes and typed relationships.
  • Builds a deterministic SQLite search and graph index.
  • Exposes project-bound CLI and MCP query surfaces.
  • Compiles versioned, generation-bound task context with cited evidence, explicit gaps, and bounded continuation.
  • Records one bounded, versioned latest-generation graph transition without creating a history database.
  • Automatically synchronizes disposable indexes before MCP work.
  • Creates, validates, diffs, and previews isolated changesets.
  • Registers complete proposals atomically without caller-managed hash chaining.
  • Applies one explicitly approved changeset hash through CLI or gated MCP.
  • Supports opt-in incremental adapters with reverse-dependency invalidation and full-build equivalence checks.
  • Detects project-local adapter implementation and configuration changes and requires a fresh project-bound process before any further MCP work.
  • Keeps function-scoped control-flow projections separate from the primary architecture graph.
  • Runs a managed loopback graph browser with neighborhood, semantic Flow, convergence Web, function-scoped Logic, source inspection, and branch-aware node hiding.
  • Supports generic documentation projects and project-owned source adapters.

DocForge never treats indexed text as instructions. It does not run shell commands, mutate Git, build applications, deploy, publish, or select projects globally.

For implementation projects, the recommended cadence is to read canonical documentation during intake, keep it read-only through implementation and focused testing, freeze and validate a release candidate, then perform one atomic documentation closeout before the final commit and tag. This keeps the manual authoritative without using it as an implementation notebook.

Release 1

DocForge 1.0.0 is the first stable product release. It combines the project-scoped graph, CLI and MCP query surfaces, reviewable hash-approved changesets, generic and project-owned adapters, declared rendering, and the complete Nodes/Flow/Web visualization model in one supported release.

The post-1.0 incremental compiler is a backward-compatible, optional enhancement. Existing Release 1 adapters that implement only load_projection() continue to use the original complete-projection path without modification. Adapters gain incremental performance only when they additionally implement the source manifest and extraction methods. Incremental adapters must retain load_projection() as their clean-rebuild fallback and equivalence oracle.

MCP bindings can additionally select --no-ast when an owner wants to preserve an existing non-AST adapter. The binding advertises that policy to clients, forbids adapter rewrites that add AST, Tree-sitter, compiler-AST, or function-Logic extraction, blocks the Logic tool, and rejects nonempty Logic publication. Complete-projection adapters continue unchanged, and non-AST incremental fingerprinting and caching remain allowed.

DocForge2 bindings may also declare --capability-mode read|proposal|application|operator. Bootstrap returns one versioned effective policy and the actual startup-gated capabilities. Existing tool surfaces and the legacy no-AST payload remain compatible.

docforge_get_task_context is an additive read tool for change, implementation, failure, ownership, test, operation, and release work. It derives a closed version-1 retrieval plan, executes it against one immutable index generation, and returns a hash-bound context capsule. Project relation names remain authoritative. DocForge classifies only its versioned alias set and preserves every unknown relation as unclassified instead of guessing semantics.

docforge_get_generation_diff reports the latest verified primary-graph transition through one bounded disposable receipt. It includes exact node and edge change counts, hash-bound retained details, and explicit truncation. Paged results use one top-level cursor and a versioned receipt_header; stored_receipt_hash identifies the complete persisted receipt. The read never exposes Logic details, loads canonical source, repairs derived state, or invents history.

Graph views

The browser presents the primary architecture graph through three complementary views and loads a fourth function-scoped view only when requested:

  • Nodes shows a bounded, relation-neutral neighborhood around the focus. It is the broad inspection view for seeing stored incoming and outgoing relationships without changing their direction. Semantic cards distinguish structure, behavior, dependencies, execution, data, evidence, context, and other relationships.
  • Flow shows semantic origin-to-destination paths that terminate at the focus. DocForge reverses prerequisite-style relationships for presentation, so imports, dependencies, reads, inheritance, definitions, and tests flow toward the thing they help create or exercise.
  • Web shows the larger convergence picture: Flow contributors plus contextual relationships, callers, containers, and direct members or execution dependencies owned by the focus.
  • Logic shows the possible static control paths inside a focused Python, JavaScript, or C++ function or method. Entry, decisions, actions, loops, convergence points, returns, and exceptions connect through explicit TRUE, FALSE, NEXT, CASE, LOOP, RETURN, and RAISE paths. Logic is stored separately and does not add statement-level noise to Nodes, Flow, Web, or search.

Graph cards show the node's readable leaf name and kind without clipping either value. The full qualified identity remains available in the tooltip, compact descriptor, and full inspector. The left browser panel can combine text, family, node-kind, language, and capability filters. Quick presets expose Logic-ready nodes, Python callables, tests, routes, and documentation without requiring users to know stable IDs. Selecting a canvas node emphasizes its directly connected neighbors and edges while muting unrelated paths.

Hide node removes noise without changing the index. In Flow and Web, hiding a contributor also removes upstream ancestors that no longer have a path to the focus. Nodes between the hidden contributor and the focus stay visible, and alternate ancestor paths remain intact. In Logic, hiding a step inserts an explicit omitted-path bridge so downstream control flow remains readable. Restore hidden restores the presentation.

Five-minute start

Requirements are Python 3.12+, uv, and Node.js/npm.

git clone forgejo@repo.andraxion.net:administrator/DocForge2.git /absolute/path/DocForge2
cd /absolute/path/DocForge2
uv sync --group dev
npm ci

PROJECT=/absolute/path/MyProject
.venv/bin/docforge --project-root "$PROJECT" validate
.venv/bin/docforge --project-root "$PROJECT" reindex
.venv/bin/docforge --project-root "$PROJECT" visualize

Install the persistent per-user graph viewer once:

.venv/bin/docforge-viewer-manager install-user-service

Start an MCP server for one project:

.venv/bin/docforge-mcp \
  --project-root "$PROJECT" \
  --proposal-writer project-editor

Add --canonical-applier project-editor only when that MCP integration should expose the hash-bound docforge_apply_changeset tool.

For an unconfigured codebase, begin with a read-only language and documentation assessment:

.venv/bin/docforge --project-root /absolute/path/MyProject onboard

Add --scaffold, a stable project ID, and a title to create, index, and render a generic starter manual. Source files are reported separately and require a validated language frontend before DocForge describes them as a source graph.

Documentation

  • User manual — features, setup, visualization, CLI, MCP, apply, adapters, and troubleshooting.
  • Core contract — invariants and security boundary.
  • Milestone 0 compatibility — preserved package, CLI, MCP, adapter, schema, changeset, rendering, and no-AST guarantees.
  • Milestone 0 baseline — validation evidence, cold and warm performance, memory, rendering and response sizes, bottlenecks, and missing coverage.
  • Milestone 0 closeout — lineage, migration, security scan, repository state, and fresh-clone proof.
  • MCP contract — exact tool and process boundary.
  • Viewer manager — native service setup and lifecycle.
  • Adapter decision — why custom adapters own canonical serialization.
  • Incremental adapter indexing — source-scoped extraction, invalidation, equivalence, relationship changes, and the lazy Logic boundary.
  • Project onboarding — repository assessment, safe manual scaffolding, language frontends, source/manual integration, proof, and MCP activation.
  • Language adapter authoring — implementation sequence, stable identities, overlap ownership, normalization, incremental equivalence, troubleshooting, and the complete adapter proof matrix.

Development

Run the complete repository-native gate:

make gate

Focused entry points are available as make contract, make test, make type, make benchmark-smoke, make benchmark, make benchmark-m1-smoke, and make benchmark-m1.

The committed 1,000-node baseline and its measurement method are under benchmarks/.

Pass --diagnostics to docforge or docforge-mcp to attach bounded request-local stage timings and compiler-work counters. Diagnostics are disabled by default and are dropped before primary MCP results when the configured output budget is tight.

See AGENTS.md before changing core boundaries.