1
0
Fork 0
Code Issues Pull requests Projects Releases 2 Packages Wiki Activity Actions Pages
DocForge2/docs/ADAPTER_AUTHORING_GUIDE.md

549 lines
22 KiB
Markdown

# Language Adapter Authoring Guide
This guide explains how to turn compiler, parser, build-system, or documentation evidence into a
deterministic DocForge project graph. It covers the design work that is easy to miss when a small
fixture is expanded into a complete repository.
Use this guide after the repository assessment in
[Project Onboarding](PROJECT_ONBOARDING.md). The onboarding checklist decides whether an adapter
is needed. This guide defines how to build and prove one.
## Public authoring surface
Adapter authors should import typed contracts, validation helpers, graph models, and the
conformance helper from `docforge.adapter_sdk`:
```python
from docforge.adapter_sdk import (
AdapterAssembly,
AdapterImplementation,
AdapterManifest,
AdapterProject,
AdapterProjectSettings,
AdapterProjection,
AdapterSource,
AdapterSourceProjection,
LogicProjection,
verify_adapter_conformance,
)
```
Language-specific modules under `docforge.adapters` are repository reference implementations, not
the general authoring namespace. Their exact, deliberately narrow behavior is documented in
[Reference Adapters](REFERENCE_ADAPTERS.md).
## Required outcome
A production adapter must provide one reproducible public graph from authoritative project
evidence. It must not guess facts from filenames, preserve unstable parser identities, publish the
same declaration from several extraction units, or let an incremental cache become a second
source of truth.
The complete and incremental paths must publish exactly the same:
- project identity, adapter identity, revision, and source hash;
- nodes and stable node identities;
- relationships and their deterministic metadata;
- function Logic projections and owners;
- source paths and anchors;
- validation failures for invalid input.
Performance does not relax this requirement. A fast graph that sometimes retains stale or
translation-unit-dependent facts is invalid.
## The four adapter layers
Keep these layers separate:
1. **Evidence discovery** finds the authoritative build and source inputs.
2. **Source extraction** converts one extraction unit into raw, deterministic evidence.
3. **Assembly** resolves overlap and assigns each published fact to one owner.
4. **DocForge publication** validates, caches, indexes, queries, and visualizes the assembled graph.
DocForge supplies the publication contracts. A language frontend owns the first three layers
because only the frontend understands the language's build semantics, identity rules, generated
evidence, and source ownership.
The contracts are:
```python
class MyAdapter:
def load_manifest(self) -> AdapterManifest: ...
def extract_source(self, source: AdapterSource) -> AdapterSourceProjection: ...
def assemble_projection(
self,
manifest: AdapterManifest,
contributions: tuple[AdapterSourceProjection, ...],
) -> AdapterAssembly: ...
def load_complete_assembly(self) -> AdapterAssembly: ...
def load_projection(self) -> AdapterProjection: ...
```
`assemble_projection()` is optional only when extraction units already have disjoint ownership.
`load_projection()` is always required. It is the clean primary-graph compatibility oracle.
`load_complete_assembly()` supplies the cache-independent complete graph-plus-Logic oracle. It is
required when incremental contributions publish Logic; without it, DocForge cannot prove Logic
parity.
## What the conformance helper proves
Run the public helper against a fresh confined cache root:
```python
report = verify_adapter_conformance(
MyAdapter(project_root),
cache_root=project_root / ".docforge" / "cache" / "adapter-conformance",
)
```
`verify_adapter_conformance()` proves:
- two repeated complete assemblies are exactly deterministic, including Logic;
- when `load_complete_assembly()` exists, its primary graph exactly matches
`load_projection()`; and
- when the incremental methods exist, the assembled incremental graph and Logic exactly match the
independent complete oracle.
The report records stable identity, counts, and an assembly hash. This helper does not by itself
prove process confinement, implementation restart behavior, corrupt-cache recovery, warm zero
parsing, retrieval, or every case in the proof matrix below. Keep those as separate focused and
integration tests.
## Step 1: define authority before parsing
Write down which tool owns each fact before implementing extraction.
Typical authorities include:
| Fact | Suitable authority |
|---|---|
| Source inventory | Build graph, workspace manifest, or declared source roots |
| Active flags and features | Build-system output |
| Declaration identity | Compiler or language-service semantic identity |
| Definition location | Compiler or parser source location |
| Calls and inheritance | Resolved semantic evidence |
| Include or module dependencies | Compiler or build graph |
| Function control-flow shape | A language-aware analyzer |
| Documentation meaning | Canonical documentation sources |
Do not let a syntax parser overrule a compiler on semantics. Do not infer a resolved call merely
because names match. When the selected frontend cannot prove a fact, omit it or label it as
explicitly derived or proposed.
Record:
- frontend and compiler versions;
- build-system and feature configuration;
- supported source and generated-source roots;
- supported declaration and relationship kinds;
- unsupported facts;
- any normalization applied to frontend output.
Changing one of these rules normally requires an extractor or adapter version change.
## Step 2: define the extraction unit
An extraction unit is the smallest input that can be fingerprinted, invalidated, and re-extracted
without hidden state.
Examples include:
- one compiler translation unit;
- one module or crate target;
- one Java source set or compilation unit;
- one canonical documentation file;
- one generated API registry snapshot.
Use the build system's real unit rather than an arbitrary file grouping. One source file is not
necessarily one semantic unit when features, generated files, module resolution, or compiler flags
change its meaning.
Every unit needs:
- a stable source ID;
- a project-confined relative path;
- a content or semantic fingerprint;
- an extractor version;
- direct dependencies whose changes can alter its evidence.
Do not include output paths, temporary directories, wall time, process IDs, pointer values, or
unordered container iteration in any identity or fingerprint.
## Step 3: prove one bounded extraction
Begin with one representative unit. Include enough language behavior to expose identity and
ownership problems:
- declaration and definition;
- nested and out-of-line ownership;
- resolved call;
- inheritance or interface implementation;
- an internal or private symbol;
- one function suitable for Logic extraction;
- one shared declaration imported by another unit.
The first proof should establish:
- repeated byte-stable extraction;
- project-root confinement;
- stable source and symbol IDs;
- exact source paths and anchors;
- relationship endpoint validity;
- exact Logic ownership;
- explicit rejection of unsafe or ambiguous build input.
Do not activate the adapter in a routine project session at this point. A correct single-unit
projection does not prove whole-project ownership.
## Step 4: design stable identities
Stable identity is a semantic design decision, not a serialization detail.
Prefer, in order:
1. a stable compiler or language-tool symbol identity;
2. a normalized semantic key built from qualified ownership, symbol kind, and signature;
3. a source-qualified identity only for symbols whose language visibility is source-local.
Avoid:
- parser object addresses or transient declaration IDs;
- traversal order;
- result-set position;
- source line as the entire identity;
- a display name without namespace, owner, or signature;
- one identity policy for both externally shared and source-local symbols.
Definitions and declarations of the same externally visible symbol must converge. Anonymous,
private-to-unit, or internal-linkage symbols must remain distinct when the language makes them
distinct.
Every normalization used inside an identity must be tested against more than one extraction unit.
Some compiler fields change only after a symbol is instantiated, referenced, or fully evaluated.
If such a field is not part of authored identity, remove or normalize it before hashing.
## Step 5: separate raw evidence from the public graph
Many semantic tools repeat the same declaration in every unit that imports a header, module,
crate, package, or generated interface. That repetition is valid raw evidence but invalid public
ownership.
Do not force raw extraction to guess the final owner before the complete dependency inventory is
known. Instead:
1. extract deterministic raw contributions;
2. inventory the complete current contribution set;
3. assign each shared source or symbol to one deterministic owner;
4. publish each node, relationship, and Logic record once;
5. reject conflicting evidence instead of silently choosing incompatible values.
The raw contribution cache may contain overlap. The assembled DocForge graph may not.
### Ownership rules
Define ownership for every published fact. A common policy is:
| Fact | Recommended owner |
|---|---|
| Shared source declaration | Deterministically selected dependent unit or declared module owner |
| Unit-private symbol | Its extraction unit |
| Function Logic | The unit owning the exact function definition |
| Containment | The owner of the contained member |
| Call relationship | The owner of the caller |
| Inheritance relationship | The owner of the derived type |
| Source declaration/definition edge | The owner of the source or symbol selected by the adapter |
The correct policy depends on the language. What matters is that it is explicit, deterministic,
and identical in complete and incremental assembly.
Relationship metadata also needs deterministic reduction. Repeated evidence locations must not
grow with the number or order of extraction units. Select one stable evidence record, or define a
bounded ordered representation with a documented reason.
## Step 6: normalize frontend-generated noise
Whole-project extraction exposes facts that a single fixture will not. Compilers and language
services may emit:
- implicit template or generic instantiations;
- synthesized methods and bridge functions;
- default constructors or defaulted functions;
- inferred exception or effect annotations;
- generated annotation-processor output;
- macro expansions;
- duplicate declarations with different amounts of semantic completion;
- declarations from libraries outside the project root.
For each category, choose one policy:
- publish as authored evidence;
- publish as derived evidence with a stable identity;
- attach to an authored owner without becoming a primary node;
- omit as compiler-use noise.
Do not keep a field merely because the frontend emits it. If it changes based on whether another
unit uses the symbol, it will break determinism or complete/incremental parity unless normalized.
Add a regression fixture for every normalization rule. The fixture should demonstrate the
unstable input and the expected stable output.
## Step 7: compose the complete reference graph
Run the frontend over the complete supported source inventory and apply the final ownership
partition.
The composition must:
- sort units and outputs deterministically;
- converge declarations and definitions;
- preserve source-local identity;
- resolve semantic parents rather than relying only on visual parser nesting;
- reject duplicate node identities with conflicting semantic metadata;
- reject conflicting Logic for one owner;
- reject duplicate source IDs;
- reject missing relationship endpoints;
- retain explicit status when optional Logic analysis cannot parse a function.
Run two independent complete extractions and compare exact serialized projections or their
canonical fingerprints.
Record:
- extraction-unit count;
- node, relationship, and Logic counts;
- relationship counts by important type;
- duration and peak memory;
- output or index size;
- exact fingerprint.
These measurements establish the reference shape. They are not performance promises across
machines.
## Step 8: add dependency-aware incremental extraction
Use authoritative dependency evidence whenever possible:
- compiler dependency output;
- module or crate graph;
- Maven or Gradle compilation graph;
- generated-source and annotation-processor inputs;
- explicitly declared documentation dependencies.
The manifest must include every unit that can invalidate a cached contribution. A dependency-only
unit may publish an empty contribution; it still needs a fingerprint and stable ID so reverse
dependents are invalidated.
The incremental path is:
1. build the current manifest;
2. compare fingerprints, extractor versions, additions, and deletions;
3. invalidate changed units and their reverse dependents;
4. reuse only valid raw contributions;
5. extract invalidated units;
6. assemble the complete current contribution set;
7. validate the candidate graph;
8. reread the manifest to detect concurrent source changes;
9. atomically publish the cache and index.
Missing, incompatible, or corrupt cache data is a cache miss. It must never become a partial graph
or replace the last valid index.
### Declare the process-stable adapter implementation boundary
An adapter object is loaded once when its project-bound process starts. Source synchronization can
refresh the graph, but it cannot safely replace already imported adapter code in place.
`AdapterProject` automatically fingerprints a project-local Python package containing the loader
class and a declared descriptor file. Declare a broader or non-Python boundary explicitly when the
adapter uses helpers, configuration, schemas, or templates outside that inferred package:
```python
settings = AdapterProjectSettings(
implementation=AdapterImplementation(
roots=(project_root / "docforge_adapter",),
files=(project_root / ".docforge" / "project.toml",),
suffixes=(".py", ".toml"),
),
)
```
The boundary is confined to the project root and limited to 4,096 files and 64,000,000 bytes.
DocForge fingerprints relative paths and bytes. An added, changed, deleted, missing, or symlinked
implementation file produces `adapter_restart_required` before another MCP operation. Restart the
project-bound server; do not use Python module reloading to mutate a live adapter graph.
The manifest remains a current source snapshot, not a Git-index snapshot. A Git-backed adapter must
omit a deleted source whether its deletion is unstaged or staged. Staging is never a required
DocForge synchronization step.
### Current aggregate bounds
DocForge limits one extraction-cache generation to 10,000 source contributions and 64,000,000
encoded bytes. Malformed, incompatible, missing, oversized, or unsafe cache data is treated as a
cache miss.
The assembled graph is bounded by the effective project `Limits.max_nodes`. When adapter settings
do not override limits, DocForge selects at least 10,000 nodes and raises that bound to the
manifest's `estimated_nodes` when larger. Aggregate assembly ceilings are:
- nodes: `max_nodes`;
- relationships: `max_nodes * 32`;
- Logic nodes across all functions: `max_nodes * 32`; and
- Logic edges across all functions: `max_nodes * 64`.
Adapters should set realistic limits and fail before retaining unbounded frontend evidence. The
cache and assembly bounds do not replace each adapter's own bounded source, command, parser, or
per-function limits.
## Step 9: keep the complete path independent
The full rebuild must not read the incremental extraction cache. Otherwise equivalence compares
the cache with itself and cannot detect stale or incorrectly owned facts.
The complete path must independently:
- rediscover the supported source inventory;
- extract every semantic unit;
- apply the same normalization rules;
- apply the same public ownership partition;
- generate the same project metadata and source hash;
- publish the same nodes, relationships, and Logic.
Raw unpartitioned frontend output is not the equivalence oracle when the incremental assembler
publishes a partitioned graph. Both paths must compare the same public projection shape.
## Step 10: integrate manual and source projections
Keep canonical manual and derived source projections independently rebuildable. Compose them at a
session boundary rather than making source extraction rewrite documentation.
Decide:
- which sessions receive source nodes;
- which documentation families remain isolated;
- whether active context includes source nodes or manual guidance only;
- how manual nodes link to implementation nodes;
- what happens when required build evidence is absent;
- whether source support is required, optional with a null fallback, or disabled.
The source graph has no authority to edit runtime code or canonical documentation. MCP, viewer,
cache, and index operation remain bound to one explicit project root.
## Required proof matrix
An adapter is not complete until these cases pass:
| Case | Required result |
|---|---|
| Repeated single-unit extraction | Exact stable output |
| Two independent complete builds | Exact public projection equality |
| Cold incremental build | Every current unit extracted once |
| Unchanged warm build | Zero reparses |
| Implementation/source change | Only the unit and declared dependents reparse |
| Shared header/module change | Every reverse dependent reparses |
| Added source | New contribution appears without stale duplicates |
| Renamed source | Old contribution disappears and new identity follows policy |
| Unstaged and staged deleted source | Owned nodes, relationships, and Logic disappear identically |
| Adapter implementation edit/add/delete | `adapter_restart_required` before synchronization |
| Adapter descriptor/configuration edit | `adapter_restart_required` before synchronization |
| Build flags/features change | Affected units invalidate |
| Extractor version change | Old contributions invalidate |
| Corrupt cache | Clean recovery without partial publication |
| Interrupted extraction | Last validated index remains active |
| Complete versus incremental | Exact equality |
| Cross-session isolation | Unconfigured sessions cannot see the source graph |
| Viewer and MCP retrieval | Exact symbols, relationships, and Logic are retrievable |
Fixtures must include overlapping shared declarations. A single-file fixture cannot prove
assembly.
## Performance review
Measure before and after adding incremental extraction:
- complete extraction wall time and peak resident memory;
- warm manifest, assembly, validation, and index time;
- raw cache and final index size;
- cache-unit count and cache-hit count;
- node, relationship, and Logic counts;
- checked query latency.
If warm operation remains expensive, identify whether the cost is dependency discovery,
fingerprinting, assembly, validation, or indexing. Do not hide staleness behind an arbitrary time
window. An optimization must retain same-operation source-change detection or replace it with an
equally explicit freshness contract.
## Troubleshooting by symptom
### Node counts change between identical complete builds
Check for unstable frontend IDs, unordered output, generated declarations, inferred type or
exception information, and source paths containing temporary directories.
### The same header or module node appears in many contributions
Keep the overlap in raw evidence and add deterministic assembly ownership. Do not publish every
copy and rely on index deduplication.
### Complete and incremental graphs differ
Compare the public ownership partition first. Confirm that the complete path does not compare
unpartitioned raw output with assembled incremental output. Then compare normalization versions,
dependency inventories, deleted units, and relationship ownership.
### Warm builds consume nearly complete-build memory
Check whether cached raw contributions are too verbose, whether the assembler retains all
frontend AST data, and whether dependency discovery reparses semantic source. Cache only the
bounded projection required for deterministic assembly.
### Relationship counts grow when another unit includes the same source
Define a relationship owner and deterministic evidence reduction. Do not concatenate repeated
evidence from every unit.
### A function has conflicting Logic
Attach Logic only to the exact semantic definition owner. A syntax analyzer may find a similar
function, but line or symbol agreement with the semantic frontend is required before publication.
### An unchanged query is slow
Measure manifest revalidation separately from SQLite lookup. Preserve freshness; optimize the
authoritative dependency and fingerprint path rather than skipping it silently.
## Change and release rules
Change the extractor version when parsing, normalization, identity, relationship, Logic, or
dependency behavior can alter a source contribution.
Change the adapter version when assembly, project metadata, session composition, or public graph
policy changes.
Change the cache schema version when older serialized contributions cannot be read safely.
For every such change:
1. add a fixture for the behavior;
2. run the complete proof matrix;
3. prove clean recovery from the previous disposable cache;
4. record before-and-after graph counts and fingerprints;
5. update the adapter's operating guide and unsupported-fact list.
## Completion checklist
- [ ] Authority for every extracted fact is recorded.
- [ ] Build evidence and source inventory are reproducible.
- [ ] Stable identities distinguish shared and source-local symbols correctly.
- [ ] One representative unit extracts deterministically.
- [ ] Shared declarations and relationships have explicit public owners.
- [ ] Frontend-generated noise has documented normalization rules.
- [ ] Two independent complete builds match exactly.
- [ ] Incremental extraction uses authoritative dependencies.
- [ ] The complete oracle is independent of the cache.
- [ ] Complete and incremental public graph and Logic projections match exactly.
- [ ] `verify_adapter_conformance()` passes, and the separate proof-matrix cases also pass.
- [ ] Corrupt, missing, and interrupted cache cases fail safely.
- [ ] Session composition and family isolation are proven.
- [ ] Viewer, query, context, and Logic retrieval are proven.
- [ ] Performance, graph shape, unsupported facts, and version rules are recorded.