--- since: 1.0.0 --- # Architecture Diagrams Detailed flow diagrams for individual astwire subsystems. For the top-level pipeline, see [`astwire-architecture.md`](astwire-architecture.md). ## 1. CLI command routing ```text astwire │ ▼ argparse ArgumentParser │ ┌────────────────┼────────────────┐ │ │ │ -v/--version --uninstall [targets...] + flags │ │ │ ▼ ▼ ▼ print version perform_uninstall() deferred imports exit(0) exit(0/1) │ ▼ --since given? │ ┌───────────────┴───────────────┐ │ │ yes no │ │ ▼ │ git_changes.get_changed_files() │ (exit 2: bad scope/ref; │ exit 1: nothing changed) │ │ │ └───────────────┬────────────────┘ ▼ full pipeline (below) (-i also on: --max-depth / --decay-depth prune or skeletonize by hop-distance; --graph renders import edges) ``` ## 2. Import resolution (`-i` / `--resolve-imports`) ```mermaid flowchart TD A["Entry target
file or directory"] --> B{"Is it a directory?"} B -->|Yes| C["rglob('*.py')
seed with every .py file"] B -->|No| D["Seed with the single file"] C --> E["BFS queue
collections.deque"] D --> E E --> F["Pop next file"] F --> G["ast.parse()
ImportVisitor collects
Import / ImportFrom nodes
"] G --> H["For each (module, level)"] H --> I{"level == 0 and
top-level module in
sys.stdlib_module_names?"} I -->|Yes| J["Skip — standard library"] I -->|No| K{"level > 0?
relative import"} K -->|Yes| L["Walk up 'level' parent dirs
from current file"] K -->|No| M["Try project_root / module.path
then current_file.parent / module.path
then each ancestor dir up to project_root"] L --> N["Resolve to .py file
or package/__init__.py"] M --> N N --> O{"Resolved path exists
and is under project_root?"} O -->|Yes, unseen| P["Add to discovered set
+ push to BFS queue"] O -->|No / already seen| Q["Discard"] P --> F Q --> F F -->|queue empty| R["depths = {path: hop-count}
edges = [(from, to), ...]
multi-source BFS: seeds are depth 0,
each import hop adds 1
"] R --> S["attach_manifests()
package.json / go.mod
always folded in at depth 0
"] S --> T{"--max-depth N set?"} T -->|Yes| U["Drop paths with depth > N
drop edges touching a dropped path"] T -->|No| V["Keep all discovered paths"] U --> W["Final file list
+ depths, edges for --analysis / --graph"] V --> W classDef entry fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485 class A,D,C entry class B,I,K,O,T decision class E,F,G,H,L,M,N,P,Q,S process class R,U,V,W output ``` `find_project_root()` (used both to bound resolution and as the config search root) walks upward from the entry path looking for the first of: `.git`, `.astwire.toml`, `pyproject.toml`, `setup.py`, `manage.py`. If none is found before the filesystem root, it falls back to the entry path itself (or its parent, if the entry is a file). Seeds normally come from the entry target(s) above. With `--since REF`, they come from `git_changes.get_changed_files()` instead (see [`§2b`](#2b-change-scoped-seeding---since) below) — everything downstream of seeding (BFS, depth/edge tracking, `--max-depth`, manifests) is unchanged either way. `--decay-depth N` doesn't change this diagram's output — it's consumed later, per file, in the transform pipeline: any file whose recorded `depth > N` gets skeletonized even without `--skeleton`. `--graph` renders `edges` (filtered to whichever files land in the same output partition) as a block in the chosen output format; see [`output-formats.md`](../user-guide/output-formats.md). ## 2b. Change-scoped seeding (`--since`) ```mermaid flowchart TD A["--since REF
+ one scope path"] --> B["get_git_root(scope)"] B --> C{"Inside a git repo?"} C -->|No| D["exit 2:
not a git repository"] C -->|Yes| E["is_ref_valid(root, REF)"] E --> F{"REF resolves to
a real commit?"} F -->|No| G["exit 2: invalid ref"] F -->|Yes| H["git diff --name-only REF -- scope
+ git ls-files --others --exclude-standard -- scope"] H --> I["Drop paths git itself
reports as ignored"] I --> J{"Any changed files?"} J -->|No| K["exit 1: no changed files under scope"] J -->|Yes| L["collect_seeds_from_paths()
registered language + skip/.gitignore
filtering, same as any other seed
"] L --> M{"Anything left
after filtering?"} M -->|No| N["exit 1: found N changed files,
but all excluded by --lang/skip/.gitignore"] M -->|Yes| O{"-i set?"} O -->|Yes| P["Feed as seeds into
§2 Import resolution BFS"] O -->|No| Q["Bundle these files only
no import expansion"] classDef entry fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b classDef error fill:#fff0f0,stroke:#d64545,stroke-width:2px,color:#5c2020 classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485 class A entry class C,F,J,M,O decision class B,E,H,I,L process class D,G,K,N error class P,Q output ``` `REF` is always used exactly as supplied — this flow never inspects remotes, never falls back to a default branch, and never resolves `origin/main` vs `origin/master` on your behalf. When `-i` is also on and `--decay-depth` isn't set explicitly, it defaults to `0` (the changed files stay full text; anything pulled in via import gets skeletonized). ## 3. Skeleton generation (`--skeleton`) ```mermaid flowchart LR A["Source file content"] --> B["ast.parse()"] B -->|"SyntaxError"| C["Return original
source unchanged"] B -->|OK| D["SkeletonTransformer
ast.NodeTransformer"] D --> E["visit_FunctionDef /
visit_AsyncFunctionDef"] E --> F{"Has a
docstring?"} F -->|Yes| G["body = [docstring, ...]"] F -->|No| H["body = [...]"] D --> I["visit_ClassDef"] I --> J{"Body became
empty?"} J -->|Yes| K["Insert a single ...
keeps class syntactically valid"] J -->|No| L["Leave nested defs
as already transformed"] G --> M["ast.fix_missing_locations()"] H --> M K --> M L --> M M --> N["ast.unparse()
returns skeleton source text"] classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a classDef process fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b classDef decision fill:#f5f3ed,stroke:#777777,stroke-width:1px,color:#444444 classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485 class A io class B,D,E,I,M process class F,J decision class C,N,G,H,K,L output ``` Type hints, decorators, class attributes, and docstrings survive untouched — only executable statement bodies inside function/method definitions are replaced. This same transformer runs whenever `--skeleton` is set *or* a file's recorded import-depth exceeds `--decay-depth N` (§2) — the trigger differs, the transformation is identical. ## 4. Waste analysis & cleansing (`--analysis` / `--strip-waste`) ```mermaid flowchart TD A["Raw file content"] --> B["clean_content()"] B --> C["Normalize \\r\\n and \\r → \\n"] C --> D["Strip trailing whitespace
per line
counts trailing_space_lines"] D --> E["Collapse runs of
2+ blank lines to 1
counts redundant_lines"] E --> F["Trim trailing blank lines
enforce single final newline"] F --> G["Cleaned content"] A --> H["count_tokens(raw)"] G --> I["count_tokens(clean)"] H --> J["WasteProfile
raw_tokens, clean_tokens,
wasted_tokens, waste %
"] I --> J J -->|"--analysis"| K["format_analysis_table()
printed to stdout, no files written"] G -->|"--strip-waste
(normal run)"| L["Content continues to
redaction + FileRecord"] classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485 class A io class B,C,D,E,F,H,I process class G,J,K,L output ``` `--analysis` and `--strip-waste` share the same `clean_content()` function but serve different purposes: `--analysis` only *reports* recoverable waste (dry run, no output file); `--strip-waste` actually applies the cleansing to the content that gets written. ## 5. Token-bounded partitioning (`--max-tokens`) ```mermaid flowchart TD A["Ordered FileRecords"] --> B{"max_tokens set
and > 0?"} B -->|No| C["Single partition
total_parts = 1, no filename suffix"] B -->|Yes| D["current_batch = []
current_tokens = header_overhead (50)"] D --> E["Next record"] E --> F["file_tokens = count_tokens(content)"] F --> G{"batch non-empty AND
current_tokens + file_tokens
> max_tokens?"} G -->|Yes| H["Close current batch as a partition
start new batch with this record"] G -->|No| I["Append record to current batch
current_tokens += file_tokens"] H --> E I --> E E -->|records exhausted| J["Close final batch"] J --> K["Partitions numbered 1..N
filename_suffix = _part_N when N > 1"] classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485 class A io class B,G decision class D,E,F,H,I,J process class C,K output ``` A single file whose own token count exceeds `max_tokens` is never split — it becomes the sole member of its own oversized partition (the "batch non-empty AND over budget" check only closes a batch that already has *other* content). ## 6. Tokenizer fallback chain ```text count_tokens(text) │ ▼ tiktoken installed? │ ┌────┴────┐ │ │ yes no │ │ ▼ ▼ cl100k_base fallback heuristic .encode() len(text) / 3.7 │ (chars-per-token, │ calibrated for ▼ Python source) count returned ``` `tiktoken` is an optional dependency — the compiled binary works with or without it, falling back to the calibrated character-ratio heuristic when it isn't bundled or fails to load. ## 7. Formatter dispatch ```text partition_records() output │ ▼ get_formatter(format_choice) │ ┌────┼────────────┬─────────────┐ │ │ │ "markdown" "llm" "json" (default) │ │ │ │ │ ▼ ▼ ▼ MarkdownFormatter LLMFormatter JSONFormatter │ │ │ ▼ ▼ ▼ ASCII tree + {"part_index", fenced code "total_parts", blocks per file … "graph": [...], "files": [...]} ``` With `-i --graph`, each branch also renders `edges` (from §2) as a block ahead of the file contents — a `## Dependency Graph` bullet list in Markdown, a `` block in the XML formatter (before ``), and the `"graph"` array shown above in JSON (always present; `[]` without `--graph`). An edge only appears if both of its endpoints landed in that same partition — edges crossing a `--max-tokens` partition boundary are dropped rather than left dangling. ## 8. Ignore resolution (`.gitignore` + `.astwireignore`) ```text IgnoreEngine(root) │ ▼ is_ignored(path) │ ▼ walk ancestor directories from path's parent up to root (root → ... → path.parent, in that order) │ ▼ for each ancestor directory: load & cache .gitignore lines load & cache .astwireignore lines combine into one pathspec.PathSpec (gitignore syntax) │ ▼ match relative path against each directory's spec — first match wins │ ▼ ignored / not ignored ``` Specs are cached per-directory (`_spec_cache`) so repeated `is_ignored()` calls across many files in the same tree don't re-read `.gitignore`/`.astwireignore` from disk.