Architecture Diagrams
Detailed flow diagrams for individual astwire subsystems. For the top-level pipeline, see astwire-architecture.md.
On this page ▾
- 1. CLI command routing
- 2. Import resolution (-i / --resolve-imports)
- 2b. Change-scoped seeding (--since)
- 3. Skeleton generation (--skeleton)
- 4. Waste analysis & cleansing (--analysis / --strip-waste)
- 5. Token-bounded partitioning (--max-tokens)
- 6. Tokenizer fallback chain
- 7. Formatter dispatch
- 8. Ignore resolution (.gitignore + .astwireignore)
1. CLI command routing
astwire
│
▼
argparse ArgumentParser
│
┌────────────────┼────────────────┐
│ │ │
-v/--version --uninstall [targets...] + flags
│ │ │
▼ ▼ ▼
print version perform_uninstall() deferred imports
exit(0) exit(0/1) │
▼
--since given?
│
┌───────────────┴───────────────┐
│ │
yes no
│ │
▼ │
git_changes.get_changed_files() │
(exit 2: bad scope/ref; │
exit 1: nothing changed) │
│ │
└───────────────┬────────────────┘
▼
full pipeline (below)
(-i also on: --max-depth / --decay-depth
prune or skeletonize by hop-distance;
--graph renders import edges)2. Import resolution (-i / --resolve-imports)
flowchart TD
A["Entry target<br/><small>file or directory</small>"] --> B{"Is it a directory?"}
B -->|Yes| C["rglob('*.py')<br/><small>seed with every .py file</small>"]
B -->|No| D["Seed with the single file"]
C --> E["BFS queue<br/><small>collections.deque</small>"]
D --> E
E --> F["Pop next file"]
F --> G["ast.parse()<br/><small>ImportVisitor collects<br/>Import / ImportFrom nodes</small>"]
G --> H["For each (module, level)"]
H --> I{"level == 0 and<br/>top-level module in<br/>sys.stdlib_module_names?"}
I -->|Yes| J["Skip — standard library"]
I -->|No| K{"level > 0?<br/><small>relative import</small>"}
K -->|Yes| L["Walk up 'level' parent dirs<br/>from current file"]
K -->|No| M["Try project_root / module.path<br/>then current_file.parent / module.path<br/>then each ancestor dir up to project_root"]
L --> N["Resolve to .py file<br/>or package/__init__.py"]
M --> N
N --> O{"Resolved path exists<br/>and is under project_root?"}
O -->|Yes, unseen| P["Add to discovered set<br/>+ push to BFS queue"]
O -->|No / already seen| Q["Discard"]
P --> F
Q --> F
F -->|queue empty| R["depths = {path: hop-count}<br/>edges = [(from, to), ...]<br/><small>multi-source BFS: seeds are depth 0,<br/>each import hop adds 1</small>"]
R --> S["attach_manifests()<br/><small>package.json / go.mod<br/>always folded in at depth 0</small>"]
S --> T{"--max-depth N set?"}
T -->|Yes| U["Drop paths with depth > N<br/>drop edges touching a dropped path"]
T -->|No| V["Keep all discovered paths"]
U --> W["Final file list<br/><small>+ depths, edges for --analysis / --graph</small>"]
V --> W
classDef entry fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A,D,C entry
class B,I,K,O,T decision
class E,F,G,H,L,M,N,P,Q,S process
class R,U,V,W outputfind_project_root() (used both to bound resolution and as the config search root) walks upward from the entry path looking for the first of: .git, .astwire.toml, pyproject.toml, setup.py, manage.py. If none is found before the filesystem root, it falls back to the entry path itself (or its parent, if the entry is a file).
Seeds normally come from the entry target(s) above. With --since REF, they come from git_changes.get_changed_files() instead (see §2b below) — everything downstream of seeding (BFS, depth/edge tracking, --max-depth, manifests) is unchanged either way.
--decay-depth N doesn't change this diagram's output — it's consumed later, per file, in the transform pipeline: any file whose recorded depth > N gets skeletonized even without --skeleton. --graph renders edges (filtered to whichever files land in the same output partition) as a block in the chosen output format; see output-formats.md.
2b. Change-scoped seeding (--since)
flowchart TD
A["--since REF<br/><small>+ one scope path</small>"] --> B["get_git_root(scope)"]
B --> C{"Inside a git repo?"}
C -->|No| D["exit 2:<br/>not a git repository"]
C -->|Yes| E["is_ref_valid(root, REF)"]
E --> F{"REF resolves to<br/>a real commit?"}
F -->|No| G["exit 2: invalid ref"]
F -->|Yes| H["git diff --name-only REF -- scope<br/>+ git ls-files --others --exclude-standard -- scope"]
H --> I["Drop paths git itself<br/>reports as ignored"]
I --> J{"Any changed files?"}
J -->|No| K["exit 1: no changed files under scope"]
J -->|Yes| L["collect_seeds_from_paths()<br/><small>registered language + skip/.gitignore<br/>filtering, same as any other seed</small>"]
L --> M{"Anything left<br/>after filtering?"}
M -->|No| N["exit 1: found N changed files,<br/>but all excluded by --lang/skip/.gitignore"]
M -->|Yes| O{"-i set?"}
O -->|Yes| P["Feed as seeds into<br/>§2 Import resolution BFS"]
O -->|No| Q["Bundle these files only<br/><small>no import expansion</small>"]
classDef entry fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef error fill:#fff0f0,stroke:#d64545,stroke-width:2px,color:#5c2020
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A entry
class C,F,J,M,O decision
class B,E,H,I,L process
class D,G,K,N error
class P,Q outputREF is always used exactly as supplied — this flow never inspects remotes, never falls back to a default branch, and never resolves origin/main vs origin/master on your behalf. When -i is also on and --decay-depth isn't set explicitly, it defaults to 0 (the changed files stay full text; anything pulled in via import gets skeletonized).
3. Skeleton generation (--skeleton)
flowchart LR
A["Source file content"] --> B["ast.parse()"]
B -->|"SyntaxError"| C["Return original<br/>source unchanged"]
B -->|OK| D["SkeletonTransformer<br/><small>ast.NodeTransformer</small>"]
D --> E["visit_FunctionDef /<br/>visit_AsyncFunctionDef"]
E --> F{"Has a<br/>docstring?"}
F -->|Yes| G["body = [docstring, ...]"]
F -->|No| H["body = [...]"]
D --> I["visit_ClassDef"]
I --> J{"Body became<br/>empty?"}
J -->|Yes| K["Insert a single ...<br/><small>keeps class syntactically valid</small>"]
J -->|No| L["Leave nested defs<br/>as already transformed"]
G --> M["ast.fix_missing_locations()"]
H --> M
K --> M
L --> M
M --> N["ast.unparse()<br/><small>returns skeleton source text</small>"]
classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef process fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef decision fill:#f5f3ed,stroke:#777777,stroke-width:1px,color:#444444
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A io
class B,D,E,I,M process
class F,J decision
class C,N,G,H,K,L outputType hints, decorators, class attributes, and docstrings survive untouched — only executable statement bodies inside function/method definitions are replaced. This same transformer runs whenever --skeleton is set or a file's recorded import-depth exceeds --decay-depth N (§2) — the trigger differs, the transformation is identical.
4. Waste analysis & cleansing (--analysis / --strip-waste)
flowchart TD
A["Raw file content"] --> B["clean_content()"]
B --> C["Normalize \\r\\n and \\r → \\n"]
C --> D["Strip trailing whitespace<br/>per line<br/><small>counts trailing_space_lines</small>"]
D --> E["Collapse runs of<br/>2+ blank lines to 1<br/><small>counts redundant_lines</small>"]
E --> F["Trim trailing blank lines<br/>enforce single final newline"]
F --> G["Cleaned content"]
A --> H["count_tokens(raw)"]
G --> I["count_tokens(clean)"]
H --> J["WasteProfile<br/><small>raw_tokens, clean_tokens,<br/>wasted_tokens, waste %</small>"]
I --> J
J -->|"--analysis"| K["format_analysis_table()<br/><small>printed to stdout, no files written</small>"]
G -->|"--strip-waste<br/>(normal run)"| L["Content continues to<br/>redaction + FileRecord"]
classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A io
class B,C,D,E,F,H,I process
class G,J,K,L output--analysis and --strip-waste share the same clean_content() function but serve different purposes: --analysis only reports recoverable waste (dry run, no output file); --strip-waste actually applies the cleansing to the content that gets written.
5. Token-bounded partitioning (--max-tokens)
flowchart TD
A["Ordered FileRecords"] --> B{"max_tokens set<br/>and > 0?"}
B -->|No| C["Single partition<br/><small>total_parts = 1, no filename suffix</small>"]
B -->|Yes| D["current_batch = []<br/>current_tokens = header_overhead (50)"]
D --> E["Next record"]
E --> F["file_tokens = count_tokens(content)"]
F --> G{"batch non-empty AND<br/>current_tokens + file_tokens<br/>> max_tokens?"}
G -->|Yes| H["Close current batch as a partition<br/>start new batch with this record"]
G -->|No| I["Append record to current batch<br/>current_tokens += file_tokens"]
H --> E
I --> E
E -->|records exhausted| J["Close final batch"]
J --> K["Partitions numbered 1..N<br/><small>filename_suffix = _part_N when N > 1</small>"]
classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A io
class B,G decision
class D,E,F,H,I,J process
class C,K outputA single file whose own token count exceeds max_tokens is never split — it becomes the sole member of its own oversized partition (the "batch non-empty AND over budget" check only closes a batch that already has other content).
6. Tokenizer fallback chain
count_tokens(text)
│
▼
tiktoken installed?
│
┌────┴────┐
│ │
yes no
│ │
▼ ▼
cl100k_base fallback heuristic
.encode() len(text) / 3.7
│ (chars-per-token,
│ calibrated for
▼ Python source)
count returnedtiktoken is an optional dependency — the compiled binary works with or without it, falling back to the calibrated character-ratio heuristic when it isn't bundled or fails to load.
7. Formatter dispatch
partition_records() output
│
▼
get_formatter(format_choice)
│
┌────┼────────────┬─────────────┐
│ │ │
"markdown" "llm" "json"
(default) │ │
│ │ │
▼ ▼ ▼
MarkdownFormatter LLMFormatter JSONFormatter
│ │ │
▼ ▼ ▼
ASCII tree + <context> {"part_index",
fenced code <file path=…> "total_parts",
blocks per file …</file> "graph": [...],
</context> "files": [...]}With -i --graph, each branch also renders edges (from §2) as a block ahead of the file contents — a ## Dependency Graph bullet list in Markdown, a <graph><edge from=… to=…/></graph> block in the XML formatter (before <files>), and the "graph" array shown above in JSON (always present; [] without --graph). An edge only appears if both of its endpoints landed in that same partition — edges crossing a --max-tokens partition boundary are dropped rather than left dangling.
8. Ignore resolution (.gitignore + .astwireignore)
IgnoreEngine(root)
│
▼
is_ignored(path)
│
▼
walk ancestor directories
from path's parent up to root
(root → ... → path.parent, in that order)
│
▼
for each ancestor directory:
load & cache .gitignore lines
load & cache .astwireignore lines
combine into one pathspec.PathSpec (gitignore syntax)
│
▼
match relative path against each
directory's spec — first match wins
│
▼
ignored / not ignoredSpecs are cached per-directory (_spec_cache) so repeated is_ignored() calls across many files in the same tree don't re-read .gitignore/.astwireignore from disk.