---
since: 1.0.0
---
# Architecture Diagrams
Detailed flow diagrams for individual astwire subsystems. For the top-level pipeline, see [`astwire-architecture.md`](astwire-architecture.md).
## 1. CLI command routing
```text
astwire
│
▼
argparse ArgumentParser
│
┌────────────────┼────────────────┐
│ │ │
-v/--version --uninstall [targets...] + flags
│ │ │
▼ ▼ ▼
print version perform_uninstall() deferred imports
exit(0) exit(0/1) │
▼
--since given?
│
┌───────────────┴───────────────┐
│ │
yes no
│ │
▼ │
git_changes.get_changed_files() │
(exit 2: bad scope/ref; │
exit 1: nothing changed) │
│ │
└───────────────┬────────────────┘
▼
full pipeline (below)
(-i also on: --max-depth / --decay-depth
prune or skeletonize by hop-distance;
--graph renders import edges)
```
## 2. Import resolution (`-i` / `--resolve-imports`)
```mermaid
flowchart TD
A["Entry target
file or directory"] --> B{"Is it a directory?"}
B -->|Yes| C["rglob('*.py')
seed with every .py file"]
B -->|No| D["Seed with the single file"]
C --> E["BFS queue
collections.deque"]
D --> E
E --> F["Pop next file"]
F --> G["ast.parse()
ImportVisitor collects
Import / ImportFrom nodes"]
G --> H["For each (module, level)"]
H --> I{"level == 0 and
top-level module in
sys.stdlib_module_names?"}
I -->|Yes| J["Skip — standard library"]
I -->|No| K{"level > 0?
relative import"}
K -->|Yes| L["Walk up 'level' parent dirs
from current file"]
K -->|No| M["Try project_root / module.path
then current_file.parent / module.path
then each ancestor dir up to project_root"]
L --> N["Resolve to .py file
or package/__init__.py"]
M --> N
N --> O{"Resolved path exists
and is under project_root?"}
O -->|Yes, unseen| P["Add to discovered set
+ push to BFS queue"]
O -->|No / already seen| Q["Discard"]
P --> F
Q --> F
F -->|queue empty| R["depths = {path: hop-count}
edges = [(from, to), ...]
multi-source BFS: seeds are depth 0,
each import hop adds 1"]
R --> S["attach_manifests()
package.json / go.mod
always folded in at depth 0"]
S --> T{"--max-depth N set?"}
T -->|Yes| U["Drop paths with depth > N
drop edges touching a dropped path"]
T -->|No| V["Keep all discovered paths"]
U --> W["Final file list
+ depths, edges for --analysis / --graph"]
V --> W
classDef entry fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A,D,C entry
class B,I,K,O,T decision
class E,F,G,H,L,M,N,P,Q,S process
class R,U,V,W output
```
`find_project_root()` (used both to bound resolution and as the config search root) walks upward from the entry path looking for the first of: `.git`, `.astwire.toml`, `pyproject.toml`, `setup.py`, `manage.py`. If none is found before the filesystem root, it falls back to the entry path itself (or its parent, if the entry is a file).
Seeds normally come from the entry target(s) above. With `--since REF`, they come from `git_changes.get_changed_files()` instead (see [`§2b`](#2b-change-scoped-seeding---since) below) — everything downstream of seeding (BFS, depth/edge tracking, `--max-depth`, manifests) is unchanged either way.
`--decay-depth N` doesn't change this diagram's output — it's consumed later, per file, in the transform pipeline: any file whose recorded `depth > N` gets skeletonized even without `--skeleton`. `--graph` renders `edges` (filtered to whichever files land in the same output partition) as a block in the chosen output format; see [`output-formats.md`](../user-guide/output-formats.md).
## 2b. Change-scoped seeding (`--since`)
```mermaid
flowchart TD
A["--since REF
+ one scope path"] --> B["get_git_root(scope)"]
B --> C{"Inside a git repo?"}
C -->|No| D["exit 2:
not a git repository"]
C -->|Yes| E["is_ref_valid(root, REF)"]
E --> F{"REF resolves to
a real commit?"}
F -->|No| G["exit 2: invalid ref"]
F -->|Yes| H["git diff --name-only REF -- scope
+ git ls-files --others --exclude-standard -- scope"]
H --> I["Drop paths git itself
reports as ignored"]
I --> J{"Any changed files?"}
J -->|No| K["exit 1: no changed files under scope"]
J -->|Yes| L["collect_seeds_from_paths()
registered language + skip/.gitignore
filtering, same as any other seed"]
L --> M{"Anything left
after filtering?"}
M -->|No| N["exit 1: found N changed files,
but all excluded by --lang/skip/.gitignore"]
M -->|Yes| O{"-i set?"}
O -->|Yes| P["Feed as seeds into
§2 Import resolution BFS"]
O -->|No| Q["Bundle these files only
no import expansion"]
classDef entry fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef error fill:#fff0f0,stroke:#d64545,stroke-width:2px,color:#5c2020
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A entry
class C,F,J,M,O decision
class B,E,H,I,L process
class D,G,K,N error
class P,Q output
```
`REF` is always used exactly as supplied — this flow never inspects remotes, never falls back to a default branch, and never resolves `origin/main` vs `origin/master` on your behalf. When `-i` is also on and `--decay-depth` isn't set explicitly, it defaults to `0` (the changed files stay full text; anything pulled in via import gets skeletonized).
## 3. Skeleton generation (`--skeleton`)
```mermaid
flowchart LR
A["Source file content"] --> B["ast.parse()"]
B -->|"SyntaxError"| C["Return original
source unchanged"]
B -->|OK| D["SkeletonTransformer
ast.NodeTransformer"]
D --> E["visit_FunctionDef /
visit_AsyncFunctionDef"]
E --> F{"Has a
docstring?"}
F -->|Yes| G["body = [docstring, ...]"]
F -->|No| H["body = [...]"]
D --> I["visit_ClassDef"]
I --> J{"Body became
empty?"}
J -->|Yes| K["Insert a single ...
keeps class syntactically valid"]
J -->|No| L["Leave nested defs
as already transformed"]
G --> M["ast.fix_missing_locations()"]
H --> M
K --> M
L --> M
M --> N["ast.unparse()
returns skeleton source text"]
classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef process fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef decision fill:#f5f3ed,stroke:#777777,stroke-width:1px,color:#444444
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A io
class B,D,E,I,M process
class F,J decision
class C,N,G,H,K,L output
```
Type hints, decorators, class attributes, and docstrings survive untouched — only executable statement bodies inside function/method definitions are replaced. This same transformer runs whenever `--skeleton` is set *or* a file's recorded import-depth exceeds `--decay-depth N` (§2) — the trigger differs, the transformation is identical.
## 4. Waste analysis & cleansing (`--analysis` / `--strip-waste`)
```mermaid
flowchart TD
A["Raw file content"] --> B["clean_content()"]
B --> C["Normalize \\r\\n and \\r → \\n"]
C --> D["Strip trailing whitespace
per line
counts trailing_space_lines"]
D --> E["Collapse runs of
2+ blank lines to 1
counts redundant_lines"]
E --> F["Trim trailing blank lines
enforce single final newline"]
F --> G["Cleaned content"]
A --> H["count_tokens(raw)"]
G --> I["count_tokens(clean)"]
H --> J["WasteProfile
raw_tokens, clean_tokens,
wasted_tokens, waste %"]
I --> J
J -->|"--analysis"| K["format_analysis_table()
printed to stdout, no files written"]
G -->|"--strip-waste
(normal run)"| L["Content continues to
redaction + FileRecord"]
classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A io
class B,C,D,E,F,H,I process
class G,J,K,L output
```
`--analysis` and `--strip-waste` share the same `clean_content()` function but serve different purposes: `--analysis` only *reports* recoverable waste (dry run, no output file); `--strip-waste` actually applies the cleansing to the content that gets written.
## 5. Token-bounded partitioning (`--max-tokens`)
```mermaid
flowchart TD
A["Ordered FileRecords"] --> B{"max_tokens set
and > 0?"}
B -->|No| C["Single partition
total_parts = 1, no filename suffix"]
B -->|Yes| D["current_batch = []
current_tokens = header_overhead (50)"]
D --> E["Next record"]
E --> F["file_tokens = count_tokens(content)"]
F --> G{"batch non-empty AND
current_tokens + file_tokens
> max_tokens?"}
G -->|Yes| H["Close current batch as a partition
start new batch with this record"]
G -->|No| I["Append record to current batch
current_tokens += file_tokens"]
H --> E
I --> E
E -->|records exhausted| J["Close final batch"]
J --> K["Partitions numbered 1..N
filename_suffix = _part_N when N > 1"]
classDef io fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef decision fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef process fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class A io
class B,G decision
class D,E,F,H,I,J process
class C,K output
```
A single file whose own token count exceeds `max_tokens` is never split — it becomes the sole member of its own oversized partition (the "batch non-empty AND over budget" check only closes a batch that already has *other* content).
## 6. Tokenizer fallback chain
```text
count_tokens(text)
│
▼
tiktoken installed?
│
┌────┴────┐
│ │
yes no
│ │
▼ ▼
cl100k_base fallback heuristic
.encode() len(text) / 3.7
│ (chars-per-token,
│ calibrated for
▼ Python source)
count returned
```
`tiktoken` is an optional dependency — the compiled binary works with or without it, falling back to the calibrated character-ratio heuristic when it isn't bundled or fails to load.
## 7. Formatter dispatch
```text
partition_records() output
│
▼
get_formatter(format_choice)
│
┌────┼────────────┬─────────────┐
│ │ │
"markdown" "llm" "json"
(default) │ │
│ │ │
▼ ▼ ▼
MarkdownFormatter LLMFormatter JSONFormatter
│ │ │
▼ ▼ ▼
ASCII tree + {"part_index",
fenced code "total_parts",
blocks per file … "graph": [...],
"files": [...]}
```
With `-i --graph`, each branch also renders `edges` (from §2) as a block ahead of the file contents — a `## Dependency Graph` bullet list in Markdown, a `` block in the XML formatter (before ``), and the `"graph"` array shown above in JSON (always present; `[]` without `--graph`). An edge only appears if both of its endpoints landed in that same partition — edges crossing a `--max-tokens` partition boundary are dropped rather than left dangling.
## 8. Ignore resolution (`.gitignore` + `.astwireignore`)
```text
IgnoreEngine(root)
│
▼
is_ignored(path)
│
▼
walk ancestor directories
from path's parent up to root
(root → ... → path.parent, in that order)
│
▼
for each ancestor directory:
load & cache .gitignore lines
load & cache .astwireignore lines
combine into one pathspec.PathSpec (gitignore syntax)
│
▼
match relative path against each
directory's spec — first match wins
│
▼
ignored / not ignored
```
Specs are cached per-directory (`_spec_cache`) so repeated `is_ignored()` calls across many files in the same tree don't re-read `.gitignore`/`.astwireignore` from disk.