---
since: 1.0.0
---
# astwire โ Architecture
astwire is a local-first CLI application that turns a Python codebase into a clean, partitionable LLM context bundle. It separates **target resolution** (which files matter), **AST transformation** (import tracing, skeletonization, waste analysis), **redaction**, **partitioning**, and **formatting** into distinct, independently testable stages. No stage executes the analyzed code โ every transformation operates on the `ast` module's parsed syntax tree.
## Architecture
```mermaid
flowchart TD
CLI["๐ค astwire CLI
cli.py โ argparse entrypoint"]
CONFIG["โ๏ธ Config Loader
config.py โ .astwire.toml"]
GITCHANGES["๐ Git Change Detection
git_changes.py
--since REF"]
TARGETING["๐ฏ Target Resolution
targeting.py
gitignore.py"]
CRAWLER["๐ Import Crawler
ast_crawler.py
-i / --resolve-imports
tracks depth + edges"]
FILES["๐ Resolved Python Files
+ depth map when -i is on"]
ANALYSIS["๐ Waste Analyzer
analyzer.py
--analysis"]
SKELETON["๐ฉป Skeletonizer
skeleton.py
--skeleton or --decay-depth"]
CLEAN["๐งน Waste Cleanser
analyzer.py
--strip-waste"]
REDACT["๐ Secret Redaction
redact.py
always on"]
RECORDS["๐ฆ FileRecords
records.py"]
SPLITTER["โ๏ธ Partition Splitter
splitter.py
--max-tokens"]
TOKENIZER["๐ข Tokenizer
tokenizer.py
tiktoken / fallback"]
FORMATTER["๐ Formatter
*
markdown / llm / json
+ graph block when --graph"]
OUTPUT["๐พ Output File(s)
context.md / .xml / .json"]
CLI --> CONFIG
CONFIG -->|"--since"| GITCHANGES
CONFIG -->|"no --since"| TARGETING
GITCHANGES -->|"-i"| CRAWLER
GITCHANGES -->|"no -i"| FILES
CLI -->|"-i"| CRAWLER
CLI -->|"no -i, no --since"| TARGETING
CRAWLER --> FILES
TARGETING --> FILES
FILES -->|"--analysis"| ANALYSIS
FILES -->|"normal run"| SKELETON
SKELETON --> CLEAN
CLEAN --> REDACT
REDACT --> RECORDS
ANALYSIS --> TOKENIZER
SPLITTER --> TOKENIZER
RECORDS --> SPLITTER
SPLITTER --> FORMATTER
FORMATTER --> OUTPUT
classDef entry fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef local fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef process fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef security fill:#fff0f0,stroke:#d64545,stroke-width:2px,color:#5c2020
classDef output fill:#eeeaff,stroke:#7655c7,stroke-width:2px,color:#4a3485
class CLI,CONFIG entry
class GITCHANGES,TARGETING,CRAWLER,FILES local
class ANALYSIS,SKELETON,CLEAN,RECORDS,SPLITTER,TOKENIZER,FORMATTER process
class REDACT security
class OUTPUT output
```
## Execution Flow
1. **CLI parsing** โ `main()` in `cli.py` parses arguments with `argparse`. `--uninstall` and `-v/--version` short-circuit before any heavy module is imported (heavy imports are deferred inside `main()` specifically so `astwire --help` and `astwire --uninstall` stay fast on the compiled binary).
2. **Project root & config** โ `find_project_root()` walks upward from the first target looking for `.git`, `.astwire.toml`, `pyproject.toml`, `setup.py`, or `manage.py`. `load_config()` then reads `.astwire.toml` (if found on the same upward walk) and merges it with CLI flags, where **CLI flags always win** over config file values.
3. **Target resolution** โ exactly one of two paths runs:
- **`-i` / `--resolve-imports`**: `discover_dependencies()` seeds a BFS queue from the given file(s) (or every `.py` file under a given directory) โ or, when `--since REF` is given, from `git_changes.get_changed_files()`'s output instead โ and recursively resolves each `import`/`from ... import` statement to a local file via `resolve_import_path()`, skipping anything in `sys.stdlib_module_names` or unresolvable to a project-local path. Each discovered file's import-hop distance from its nearest seed is recorded, along with every import edge walked (including edges into files reached earlier by a different path). `--max-depth N` then drops files beyond that distance; manifests (`package.json`/`go.mod`) are exempt.
- **default**: `resolve_named_targets()` walks the given sources, filtering by `skip` patterns from config and by `IgnoreEngine` (`.gitignore` + `.astwireignore`, nested per-directory) โ or, with `--since REF` and no `-i`, the changed-file list from `git_changes.get_changed_files()` is filtered the same way instead of being walked.
- **`--since REF`** (either path): seeds come from `git diff --name-only REF` plus new untracked files under a single scope path, gitignore-aware, via `git_changes.py`. `REF` is always used exactly as given โ never resolved against a remote or auto-detected base branch. Requires the scope to be inside a git repository and `REF` to resolve to a real commit.
4. **Read & classify** โ resolved paths are deduplicated, filtered to `.py` files only (non-Python files are counted and reported, not processed), and sorted for deterministic output.
5. **Per-file transform pipeline** โ for each file, in order:
- `--skeleton`, or `--decay-depth N` where the file's recorded depth exceeds `N` โ `generate_skeleton()` parses the file's AST and replaces every function/method body with `...` (preserving a leading docstring, if any), then `ast.unparse()`s the result. Falls back to returning the original source unchanged on `SyntaxError`. (`--decay-depth` without `-i` is a no-op with a stderr note; when `--since` is used with `-i` and `--decay-depth` isn't set explicitly, it defaults to `0`.)
- `--strip-waste` โ `clean_content()` normalizes CRLF/CR to LF, strips trailing whitespace per line, collapses runs of more than one blank line, and enforces exactly one trailing newline.
- **always** โ `redact_secrets()` scans the (possibly transformed) content against four regex patterns and replaces matches with `[REDACTED:SECRET]`.
6. **Analysis mode** (`--analysis`) is a separate branch: it calls `profile_file()` for every resolved file (running `clean_content()` internally to compute recoverable waste), never invokes the formatter, and prints a table via `format_analysis_table()` instead of writing output โ the table gains a **Depth** column whenever `-i` was used. No files are written in this mode.
7. **Partitioning** โ `partition_records()` accumulates `FileRecord`s into partitions bounded by `--max-tokens` (via `count_tokens()`), never splitting a single file across two partitions. A file whose own size exceeds the budget gets an isolated partition. `--max-tokens` unset (or โค0) produces a single partition.
8. **Formatting** โ one of three `BaseFormatter` implementations renders each partition: `MarkdownFormatter` (ASCII directory tree + fenced code blocks), `LLMFormatter` (minimal ``/`` XML tags), or `JSONFormatter` (structured payload with per-file redaction counts). With `-i --graph`, each formatter also renders the import edges landing inside that partition as a graph block (`## Dependency Graph` in Markdown, `` in the XML formatter, a top-level `"graph"` array in JSON) โ an edge is only rendered when both of its endpoints are present in that same partition. Without `-i`, `--graph` is a no-op with a stderr note.
9. **Write** โ each partition is written to `