astwire / Architecture / astwire — Architecture
DocsastwireArchitectureastwire — Architecture

astwire — Architecture

astwire is a local-first CLI application that turns a Python codebase into a clean, partitionable LLM context bundle. It separates target resolution (which files matter), AST transformation (import...

7 min readApplies to v1.0.0
On this page ▾
  1. Architecture
  2. Execution Flow
  3. Main Components
  4. Design Boundaries
  5. Configuration Precedence
  6. Related Documents

Architecture

Execution Flow

  1. CLI parsing — main() in cli.py parses arguments with argparse. --uninstall and -v/--version short-circuit before any heavy module is imported (heavy imports are deferred inside main() specifically so astwire --help and astwire --uninstall stay fast on the compiled binary).
  2. Project root & config — find_project_root() walks upward from the first target looking for .git, .astwire.toml, pyproject.toml, setup.py, or manage.py. load_config() then reads .astwire.toml (if found on the same upward walk) and merges it with CLI flags, where CLI flags always win over config file values.
  3. Target resolution — exactly one of two paths runs:
    • -i / --resolve-imports: discover_dependencies() seeds a BFS queue from the given file(s) (or every .py file under a given directory) — or, when --since REF is given, from git_changes.get_changed_files()'s output instead — and recursively resolves each import/from ... import statement to a local file via resolve_import_path(), skipping anything in sys.stdlib_module_names or unresolvable to a project-local path. Each discovered file's import-hop distance from its nearest seed is recorded, along with every import edge walked (including edges into files reached earlier by a different path). --max-depth N then drops files beyond that distance; manifests (package.json/go.mod) are exempt.
    • default: resolve_named_targets() walks the given sources, filtering by skip patterns from config and by IgnoreEngine (.gitignore + .astwireignore, nested per-directory) — or, with --since REF and no -i, the changed-file list from git_changes.get_changed_files() is filtered the same way instead of being walked.
    • --since REF (either path): seeds come from git diff --name-only REF plus new untracked files under a single scope path, gitignore-aware, via git_changes.py. REF is always used exactly as given — never resolved against a remote or auto-detected base branch. Requires the scope to be inside a git repository and REF to resolve to a real commit.
  4. Read & classify — resolved paths are deduplicated, filtered to .py files only (non-Python files are counted and reported, not processed), and sorted for deterministic output.
  5. Per-file transform pipeline — for each file, in order:
    • --skeleton, or --decay-depth N where the file's recorded depth exceeds N → generate_skeleton() parses the file's AST and replaces every function/method body with ... (preserving a leading docstring, if any), then ast.unparse()s the result. Falls back to returning the original source unchanged on SyntaxError. (--decay-depth without -i is a no-op with a stderr note; when --since is used with -i and --decay-depth isn't set explicitly, it defaults to 0.)
    • --strip-waste → clean_content() normalizes CRLF/CR to LF, strips trailing whitespace per line, collapses runs of more than one blank line, and enforces exactly one trailing newline.
    • always → redact_secrets() scans the (possibly transformed) content against four regex patterns and replaces matches with [REDACTED:SECRET].
  6. Analysis mode (--analysis) is a separate branch: it calls profile_file() for every resolved file (running clean_content() internally to compute recoverable waste), never invokes the formatter, and prints a table via format_analysis_table() instead of writing output — the table gains a Depth column whenever -i was used. No files are written in this mode.
  7. Partitioning — partition_records() accumulates FileRecords into partitions bounded by --max-tokens (via count_tokens()), never splitting a single file across two partitions. A file whose own size exceeds the budget gets an isolated partition. --max-tokens unset (or ≤0) produces a single partition.
  8. Formatting — one of three BaseFormatter implementations renders each partition: MarkdownFormatter (ASCII directory tree + fenced code blocks), LLMFormatter (minimal <context>/<file> XML tags), or JSONFormatter (structured payload with per-file redaction counts). With -i --graph, each formatter also renders the import edges landing inside that partition as a graph block (## Dependency Graph in Markdown, <graph> in the XML formatter, a top-level "graph" array in JSON) — an edge is only rendered when both of its endpoints are present in that same partition. Without -i, --graph is a no-op with a stderr note.
  9. Write — each partition is written to <output stem><suffix><output ext> (suffix is _part_N when there is more than one partition), and a one-line summary (file count, ~tokens, active flags) is printed per partition to stdout.

Main Components

Component Responsibility
cli.py Argument parsing, uninstall handling, pipeline orchestration
config.py .astwire.toml discovery (upward search) and merge into astwireConfig
targeting.py Directory-walk target resolution against skip patterns
gitignore.py Nested .gitignore / .astwireignore matching via pathspec
records.py FileRecord dataclass — the unit passed between every downstream stage
ast_crawler.py AST-based import discovery and project-root detection (-i)
git_changes.py Git-diff-based change detection for --since — changed-file listing, git-root/ref checks, gitignore-aware filtering
skeleton.py ast.NodeTransformer that strips function/method bodies to ...
analyzer.py Whitespace cleansing (clean_content) and token-waste profiling
splitter.py Token-bounded partitioning that preserves file boundaries
redact.py Regex-based secret detection and redaction
tokenizer.py tiktoken (cl100k_base) token counting with a calibrated character-ratio fallback
base.py BaseFormatter ABC and ASCII directory-tree generator
markdown.py Markdown output with directory tree + fenced code blocks
llm.py Minimal XML-style output tuned for LLM prompt templates
json_fmt.py Structured JSON output with partition and redaction metadata

Design Boundaries

  • Static analysis only. Every transformation (ast_crawler, skeleton) operates on ast.parse() output. astwire never exec()s, eval()s, or imports the code it analyzes.
  • File boundaries are never split. The partition splitter (splitter.py) treats each file as an atomic unit; a partition either contains a whole file or the file gets its own oversized partition.
  • Redaction is best-effort, not a secret scanner. redact_secrets() matches four specific patterns (AWS keys, generic key/secret/token/password = assignments, bearer tokens, PEM private-key headers). See security-and-privacy.md for what this does and does not catch.
  • CLI flags always override .astwire.toml. load_config() supplies defaults; main() applies args.X or cfg.X for every overlapping setting.
  • Heavy imports are deferred. config, *, *, *, and redact are imported inside main(), after argument parsing, specifically so --help, -v, and --uninstall remain fast when astwire runs as a compiled Nuitka binary.
  • --since never guesses. git_changes.py takes REF exactly as given — it never auto-detects a base branch or remote, makes no AI calls, and caches nothing across invocations; every run re-derives its answer from git. It is a deliberate fork of a narrow slice of gitaiflow's git-plumbing rather than a dependency on it, so astwire does not automatically inherit fixes made there.

Configuration Precedence

text
CLI flags (-i, --skeleton, --strip-waste, --max-tokens, -f, -o)
        │
        │  wins when both are set
        ▼
.astwire.toml (general / targeting / ai sections)
        │
        │  falls back when neither is set
        ▼
astwireConfig defaults (markdown, no skeleton, no strip, show_tree=True)

For the full .astwire.toml schema, see api/configuration.md.