astwire / Architecture / Architecture Diagrams
DocsastwireArchitectureArchitecture Diagrams

Architecture Diagrams

Detailed flow diagrams for individual astwire subsystems. For the top-level pipeline, see astwire-architecture.md.

7 min readApplies to v1.0.0
On this page ▾
  1. 1. CLI command routing
  2. 2. Import resolution (-i / --resolve-imports)
  3. 2b. Change-scoped seeding (--since)
  4. 3. Skeleton generation (--skeleton)
  5. 4. Waste analysis & cleansing (--analysis / --strip-waste)
  6. 5. Token-bounded partitioning (--max-tokens)
  7. 6. Tokenizer fallback chain
  8. 7. Formatter dispatch
  9. 8. Ignore resolution (.gitignore + .astwireignore)

1. CLI command routing

text
                                astwire
                                   │
                                   ▼
                          argparse ArgumentParser
                                   │
                  ┌────────────────┼────────────────┐
                  │                │                 │
             -v/--version    --uninstall      [targets...] + flags
                  │                │                 │
                  ▼                ▼                 ▼
           print version    perform_uninstall()   deferred imports
           exit(0)          exit(0/1)                   │
                                                          ▼
                                                  --since given?
                                                          │
                                          ┌───────────────┴───────────────┐
                                          │                                │
                                         yes                               no
                                          │                                │
                                          ▼                                │
                              git_changes.get_changed_files()              │
                              (exit 2: bad scope/ref;                      │
                               exit 1: nothing changed)                    │
                                          │                                │
                                          └───────────────┬────────────────┘
                                                           ▼
                                                  full pipeline (below)
                                       (-i also on: --max-depth / --decay-depth
                                        prune or skeletonize by hop-distance;
                                        --graph renders import edges)

2. Import resolution (-i / --resolve-imports)

find_project_root() (used both to bound resolution and as the config search root) walks upward from the entry path looking for the first of: .git, .astwire.toml, pyproject.toml, setup.py, manage.py. If none is found before the filesystem root, it falls back to the entry path itself (or its parent, if the entry is a file).

Seeds normally come from the entry target(s) above. With --since REF, they come from git_changes.get_changed_files() instead (see §2b below) — everything downstream of seeding (BFS, depth/edge tracking, --max-depth, manifests) is unchanged either way.

--decay-depth N doesn't change this diagram's output — it's consumed later, per file, in the transform pipeline: any file whose recorded depth > N gets skeletonized even without --skeleton. --graph renders edges (filtered to whichever files land in the same output partition) as a block in the chosen output format; see output-formats.md.

2b. Change-scoped seeding (--since)

REF is always used exactly as supplied — this flow never inspects remotes, never falls back to a default branch, and never resolves origin/main vs origin/master on your behalf. When -i is also on and --decay-depth isn't set explicitly, it defaults to 0 (the changed files stay full text; anything pulled in via import gets skeletonized).

3. Skeleton generation (--skeleton)

Type hints, decorators, class attributes, and docstrings survive untouched — only executable statement bodies inside function/method definitions are replaced. This same transformer runs whenever --skeleton is set or a file's recorded import-depth exceeds --decay-depth N (§2) — the trigger differs, the transformation is identical.

4. Waste analysis & cleansing (--analysis / --strip-waste)

--analysis and --strip-waste share the same clean_content() function but serve different purposes: --analysis only reports recoverable waste (dry run, no output file); --strip-waste actually applies the cleansing to the content that gets written.

5. Token-bounded partitioning (--max-tokens)

A single file whose own token count exceeds max_tokens is never split — it becomes the sole member of its own oversized partition (the "batch non-empty AND over budget" check only closes a batch that already has other content).

6. Tokenizer fallback chain

text
count_tokens(text)
        │
        ▼
tiktoken installed?
        │
   ┌────┴────┐
   │         │
  yes        no
   │         │
   ▼         ▼
cl100k_base   fallback heuristic
 .encode()    len(text) / 3.7
   │             (chars-per-token,
   │              calibrated for
   ▼              Python source)
count returned

tiktoken is an optional dependency — the compiled binary works with or without it, falling back to the calibrated character-ratio heuristic when it isn't bundled or fails to load.

7. Formatter dispatch

text
partition_records() output
        │
        ▼
  get_formatter(format_choice)
        │
   ┌────┼────────────┬─────────────┐
   │                 │             │
"markdown"          "llm"        "json"
 (default)            │             │
   │                  │             │
   ▼                  ▼             ▼
MarkdownFormatter  LLMFormatter  JSONFormatter
  │                  │             │
  ▼                  ▼             ▼
ASCII tree +      <context>      {"part_index",
fenced code       <file path=…>   "total_parts",
blocks per file   …</file>        "graph": [...],
                  </context>      "files": [...]}

With -i --graph, each branch also renders edges (from §2) as a block ahead of the file contents — a ## Dependency Graph bullet list in Markdown, a <graph><edge from=… to=…/></graph> block in the XML formatter (before <files>), and the "graph" array shown above in JSON (always present; [] without --graph). An edge only appears if both of its endpoints landed in that same partition — edges crossing a --max-tokens partition boundary are dropped rather than left dangling.

8. Ignore resolution (.gitignore + .astwireignore)

text
IgnoreEngine(root)
        │
        ▼
is_ignored(path)
        │
        ▼
walk ancestor directories
  from path's parent up to root
  (root → ... → path.parent, in that order)
        │
        ▼
for each ancestor directory:
  load & cache .gitignore lines
  load & cache .astwireignore lines
  combine into one pathspec.PathSpec (gitignore syntax)
        │
        ▼
match relative path against each
directory's spec — first match wins
        │
        ▼
   ignored / not ignored

Specs are cached per-directory (_spec_cache) so repeated is_ignored() calls across many files in the same tree don't re-read .gitignore/.astwireignore from disk.