djangoplay-web / Apps / aicore — In-App AI Assistant
DocsDjangoPlay WebAppsaicore — In-App AI Assistant

aicore — In-App AI Assistant

aicore is a multi-provider, streaming (SSE) chat assistant exposed to authenticated DjangoPlay users. It is the product UI chatbot (a floating widget on the main site), distinct from the separate d...

17 min readApplies to v1.2.2Added in 1.2.2
On this page ▾
  1. 1. Summary
  2. 2. Architecture
  3. 2.1 Module layout
  4. 2.2 Request flow (streaming chat)
  5. 2.3 Provider abstraction
  6. 2.4 Free-model catalog (services/model_catalog.py)
  7. 2.5 BYOK (services/byok.py)
  8. 2.6 Conversation history summarization
  9. 3. Data Model
  10. ChatSession (aicore/models/chat.py)
  11. ChatMessage
  12. ChatUsageEvent (new)
  13. 3.1 Inherited lifecycle behavior (core.models.TimeStampedModel)
  14. 4. Service Layer
  15. ChatService (aicore/services/chat_service.py) — static/classmethods only
  16. Rate limiting & budgets (services/rate_limit.py)
  17. System prompt loading (services/knowledge.py)
  18. Usage telemetry (services/usage_stats.py)
  19. 5. HTTP API
  20. 5.1 POST /api/v1/ai/chat/stream/ — example stream
  21. 5.2 Admin usage dashboard (not a public API — /admin/aicore/chatusageevent/dashboard/)
  22. 6. Configuration (paystream/app_settings/ai.py)
  23. DjangoPlay AI Assistant (aicore)
  24. Quick local setup
  25. 7. Security & Failure Modes
  26. 8. Integration Points
  27. 9. Known Gaps / Notes for Future Work

Verbose name: AI Core (Django app label) · displayed as AI Assistant in the admin dashboard UI only (display-label override, not an app rename) · registered as "aicore.apps.AicoreConfig" in INSTALLED_APPS

Current as of v1.2.2. Design docs for each change in this release live in the AI Core section — this page is the living reference, updated as the app evolves.

1. Summary

Core design idea: the app never hardcodes a single LLM vendor or a single key. A small provider registry lets DjangoPlay talk to xAI (Grok), OpenAI, any OpenAI-compatible endpoint (Ollama, vLLM, LiteLLM, etc., via AI_PROVIDER=custom — local Ollama is the recommended zero-cost default), a free-tier OpenRouter model catalog the user picks from per conversation, or the user's own bring-your-own-key (BYOK) credentials supplied for a single browser tab. Provider/model selection happens per-session, not just per-deployment.

It owns three models (ChatSession, ChatMessage, ChatUsageEvent), a service layer (ChatService, rate limiting + tiered token budgets, catalog/BYOK services, usage telemetry, history summarization), a provider abstraction, six endpoints under /api/v1/ai/, and an admin-side usage dashboard.


2. Architecture

2.1 Module layout

plaintext
aicore/
├── admin.py                  # ChatSession/ChatMessage/ChatUsageEvent admin + usage dashboard view
├── apps.py                   # AicoreConfig
├── constants.py               # MessageRole, ProviderName, prompt/limits constants
├── exceptions.py               # AICoreError hierarchy (incl. AIByokError)
├── knowledge/
│   └── platform.md             # System prompt / knowledge base fed to the LLM
├── migrations/                 # 0001_initial … 0010_alter_chatmessage_options… (see §3)
├── models/
│   ├── __init__.py
│   └── chat.py                  # ChatSession, ChatMessage, ChatUsageEvent
├── providers/
│   ├── base.py                   # ChatProvider ABC, ChatMessage/StreamChunk dataclasses
│   ├── openai_compat.py          # OpenAICompatProvider — used for every provider name
│   └── registry.py               # resolve_provider_name() / get_provider(name)
├── services/
│   ├── access.py                  # ai_access_for() / user_ai_enabled() — role gating
│   ├── byok.py                    # decode/validate/test BYOK credentials, ephemeral provider
│   ├── chat_service.py            # ChatService — session + orchestration
│   ├── context_gate.py            # decides whether a message needs the full knowledge base
│   ├── knowledge.py               # load_system_prompt()
│   ├── model_catalog.py           # OpenRouter free-model catalog, cached
│   ├── rate_limit.py              # RPM/RPD, per-tab BYOK throttle, per-tier token budgets
│   ├── status.py                  # GET /status/ payload
│   └── usage_stats.py             # usage-dashboard aggregation queries
├── tasks.py                    # update_history_summary (Celery)
├── urls.py
└── views/
    ├── byok.py                    # POST /byok/validate/
    ├── chat.py                    # ChatStreamView + session list/detail
    ├── models.py                  # GET /models/ (free-catalog)
    └── status.py                  # GET /status/

2.2 Request flow (streaming chat)

plaintext
Browser (chat widget)
   │  POST /api/v1/ai/chat/stream/  { message, session_id?, model_id? }
   │  optional headers: X-AI-BYOK-Token, X-AI-Tab-Id
   ▼
ChatStreamView (login_required)
   │  parses JSON body + optional BYOK header, delegates entirely to the service layer
   ▼
ChatService.stream_reply()
   ├─ 1. truncate message to MAX_USER_MESSAGE_CHARS (4000)
   ├─ 2. ai_access_for(user)                 → gate (AI_ENABLED / SSO rollout)
   ├─ 3. check_request_limits()              → account-level RPM/RPD
   ├─ 4. check_tab_request_limit()           → additional per-tab throttle, BYOK only
   ├─ 5. get_or_create_session()             → reuses active session, or creates one
   │       (model_id validated against the free catalog if provided; BYOK sessions
   │        persist only the non-secret model name)
   ├─ 6. persist user ChatMessage            → inside transaction.atomic()
   ├─ 7. yield StreamEvent(type="session")   → SSE "session" event (id/provider/model)
   ├─ 8. list_history()                      → recent window + prepended history_summary
   ├─ 9. resolve provider:
   │       • BYOK       → build_ephemeral_provider(credentials), used once, discarded
   │       • otherwise  → get_provider(session.provider) from the registry
   ├─ 10. token_budget_tier_for(session.provider) → platform / openrouter / byok
   ├─ 11. check_token_budget(tier=...)       → tier-specific monthly ceiling
   ├─ 12. provider.stream_chat(...)          → yields StreamChunk per SSE token from upstream
   │        each chunk → yield StreamEvent(type="token")
   ├─ 13. persist assistant ChatMessage (full concatenated text)
   ├─ 14. record_tokens_used(tier=...) / record_usage_event(...)
   ├─ 15. trigger update_history_summary.delay(session_id) (skipped for BYOK sessions)
   └─ 16. yield StreamEvent(type="done")
   ▼
views/chat.py `_sse()` formats each StreamEvent as
`event: <type>\ndata: <json>\n\n` → StreamingHttpResponse
(Content-Type: text/event-stream, Cache-Control: no-cache, X-Accel-Buffering: no)

Event types on the wire: session (once) → token (many) → done (final full content) — or error at any point (access denied, rate/budget limit, provider failure, empty message). Errors are delivered inside the SSE stream, not as an HTTP error status — the one exception is a malformed/missing JSON body or missing message field, which short-circuits before streaming starts and returns a plain 400 JsonResponse.

2.3 Provider abstraction

plaintext
ChatProvider (ABC, providers/base.py)
   └── stream_chat(messages, model, temperature, max_tokens) -> Iterator[StreamChunk]

OpenAICompatProvider (providers/openai_compat.py)
   — the single concrete implementation, used for EVERY provider name below,
     including BYOK and OpenRouter — none of them need a bespoke client since
     they're all OpenAI-compatible /chat/completions endpoints
   — POSTs to {base_url}/chat/completions with stream=True
   — parses `data: {...}` SSE lines from the upstream API, stops on `data: [DONE]`

get_provider() (providers/registry.py) turns a provider name into a configured OpenAICompatProvider. resolve_provider_name()'s default fallback chain is xai -> custom -> openai — openrouter and byok are deliberately excluded from it; they're only ever selected explicitly (a user picking a catalog model, or a live BYOK header), never silently used as the platform default.

AI_PROVIDER / session provider Base URL default API key setting Model default Notes
xai https://api.x.ai/v1 AI_XAI_API_KEY (env XAI_API_KEY) grok-4.3 Platform tier
openai https://api.openai.com/v1 AI_OPENAI_API_KEY (env OPENAI_API_KEY) gpt-4.1-mini Platform tier
custom AI_CUSTOM_BASE_URL (default http://127.0.0.1:11434/v1, i.e. Ollama) AI_CUSTOM_API_KEY (default "ollama") AI_CUSTOM_MODEL Platform tier. Note: "custom" is a bare string handled by the registry, not a ProviderName.TextChoices member — works fine since Django doesn't enforce CharField(choices=...) at save() time, just a quirk worth knowing if you're grepping for it.
openrouter AI_OPENROUTER_BASE_URL (default https://openrouter.ai/api/v1) AI_OPENROUTER_API_KEY (env OPENROUTER_API_KEY) User-selected from the free catalog Own budget tier — see §6
byok (marker only) user-supplied, per request user-supplied, per request, never persisted user-supplied Never resolvable from settings — registry.get_provider("byok") raises defensively; ChatService must branch to build_ephemeral_provider() instead. Own budget tier.

Because OpenAICompatProvider.__init__ raises AIConfigError if api_key or base_url is empty, a misconfigured provider fails fast at request time — surfaces as an error SSE event, never a silent no-op or an unhandled 500.

2.4 Free-model catalog (services/model_catalog.py)

  • get_free_models() fetches GET {AI_OPENROUTER_BASE_URL}/models, filters to free (pricing.prompt == 0 and pricing.completion == 0) and text-only ("text" in input_modalities, output_modalities == ["text"]) models.
  • Result cached server-side (aicore:model_catalog:free, TTL AI_MODEL_CATALOG_CACHE_TTL, default 30 min) — never re-fetched from OpenRouter on every dropdown open. Falls back to a much longer-lived stale cache on upstream failure; returns an empty list with a short error string rather than crashing the widget.
  • Optional admin curation on top: AI_MODEL_CATALOG_ALLOWLIST (wins if non-empty) or AI_MODEL_CATALOG_BLOCKLIST.
  • is_known_free_model(model_id) validates any client-supplied model id against the current catalog before ChatService.get_or_create_session trusts it.

2.5 BYOK (services/byok.py)

Two independent paths, covering two very different threat models — see the dedicated production BYOK and local dev BYOK docs for the full design reasoning:

  • Production, session-only: BYOKCredentials (base_url, model, api_key) decoded fresh from the X-AI-BYOK-Token header on every chat request — base64-encoded JSON, never cached, never logged, never written to ChatSession. build_ephemeral_provider() constructs an OpenAICompatProvider for exactly one request. test_credentials() powers POST /aicore/byok/validate/, the modal's pre-acceptance test call — calls the provider directly with requests (not stream_chat) so it can return a field-attributable message (bad key / bad URL / rate-limited / server error) instead of a bare status code, and tolerates both the standard OpenAI error shape and Gemini's non-standard one.
  • Local development: ~/.dplay/.secrets can define up to two named provider profiles (AI_LOCAL_PROFILE_<name>__<SETTING>), switched with a one-line AI_LOCAL_ACTIVE_PROFILE edit. Gated on DJANGO_SETTINGS_MODULE == "paystream.settings.dev" (not DEBUG, since this loads before DEBUG is resolved) — never read under staging/prod settings modules.

2.6 Conversation history summarization

ChatSession.history_summary / summarized_message_count track a running summary of everything older than the most recent AI_HISTORY_WINDOW_SIZE (default 20) messages. aicore/tasks.py::update_history_summary (Celery, fire-and-forget after each reply) folds in newly-eligible older messages once enough have piled up (AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES, default 10), using the platform's default provider — never for BYOK sessions, since there's no user key available in a background task. list_history() prepends the summary (if any) ahead of the verbatim recent window.


3. Data Model

ChatSession (aicore/models/chat.py)

Extends core.models.TimeStampedModel — gets created_at, updated_at, deleted_at, is_active, soft_delete(), restore() for free (see §3.1).

Field Type Notes
id UUIDField (PK) default=uuid.uuid4
user FK → AUTH_USER_MODEL on_delete=CASCADE, related_name="ai_chat_sessions"
title CharField(200) Auto-set to first 80 chars of the first user message
provider CharField choices=ProviderName xai / openai / openrouter / byok, or the unlisted "custom" string (see §2.3)
model_name CharField(128) Resolved once per session
history_summary TextField, blank Running summary of older messages; capped at AI_HISTORY_SUMMARY_MAX_CHARS
summarized_message_count PositiveIntegerField, default 0 How many older messages are already folded into history_summary

Meta: ordering = ("-updated_at",), proper verbose_name/verbose_name_plural (fixed in 0010_…, previously rendered lowercase in the admin), indexed on (user, -updated_at).

ChatMessage

Field Type Notes
id UUIDField (PK)
session FK → ChatSession related_name="messages", CASCADE
role CharField choices=MessageRole (system/user/assistant)
content TextField
token_estimate PositiveIntegerField, nullable Still present in schema but not populated anywhere — see §9

Meta: ordering = ("created_at",), indexed on (session, created_at).

ChatUsageEvent (new)

Telemetry row written best-effort at each meaningful chat outcome — never for requests rejected before a session/provider is resolved (access/rate-limit rejections aren't model-usage events).

Field Type Notes
provider, model_name CharField What actually served the request
status choices success / error
error_message CharField(500), blank Full message (see §5's dashboard note on why this isn't truncated in the UI)
session, user FK

Schema went through one revision mid-flight (0008_delete_chatusageevent / 0009_chatusageevent) — if you're reading migration history, that's a deliberate drop + recreate, not two independent features.

3.1 Inherited lifecycle behavior (core.models.TimeStampedModel)

All three models get soft-delete semantics for free, though aicore itself never calls .soft_delete() — inherited infrastructure, not domain logic exercised here. aicore is not in AUDIT_TRACKED_MODELS (see §7).


4. Service Layer

ChatService (aicore/services/chat_service.py) — static/classmethods only

Method Responsibility
get_or_create_session Reuses an active session owned by user, or creates one (validates any client-supplied model_id against the free catalog first)
token_budget_tier_for(provider) Maps a provider to its budget tier: byok / openrouter / platform
list_history(session) Recent-window messages, with history_summary prepended if present
stream_reply Full orchestration — see §2.2

StreamEvent — type ∈ {session, token, done, error}.

Rate limiting & budgets (services/rate_limit.py)

All cache-backed (Redis in prod), fixed-window, explicitly a soft guardrail (worst case on the create/incr race: one extra request slips through) rather than a hard atomic limiter — acceptable for a demo/portfolio app, worth knowing if reused elsewhere.

Function Scope Key shape
check_request_limits() Account-level RPM + RPD aicore:rl:rpm:{user_id}:{minute} / ...rpd...:{date}
check_tab_request_limit() Additional per-browser-tab throttle, BYOK only aicore:rl:byoktab:{user_id}:{tab_id}:{minute}
check_token_budget(tier=...) / record_tokens_used(tier=...) Monthly token ceiling, partitioned by tier aicore:rl:tokmo:{tier}:{user_id}:{month}

A missing/malformed X-AI-Tab-Id never blocks a request — the per-tab check is simply skipped; the account-level limits and token budget still apply regardless.

System prompt loading (services/knowledge.py)

load_system_prompt(name="platform"), @lru_cache(maxsize=4): AI_SYSTEM_PROMPT override → aicore/knowledge/{name}.md → hardcoded fallback. knowledge/platform.md explicitly instructs the assistant to never invent endpoints/settings/models, never handle secrets, and frames the product as a demo/learning/design reference.

Usage telemetry (services/usage_stats.py)

record_usage_event() (best-effort, exceptions swallowed — telemetry must never break a chat response), get_usage_summary(since, user) (aggregates by (provider, model_name); non-admins see only their own usage), get_recent_errors(), usage_window_start(days) for the dashboard's window selector.


5. HTTP API

Mounted at path("api/v1/ai/", include(("aicore.urls", "aicore"), namespace="aicore")).

Method Path View Notes
POST /api/v1/ai/chat/stream/ ChatStreamView SSE stream; body {"message", "session_id"?, "model_id"?}; optional X-AI-BYOK-Token / X-AI-Tab-Id headers
POST /api/v1/ai/byok/validate/ byok_validate One-off test call for the BYOK modal; body read once, discarded
GET /api/v1/ai/chat/sessions/ chat_session_list Last 20 active sessions for request.user
GET /api/v1/ai/chat/sessions/<uuid>/messages/ chat_session_messages Full history for one session (404 if not owned/found)
GET /api/v1/ai/status/ ai_status Widget bootstrap payload (AI_ENABLED, access, etc.)
GET /api/v1/ai/models/ ai_model_catalog Cached free OpenRouter catalog

All plain Django views (login_required, hand-built JsonResponse), not DRF — deliberate, since SSE streaming doesn't fit DRF's renderer model cleanly. Means this app doesn't get automatic OpenAPI schema generation via drf-spectacular.

5.1 POST /api/v1/ai/chat/stream/ — example stream

plaintext
event: session
data: {"session_id": "5b1e...", "provider": "openrouter", "model": "meta-llama/llama-3.2-3b-instruct:free"}

event: token
data: {"text": "DjangoPlay"}

event: done
data: {"session_id": "5b1e...", "content": "DjangoPlay tracks changes via..."}

On failure:

plaintext
event: error
data: {"message": "Monthly AI usage budget for your account has been reached. Contact an admin if you need more."}

5.2 Admin usage dashboard (not a public API — /admin/aicore/chatusageevent/dashboard/)

Provider/model breakdown (requests, successes, errors, error rate, last-used — shown with an explicit timezone abbreviation, e.g. 2026-08-24 13:05 UTC) and a recent-errors table with the full error message (copy-to-clipboard button) rather than a truncated tooltip. ChatSession/ChatMessage are excluded from generic admin CRUD discovery (configs/admin_registry.json) — this dashboard, backed by ChatUsageEvent, is the intended admin surface for chat activity, not raw transcript browsing.


6. Configuration (paystream/app_settings/ai.py)

All values loaded via get_decrypted_value() (env var first, then ~/.dplay/.secrets under local dev, then Fernet-encrypted .env).

Setting Default Purpose
AI_ENABLED / AI_ENABLED_FOR_SSO_USERS true / — Master switch + SSO-rollout gate (ai_access_for())
AI_PROVIDER custom Platform default: xai | openai | custom
AI_XAI_* / AI_OPENAI_* / AI_CUSTOM_* see §2.3 Platform-tier provider config
AI_OPENROUTER_API_KEY (env OPENROUTER_API_KEY) / AI_OPENROUTER_BASE_URL / AI_OPENROUTER_MODEL — / https://openrouter.ai/api/v1 / — Free-catalog tier. AI_OPENROUTER_MODEL is only a fallback — real model IDs come from the catalog
AI_MODEL_CATALOG_CACHE_TTL / _ALLOWLIST / _BLOCKLIST 1800 / — / — Catalog caching + admin curation
AI_TEMPERATURE / AI_MAX_TOKENS / AI_REQUEST_TIMEOUT 0.4 / 2048 / 120s Passed to stream_chat()
AI_CHAT_RATE_LIMIT_PER_MINUTE / _PER_DAY 5 / — Account-level RPM/RPD
AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE 10 Additional per-tab throttle, BYOK only — deliberately looser than the account RPM
AI_CHAT_TOKEN_BUDGET_PER_MONTH 50000 Platform-tier (xai/openai/custom) monthly token budget
AI_CHAT_TOKEN_BUDGET_OPENROUTER_PER_MONTH 50000 Free-catalog tier
AI_CHAT_TOKEN_BUDGET_BYOK_PER_MONTH higher than the other two Pure abuse backstop, not a cost control — BYOK tokens are billed to the user's own key
AI_BYOK_ENABLED true Independent BYOK kill switch
AI_BYOK_REQUEST_TIMEOUT / _VALIDATION_TIMEOUT — Ephemeral-provider vs. modal-test-call timeouts
AI_BYOK_MAX_BASE_URL_LENGTH / _MAX_MODEL_LENGTH / _MAX_API_KEY_LENGTH — Sanity caps on the per-request header fields
AI_HISTORY_WINDOW_SIZE 20 Verbatim recent-message window
AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES 10 Older messages must pile up this much before another summarization LLM call runs
AI_HISTORY_SUMMARY_MAX_CHARS 2000 Cap on history_summary
AI_SYSTEM_PROMPT "" Full override of knowledge/platform.md

Recommended free/local setup:

DjangoPlay AI Assistant (aicore)

DjangoPlay includes a multi-provider, streaming chat assistant for authenticated users — business logic stays in services, API keys stay server-side (or never touch the server at all, for BYOK), and the LLM backend is swappable per-session, not just per-deployment.

plaintext
Browser chat widget (authenticated)
   │  POST /api/v1/ai/chat/stream/  (SSE)
   ▼
aicore (ChatService)  →  rate limits + tiered monthly token budgets → history (summarized)
   ▼
Provider registry  →  xAI | OpenAI | free OpenRouter catalog | bring-your-own-key | custom (Ollama / vLLM / …)

Highlights: a free-model picker sourced live from OpenRouter's catalog, bring-your-own-key for both local dev (two switchable profiles) and production (session/tab-scoped, key never persisted), an admin usage dashboard, and usage guardrails (per-tab throttling, per-provider-tier token budgets) that scale with the added flexibility.

Quick local setup

For example, Ollama at zero cost

bash
# Install from https://ollama.com , then:
ollama pull llama3.2:3b   # name must match `ollama list` exactly

AI_ENABLED=true
AI_PROVIDER=custom
AI_CUSTOM_BASE_URL=http://127.0.0.1:11434/v1
AI_CUSTOM_API_KEY=ollama
AI_CUSTOM_MODEL=llama3.2:3b
bash
cd webapp
python manage.py migrate aicore
# restart Django, sign in → robot button bottom-right

This runs fully local — no per-token bill unless you point AI_PROVIDER at a paid cloud API, or a user opts into the free OpenRouter catalog / brings their own key from the chat widget.

Full documentation — architecture, request flow, every setting, the BYOK security model, rate-limit/token-budget design, and the usage dashboard — lives in the docs repo, not here:

App reference docs.djangoplay.org → apps/aicore
v1.2.2 feature docs (model catalog, local/production BYOK, usage controls) docs.djangoplay.org → apps/aicore

Want two personal keys locally, switchable without editing .env? See Local Dev BYOK.


7. Security & Failure Modes

  • AuthN: every endpoint requires an authenticated session (login_required); no anonymous/API-token access path.
  • Ownership isolation: session-list/message-history queries filter by user=request.user — 404 (not 403) on mismatch, doesn't leak existence.
  • Input bounds: user message hard-capped at 4000 chars; empty/whitespace-only rejected.
  • Layered rate limiting: account-level RPM/RPD apply to every provider including BYOK (explicitly kept as the platform's only backstop against being used as a free relay, even though BYOK tokens are billed to the user's own key); BYOK additionally gets a per-tab throttle; monthly token budgets are partitioned per tier so one tier running hot doesn't starve the others.
  • BYOK key never persisted: lives only in the browser tab's sessionStorage, sent as a header (never the request body), decoded per-request, discarded after use. Verified against APIRequestLoggingMiddleware, which only ever persists path/method/status/user-agent, never headers or bodies. A ChatSession with provider="byok" stores only the non-secret model name.
  • Fail-closed provider config: OpenAICompatProvider refuses to construct without both an API key and base URL — surfaces as a caught AIConfigError → SSE error event, never a silent no-op or unhandled 500.
  • Upstream failures: normalized into AIProviderError; raw response body logged server-side (truncated to 500 chars), never returned to the client.
  • No secrets exposure via the model: knowledge/platform.md instructs the model to never request or handle secrets — a prompt-level guard, advisory rather than a hard boundary.
  • Not in AUDIT_TRACKED_MODELS: chat history isn't captured in the platform's append-only audit trail the way e.g. invoices.Invoice edits are — unchanged from before this release.
  • Admin surface restricted: ChatSession/ChatMessage are excluded from generic admin CRUD discovery (configs/admin_registry.json); the usage dashboard (ChatUsageEvent-backed) is the intended admin-facing surface.

8. Integration Points

  • core.models.TimeStampedModel — soft-delete/restore/is_active plumbing (§3.1).
  • audit.signals.events — indirectly, via inherited signals (unused in practice).
  • Celery — aicore.tasks.update_history_summary (fire-and-forget history summarization, §2.6) is the app's first real use of the platform's existing Celery wiring.
  • Frontend widget — floating chat drawer + model dropdown + BYOK modal, gated by AI_ENABLED, in webapp/frontend/templates/base/chat_widget.html / frontend/static/assets/js/components/chat_widget.js. Widget markup/behavior itself belongs to the frontend app's documentation, not here.
  • External LLM providers — xAI, OpenAI, OpenRouter (free catalog), any self-hosted OpenAI-compatible server, or a user's own BYOK endpoint.
  • shared/admin — MODEL_ICON_MAP / APP_ICON_MAP (distinct icons per model, an aicore app icon that was previously missing entirely) and shared/ui/dashboard/services/dashboard_stats_service.py's DASHBOARD_APP_LABEL_OVERRIDES (the "AI Assistant" display label, UI-only).

9. Known Gaps / Notes for Future Work

  • ChatMessage.token_estimate is still defined in the schema but never populated — no per-message token-counting logic exists (input-token estimation for budget/rate-limit purposes happens separately, at the ChatService level, not per-stored-message).
  • Rate limiting and token budgets are fixed-window, not sliding/token-bucket — minor over-allowance possible at window boundaries; low-risk given the default limits.
  • Per-tier (platform/openrouter/byok) usage is trackable but not yet broken out on the admin usage dashboard — get_usage_summary() aggregates by (provider, model_name), which is enough to derive tier totals manually today. A dedicated tier rollup on the dashboard itself would be a small follow-up.
  • aicore models remain excluded from AUDIT_TRACKED_MODELS — worth periodically confirming this is still the intended posture as chat usage grows.
  • Views are plain Django, not DRF — won't appear in Swagger/ReDoc without manual documentation (see apidocs app).
  • "custom" as a provider value isn't a member of ProviderName.TextChoices (see §2.3) — works today because Django doesn't enforce choices= at save() time, but it's an inconsistency worth cleaning up if ProviderName is ever iterated over exhaustively somewhere new.