--- since: 1.2.2 --- # Usage Controls & Platform Hardening Depends on: free-model-catalog-and-selection.md (provider tiers), production-byok-session-only.md (BYOK header flow, per-request ephemeral provider) Covers four changes, built in this order: conversation history summarization, the admin AI usage dashboard, BYOK per-tab throttling, and per-tier token budgets. --- ## 1. Conversation History Summarization ### Goal Replace the flat sliding-window history (`MAX_HISTORY_MESSAGES`, no compression) with something that retains earlier context in long conversations without sending the entire transcript on every request. ### Design - `ChatSession` gains two fields: `history_summary` (text, running summary) and `summarized_message_count` (how many older messages have already been folded into it). - `ChatService.list_history` keeps a bounded recent-messages window (`AI_HISTORY_WINDOW_SIZE`, default 20) verbatim, and — if `session.history_summary` is non-empty — prepends it as a labeled "already condensed, older than the recent window" system-style message ahead of the recent window. - `aicore/tasks.py::update_history_summary` (Celery task, fire-and-forget, triggered after every assistant reply) does the actual folding: - Computes `pending = (total_non_system_messages - window_size) - summarized_message_count`; skips (cheap DB count only, no LLM call) if `pending < AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES` (default 10) — so summarization only costs an LLM call once enough new older messages have piled up, not on every single message. - Builds a transcript of the newly-eligible older messages, sends it plus the prior summary to the platform's default provider (`resolve_provider_name()` with no argument) with a dedicated system prompt instructing a concise, factual, context-preserving summary. - Truncates to `AI_HISTORY_SUMMARY_MAX_CHARS` (default 2000) and saves `history_summary` + advances `summarized_message_count`. - Retries up to 3 times with backoff on any exception (`autoretry_for`). - **Never runs for BYOK sessions** — there's no user API key available in a background task, and spending a platform key on a BYOK user's behalf without their knowledge would be an undisclosed behavior change. The task checks `session.provider == ProviderName.BYOK` and skips. ### Settings - `AI_HISTORY_WINDOW_SIZE` (default 20) - `AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES` (default 10) - `AI_HISTORY_SUMMARY_MAX_CHARS` (default 2000) ### Acceptance criteria - [x] Conversations longer than the window still have earlier context available to the model via the summary, without resending the full transcript. - [x] Summarization is fire-and-forget and doesn't block the chat response. - [x] BYOK sessions are never summarized using a platform key. - [x] A transient summarization failure doesn't affect the chat response that triggered it (task runs after the reply, independently retried). ### Related Files - `aicore/models/chat.py` — `history_summary`, `summarized_message_count` fields. - `aicore/migrations/0005_chatsession_history_summary.py`. - `aicore/services/chat_service.py` — `list_history` prepends the summary; triggers the task after each reply. - `aicore/tasks.py` — `update_history_summary`. - `paystream/app_settings/ai.py` — three new settings. --- ## 2. Admin AI Usage Dashboard ### Goal Give admins visibility into which models are actually being used, request counts, and error rates — before deciding what (if anything) to further restrict, per-model or per-tier. ### Design - New `ChatUsageEvent` model: `provider`, `model_name`, `status` (success/error), `error_message` (capped at 500 chars), `session`, `user`, standard audit fields. - `aicore/services/usage_stats.py::record_usage_event()` — best-effort write from `ChatService.stream_reply` at each meaningful outcome. Wrapped in a bare `except Exception` (logged, not raised) since a telemetry write must never break the actual chat response. Deliberately **not** called for requests rejected before a session/provider is resolved (e.g. access or rate-limit rejections) — those are access/rate-limit concerns, not model-usage ones. - `get_usage_summary(since=None, user=None)` — aggregates by `(provider, model_name)`: total requests, successes, errors, error rate, last-used timestamp. Non-admin users see only their own usage (`PlatformAdminService.is_platform_admin` gate); platform admins see aggregate usage across all users. - `get_recent_errors(limit=20, user=None)` — most recent failed requests, same user-scoping. - `usage_window_start(days)` — helper for the dashboard's `7 / 30 / 90 / All time` window selector. - Dashboard template (`chatusageevent/dashboard.html`): a "By provider/model" table (requests, successes, errors, error rate, last used) and a "Recent errors" table (when, user, provider, model, error), plus a "View raw log" link into the regular changelist. ### Follow-on fixes, discovered and applied after initial ship - **Timestamps show an explicit timezone abbreviation** instead of a bare naive-looking datetime — `{{ row.last_used_at|localtime|date:"Y-m-d H:i T" }}` (`{% load tz %}`), e.g. `2026-08-24 13:05 UTC`, so the same dashboard reads correctly regardless of the viewing admin's locale/timezone activation. - **Full error messages, not a truncated tooltip.** A global, app-wide script (`table-truncate.js`) truncates any admin table cell over 32 chars to `text + "..."` with a native `title` tooltip — and does this by overwriting `td.innerText`, which silently destroys any child DOM (it already special-cased `` tags for changelist action columns, but not buttons). This wiped out the dashboard's copy-to-clipboard button entirely. Fixed by extending the same exclusion pattern: ```js if (td.querySelector("a")) return; if (td.querySelector("button") || td.classList.contains("usage-error-cell")) return; ``` Once exempted, the cell's own CSS (`white-space: pre-wrap`, `overflow-wrap: anywhere`) renders the full message wrapping naturally — no truncation, no scroll hack needed, and the copy button now actually renders. - **Copy-to-clipboard** on the (now full-length) error message cell, so the exact error text is available for debugging without depending on a hover tooltip. - **`ChatSession`/`ChatMessage` excluded from generic admin discovery** — `configs/admin_registry.json`'s `aicore` entry sets `primary_model: "chatusageevent"` and lists `chatmessage`/`chatsession` in both `restricted_models` and `model_exclusions`, so the intended admin surface is the usage-event-centric dashboard, not raw CRUD on chat transcripts. - **Distinct icons** for `ChatSession` / `ChatMessage` / `ChatUsageEvent` in `shared/admin/admin_metadata.py::MODEL_ICON_MAP` (previously all three fell back to the same generic `fas fa-database`), and an `aicore` entry (`fas fa-robot`) in `APP_ICON_MAP` (previously missing entirely, so the Applications dashboard card had no icon at all). - **Admin card titles respect proper title casing** — a migration (`0010_alter_chatmessage_options_alter_chatsession_options_and_more`) sets explicit `verbose_name` / `verbose_name_plural` on all three models (they were rendering lowercase by default). - **`aicore` displays as "AI Assistant" in the dashboard UI only** — Django app label itself is unchanged; `DASHBOARD_APP_LABEL_OVERRIDES` in `shared/ui/dashboard/services/dashboard_stats_service.py` maps `"aicore" -> "AI Assistant"` for display purposes, the same centralized mechanism used for any other app that needs a friendlier dashboard label. ### Acceptance criteria - [x] Dashboard shows per-provider/model request counts, success/error counts, error rate, and last-used time, for configurable windows. - [x] Non-admin users see only their own usage; platform admins see aggregate usage. - [x] Recording usage telemetry never breaks the actual chat response on failure. - [x] Full error messages are visible and copyable, not just a truncated tooltip. - [x] Timestamps show their timezone explicitly. - [x] `aicore` has a proper app icon; `ChatSession`/`ChatMessage`/ `ChatUsageEvent` have distinct model icons; admin card titles are properly cased. - [x] Raw chat transcripts (`ChatSession`/`ChatMessage`) aren't exposed through generic admin discovery. ### Related Files - `aicore/models/chat.py` — `ChatUsageEvent`. - `aicore/migrations/0006_chatusageevent.py`, `0007_alter_chatusageevent_provider.py`, `0008_delete_chatusageevent.py` / `0009_chatusageevent.py` (schema revision), `0010_alter_chatmessage_options_alter_chatsession_options_and_more.py`. - `aicore/services/usage_stats.py` — new. - `aicore/admin.py` — dashboard view registration. - `frontend/templates/admin/custom/aicore/chatusageevent/dashboard.html`. - `frontend/static/assets/css/admin/07_pages/changelist.css` — `.usage-error-cell` / `.usage-error-copy` styles. - `frontend/static/assets/js/admin/table-truncate.js` — interactive-cell exclusion fix. - `shared/admin/admin_metadata.py` — `MODEL_ICON_MAP` / `APP_ICON_MAP` entries. - `configs/admin_registry.json` — `aicore` restriction config. - `shared/ui/dashboard/services/dashboard_stats_service.py` — `DASHBOARD_APP_LABEL_OVERRIDES`. --- ## 3. BYOK Per-Tab Throttling ### Goal Stop a single browser tab from hammering an arbitrary user-supplied OpenAI-compatible endpoint via BYOK, as an *additional* layer on top of the existing account-level RPM/RPD limits (which already apply to BYOK, per the production-BYOK decision, and are unchanged). ### Design - **Tab id**: `chat_widget.js` lazily generates `crypto.randomUUID()` (with a fallback for older browsers) on first BYOK use, stores it in `sessionStorage` under `djangoplay:ai:tabId` — same per-tab lifetime as the BYOK credential storage (dies on tab close, never shared cross-tab). Not a secret; just an opaque bucket key. - **Transport**: sent as `X-AI-Tab-Id`, only alongside `X-AI-BYOK-Token` (never on platform-provider requests, where it's meaningless). - **Backend**: `aicore/views/chat.py` reads the header and passes `tab_id` into `ChatService.stream_reply`. New `check_tab_request_limit()` in `rate_limit.py`, keyed `aicore:rl:byoktab:{user_id}:{tab_id}:{minute_bucket}` — namespaced by `user_id` too, so a copied/guessed tab id from another account can't be used to grief someone else's bucket. Enforced only when `byok_credentials is not None`, immediately after the existing account-level `check_request_limits` call, same error-handling shape (raises `AIRateLimitError`, caught, yields a `StreamEvent(type="error", ...)`, no usage event recorded). - **Missing/malformed tab id**: never blocks the request — the check is simply skipped. Account-level limits and the token budget still apply regardless, so there's no security regression if an old cached frontend build doesn't send the header yet. ### Settings - `AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE` (default `10`) — deliberately looser than the account-level `AI_CHAT_RATE_LIMIT_PER_MINUTE` (default `5`), since it's a secondary backstop layered on top of, not replacing, the account limit. ### Acceptance criteria - [x] A tab id is generated once per tab, stored in `sessionStorage`, never in `localStorage`, never sent on non-BYOK requests. - [x] Exceeding `AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE` from one tab returns a clear error without touching another tab's or another user's bucket. - [x] A request missing `X-AI-Tab-Id` still succeeds (subject to the account-level limits) — additive, not a hard requirement. - [x] No usage event or DB write happens when this check rejects a request. ### Related Files - `aicore/services/rate_limit.py` — `check_tab_request_limit()`. - `aicore/services/chat_service.py` — `tab_id` param on `stream_reply`. - `aicore/views/chat.py` — reads `X-AI-Tab-Id`. - `paystream/app_settings/ai.py` — new setting. - `frontend/static/assets/js/components/chat_widget.js` — `getOrCreateTabId()`, header attachment. --- ## 4. Per-Tier Token Budgets ### Goal Replace the single shared `AI_CHAT_TOKEN_BUDGET_PER_MONTH` with independent monthly budgets per provider tier, so unrelated usage doesn't share a ceiling. ### Decisions 1. **BYOK gets its own, separate ceiling** — explicitly a pure abuse backstop, not a cost control, since BYOK tokens are billed to the user's own provider key, not the platform. 2. **Platform providers (xai/openai/custom) and OpenRouter free-catalog models are also split from each other**, not just from BYOK — OpenRouter is free per-token but still shares OpenRouter's own account-level rate limits, so it merits its own ceiling distinct from directly-configured platform providers. ### Design - `aicore/services/chat_service.py::token_budget_tier_for(provider)` maps `ProviderName.BYOK -> "byok"`, `ProviderName.OPENROUTER -> "openrouter"`, everything else (`xai`/`openai`/`custom`) -> `"platform"`. - `rate_limit.py`'s cache key for the monthly counter includes the tier: `aicore:rl:tokmo:{tier}:{user_id}:{month_bucket}`. `check_token_budget()`, `record_tokens_used()`, and `tokens_used_this_month()` all take a required `tier` kwarg. - `ChatService.stream_reply` resolves `budget_tier = token_budget_tier_for(session.provider)` once (after session creation, since that's when `session.provider` is known), looks up the matching limit, and uses it for both the pre-flight `check_token_budget` call and the post-response `record_tokens_used` call. - Purely cache-based, like the pre-existing budget mechanism — no migration required. ### Settings - `AI_CHAT_TOKEN_BUDGET_PER_MONTH` (existing, unchanged default `50000`) — now explicitly scoped to the `"platform"` tier (xai/openai/custom). - `AI_CHAT_TOKEN_BUDGET_OPENROUTER_PER_MONTH` (new) — free-catalog tier. - `AI_CHAT_TOKEN_BUDGET_BYOK_PER_MONTH` (new) — BYOK tier, intended to be higher than the other two since it's an abuse backstop, not a spend control. ### Acceptance criteria - [x] Exhausting the platform-tier budget does not affect a user's OpenRouter or BYOK budget, and vice versa (three fully independent monthly counters per user). - [x] Existing behavior for platform-provider sessions is otherwise unchanged (same default number, same error message). - [x] No migration required. - [ ] Not yet done: surfacing per-tier usage on the admin usage dashboard. `get_usage_summary()` already aggregates by `(provider, model_name)`, which is enough to derive tier-level totals manually today; a dedicated tier rollup would be a small follow-up if wanted directly on the dashboard. ### Related Files - `aicore/services/rate_limit.py` — tiered token-budget keys/functions. - `aicore/services/chat_service.py` — `token_budget_tier_for()`, tier-aware budget calls. - `paystream/app_settings/ai.py` — two new settings (verify these are actually present — see note below). --- ## 5. Open item to verify before shipping `chat_service.py` references `AI_CHAT_TOKEN_BUDGET_OPENROUTER_PER_MONTH`, `AI_CHAT_TOKEN_BUDGET_BYOK_PER_MONTH`, and `AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE` via `getattr(settings, ..., default)`. Confirm all three are actually declared in `paystream/app_settings/ai.py` (not just relied on via the `getattr` fallback default) so they're tunable through the normal encrypted-env pattern like every other `AI_*` setting, and confirm the BYOK budget default matches the intended "deliberately higher than the other two tiers" design above.