Usage Controls & Platform Hardening
Depends on: free-model-catalog-and-selection.md (provider tiers), production-byok-session-only.md (BYOK header flow, per-request ephemeral provider)
On this page ▾
- 1. Conversation History Summarization
- Goal
- Design
- Settings
- Acceptance criteria
- Related Files
- 2. Admin AI Usage Dashboard
- Goal
- Design
- Follow-on fixes, discovered and applied after initial ship
- Acceptance criteria
- Related Files
- 3. BYOK Per-Tab Throttling
- Goal
- Design
- Settings
- Acceptance criteria
- Related Files
- 4. Per-Tier Token Budgets
- Goal
- Decisions
- Design
- Settings
- Acceptance criteria
- Related Files
- 5. Open item to verify before shipping
Covers four changes, built in this order: conversation history summarization, the admin AI usage dashboard, BYOK per-tab throttling, and per-tier token budgets.
1. Conversation History Summarization
Goal
Replace the flat sliding-window history (MAX_HISTORY_MESSAGES, no
compression) with something that retains earlier context in long
conversations without sending the entire transcript on every request.
Design
ChatSessiongains two fields:history_summary(text, running summary) andsummarized_message_count(how many older messages have already been folded into it).ChatService.list_historykeeps a bounded recent-messages window (AI_HISTORY_WINDOW_SIZE, default 20) verbatim, and — ifsession.history_summaryis non-empty — prepends it as a labeled "already condensed, older than the recent window" system-style message ahead of the recent window.aicore/tasks.py::update_history_summary(Celery task, fire-and-forget, triggered after every assistant reply) does the actual folding:- Computes
pending = (total_non_system_messages - window_size) - summarized_message_count; skips (cheap DB count only, no LLM call) ifpending < AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES(default 10) — so summarization only costs an LLM call once enough new older messages have piled up, not on every single message. - Builds a transcript of the newly-eligible older messages, sends it plus
the prior summary to the platform's default provider
(
resolve_provider_name()with no argument) with a dedicated system prompt instructing a concise, factual, context-preserving summary. - Truncates to
AI_HISTORY_SUMMARY_MAX_CHARS(default 2000) and saveshistory_summary+ advancessummarized_message_count. - Retries up to 3 times with backoff on any exception (
autoretry_for).
- Computes
- Never runs for BYOK sessions — there's no user API key available in
a background task, and spending a platform key on a BYOK user's behalf
without their knowledge would be an undisclosed behavior change. The
task checks
session.provider == ProviderName.BYOKand skips.
Settings
AI_HISTORY_WINDOW_SIZE(default 20)AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES(default 10)AI_HISTORY_SUMMARY_MAX_CHARS(default 2000)
Acceptance criteria
- Conversations longer than the window still have earlier context available to the model via the summary, without resending the full transcript.
- Summarization is fire-and-forget and doesn't block the chat response.
- BYOK sessions are never summarized using a platform key.
- A transient summarization failure doesn't affect the chat response that triggered it (task runs after the reply, independently retried).
Related Files
aicore/models/chat.py—history_summary,summarized_message_countfields.aicore/migrations/0005_chatsession_history_summary.py.aicore/services/chat_service.py—list_historyprepends the summary; triggers the task after each reply.aicore/tasks.py—update_history_summary.paystream/app_settings/ai.py— three new settings.
2. Admin AI Usage Dashboard
Goal
Give admins visibility into which models are actually being used, request counts, and error rates — before deciding what (if anything) to further restrict, per-model or per-tier.
Design
- New
ChatUsageEventmodel:provider,model_name,status(success/error),error_message(capped at 500 chars),session,user, standard audit fields. aicore/services/usage_stats.py::record_usage_event()— best-effort write fromChatService.stream_replyat each meaningful outcome. Wrapped in a bareexcept Exception(logged, not raised) since a telemetry write must never break the actual chat response. Deliberately not called for requests rejected before a session/provider is resolved (e.g. access or rate-limit rejections) — those are access/rate-limit concerns, not model-usage ones.get_usage_summary(since=None, user=None)— aggregates by(provider, model_name): total requests, successes, errors, error rate, last-used timestamp. Non-admin users see only their own usage (PlatformAdminService.is_platform_admingate); platform admins see aggregate usage across all users.get_recent_errors(limit=20, user=None)— most recent failed requests, same user-scoping.usage_window_start(days)— helper for the dashboard's7 / 30 / 90 / All timewindow selector.- Dashboard template (
chatusageevent/dashboard.html): a "By provider/model" table (requests, successes, errors, error rate, last used) and a "Recent errors" table (when, user, provider, model, error), plus a "View raw log" link into the regular changelist.
Follow-on fixes, discovered and applied after initial ship
- Timestamps show an explicit timezone abbreviation instead of a bare
naive-looking datetime —
{{ row.last_used_at|localtime|date:"Y-m-d H:i T" }}({% load tz %}), e.g.2026-08-24 13:05 UTC, so the same dashboard reads correctly regardless of the viewing admin's locale/timezone activation. - Full error messages, not a truncated tooltip. A global,
app-wide script (
table-truncate.js) truncates any admin table cell over 32 chars totext + "..."with a nativetitletooltip — and does this by overwritingtd.innerText, which silently destroys any child DOM (it already special-cased<a>tags for changelist action columns, but not buttons). This wiped out the dashboard's copy-to-clipboard button entirely. Fixed by extending the same exclusion pattern:
Once exempted, the cell's own CSS (if (td.querySelector("a")) return; if (td.querySelector("button") || td.classList.contains("usage-error-cell")) return;white-space: pre-wrap,overflow-wrap: anywhere) renders the full message wrapping naturally — no truncation, no scroll hack needed, and the copy button now actually renders. - Copy-to-clipboard on the (now full-length) error message cell, so the exact error text is available for debugging without depending on a hover tooltip.
ChatSession/ChatMessageexcluded from generic admin discovery —configs/admin_registry.json'saicoreentry setsprimary_model: "chatusageevent"and listschatmessage/chatsessionin bothrestricted_modelsandmodel_exclusions, so the intended admin surface is the usage-event-centric dashboard, not raw CRUD on chat transcripts.- Distinct icons for
ChatSession/ChatMessage/ChatUsageEventinshared/admin/admin_metadata.py::MODEL_ICON_MAP(previously all three fell back to the same genericfas fa-database), and anaicoreentry (fas fa-robot) inAPP_ICON_MAP(previously missing entirely, so the Applications dashboard card had no icon at all). - Admin card titles respect proper title casing — a migration
(
0010_alter_chatmessage_options_alter_chatsession_options_and_more) sets explicitverbose_name/verbose_name_pluralon all three models (they were rendering lowercase by default). aicoredisplays as "AI Assistant" in the dashboard UI only — Django app label itself is unchanged;DASHBOARD_APP_LABEL_OVERRIDESinshared/ui/dashboard/services/dashboard_stats_service.pymaps"aicore" -> "AI Assistant"for display purposes, the same centralized mechanism used for any other app that needs a friendlier dashboard label.
Acceptance criteria
- Dashboard shows per-provider/model request counts, success/error counts, error rate, and last-used time, for configurable windows.
- Non-admin users see only their own usage; platform admins see aggregate usage.
- Recording usage telemetry never breaks the actual chat response on failure.
- Full error messages are visible and copyable, not just a truncated tooltip.
- Timestamps show their timezone explicitly.
-
aicorehas a proper app icon;ChatSession/ChatMessage/ChatUsageEventhave distinct model icons; admin card titles are properly cased. - Raw chat transcripts (
ChatSession/ChatMessage) aren't exposed through generic admin discovery.
Related Files
aicore/models/chat.py—ChatUsageEvent.aicore/migrations/0006_chatusageevent.py,0007_alter_chatusageevent_provider.py,0008_delete_chatusageevent.py/0009_chatusageevent.py(schema revision),0010_alter_chatmessage_options_alter_chatsession_options_and_more.py.aicore/services/usage_stats.py— new.aicore/admin.py— dashboard view registration.frontend/templates/admin/custom/aicore/chatusageevent/dashboard.html.frontend/static/assets/css/admin/07_pages/changelist.css—.usage-error-cell/.usage-error-copystyles.frontend/static/assets/js/admin/table-truncate.js— interactive-cell exclusion fix.shared/admin/admin_metadata.py—MODEL_ICON_MAP/APP_ICON_MAPentries.configs/admin_registry.json—aicorerestriction config.shared/ui/dashboard/services/dashboard_stats_service.py—DASHBOARD_APP_LABEL_OVERRIDES.
3. BYOK Per-Tab Throttling
Goal
Stop a single browser tab from hammering an arbitrary user-supplied OpenAI-compatible endpoint via BYOK, as an additional layer on top of the existing account-level RPM/RPD limits (which already apply to BYOK, per the production-BYOK decision, and are unchanged).
Design
- Tab id:
chat_widget.jslazily generatescrypto.randomUUID()(with a fallback for older browsers) on first BYOK use, stores it insessionStorageunderdjangoplay:ai:tabId— same per-tab lifetime as the BYOK credential storage (dies on tab close, never shared cross-tab). Not a secret; just an opaque bucket key. - Transport: sent as
X-AI-Tab-Id, only alongsideX-AI-BYOK-Token(never on platform-provider requests, where it's meaningless). - Backend:
aicore/views/chat.pyreads the header and passestab_idintoChatService.stream_reply. Newcheck_tab_request_limit()inrate_limit.py, keyedaicore:rl:byoktab:{user_id}:{tab_id}:{minute_bucket}— namespaced byuser_idtoo, so a copied/guessed tab id from another account can't be used to grief someone else's bucket. Enforced only whenbyok_credentials is not None, immediately after the existing account-levelcheck_request_limitscall, same error-handling shape (raisesAIRateLimitError, caught, yields aStreamEvent(type="error", ...), no usage event recorded). - Missing/malformed tab id: never blocks the request — the check is simply skipped. Account-level limits and the token budget still apply regardless, so there's no security regression if an old cached frontend build doesn't send the header yet.
Settings
AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE(default10) — deliberately looser than the account-levelAI_CHAT_RATE_LIMIT_PER_MINUTE(default5), since it's a secondary backstop layered on top of, not replacing, the account limit.
Acceptance criteria
- A tab id is generated once per tab, stored in
sessionStorage, never inlocalStorage, never sent on non-BYOK requests. - Exceeding
AI_BYOK_TAB_RATE_LIMIT_PER_MINUTEfrom one tab returns a clear error without touching another tab's or another user's bucket. - A request missing
X-AI-Tab-Idstill succeeds (subject to the account-level limits) — additive, not a hard requirement. - No usage event or DB write happens when this check rejects a request.
Related Files
aicore/services/rate_limit.py—check_tab_request_limit().aicore/services/chat_service.py—tab_idparam onstream_reply.aicore/views/chat.py— readsX-AI-Tab-Id.paystream/app_settings/ai.py— new setting.frontend/static/assets/js/components/chat_widget.js—getOrCreateTabId(), header attachment.
4. Per-Tier Token Budgets
Goal
Replace the single shared AI_CHAT_TOKEN_BUDGET_PER_MONTH with independent
monthly budgets per provider tier, so unrelated usage doesn't share a
ceiling.
Decisions
- BYOK gets its own, separate ceiling — explicitly a pure abuse backstop, not a cost control, since BYOK tokens are billed to the user's own provider key, not the platform.
- Platform providers (xai/openai/custom) and OpenRouter free-catalog models are also split from each other, not just from BYOK — OpenRouter is free per-token but still shares OpenRouter's own account-level rate limits, so it merits its own ceiling distinct from directly-configured platform providers.
Design
aicore/services/chat_service.py::token_budget_tier_for(provider)mapsProviderName.BYOK -> "byok",ProviderName.OPENROUTER -> "openrouter", everything else (xai/openai/custom) ->"platform".rate_limit.py's cache key for the monthly counter includes the tier:aicore:rl:tokmo:{tier}:{user_id}:{month_bucket}.check_token_budget(),record_tokens_used(), andtokens_used_this_month()all take a requiredtierkwarg.ChatService.stream_replyresolvesbudget_tier = token_budget_tier_for(session.provider)once (after session creation, since that's whensession.provideris known), looks up the matching limit, and uses it for both the pre-flightcheck_token_budgetcall and the post-responserecord_tokens_usedcall.- Purely cache-based, like the pre-existing budget mechanism — no migration required.
Settings
AI_CHAT_TOKEN_BUDGET_PER_MONTH(existing, unchanged default50000) — now explicitly scoped to the"platform"tier (xai/openai/custom).AI_CHAT_TOKEN_BUDGET_OPENROUTER_PER_MONTH(new) — free-catalog tier.AI_CHAT_TOKEN_BUDGET_BYOK_PER_MONTH(new) — BYOK tier, intended to be higher than the other two since it's an abuse backstop, not a spend control.
Acceptance criteria
- Exhausting the platform-tier budget does not affect a user's OpenRouter or BYOK budget, and vice versa (three fully independent monthly counters per user).
- Existing behavior for platform-provider sessions is otherwise unchanged (same default number, same error message).
- No migration required.
- Not yet done: surfacing per-tier usage on the admin usage dashboard.
get_usage_summary()already aggregates by(provider, model_name), which is enough to derive tier-level totals manually today; a dedicated tier rollup would be a small follow-up if wanted directly on the dashboard.
Related Files
aicore/services/rate_limit.py— tiered token-budget keys/functions.aicore/services/chat_service.py—token_budget_tier_for(), tier-aware budget calls.paystream/app_settings/ai.py— two new settings (verify these are actually present — see note below).
5. Open item to verify before shipping
chat_service.py references AI_CHAT_TOKEN_BUDGET_OPENROUTER_PER_MONTH,
AI_CHAT_TOKEN_BUDGET_BYOK_PER_MONTH, and
AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE via getattr(settings, ..., default).
Confirm all three are actually declared in paystream/app_settings/ai.py
(not just relied on via the getattr fallback default) so they're tunable
through the normal encrypted-env pattern like every other AI_* setting,
and confirm the BYOK budget default matches the intended "deliberately
higher than the other two tiers" design above.