djangoplay-web / Apps / Usage Controls & Platform Hardening
DocsDjangoPlay WebAppsUsage Controls & Platform Hardening

Usage Controls & Platform Hardening

Depends on: free-model-catalog-and-selection.md (provider tiers), production-byok-session-only.md (BYOK header flow, per-request ephemeral provider)

9 min readApplies to v1.2.2Added in 1.2.2
On this page ▾
  1. 1. Conversation History Summarization
  2. Goal
  3. Design
  4. Settings
  5. Acceptance criteria
  6. Related Files
  7. 2. Admin AI Usage Dashboard
  8. Goal
  9. Design
  10. Follow-on fixes, discovered and applied after initial ship
  11. Acceptance criteria
  12. Related Files
  13. 3. BYOK Per-Tab Throttling
  14. Goal
  15. Design
  16. Settings
  17. Acceptance criteria
  18. Related Files
  19. 4. Per-Tier Token Budgets
  20. Goal
  21. Decisions
  22. Design
  23. Settings
  24. Acceptance criteria
  25. Related Files
  26. 5. Open item to verify before shipping

Covers four changes, built in this order: conversation history summarization, the admin AI usage dashboard, BYOK per-tab throttling, and per-tier token budgets.


1. Conversation History Summarization

Goal

Replace the flat sliding-window history (MAX_HISTORY_MESSAGES, no compression) with something that retains earlier context in long conversations without sending the entire transcript on every request.

Design

  • ChatSession gains two fields: history_summary (text, running summary) and summarized_message_count (how many older messages have already been folded into it).
  • ChatService.list_history keeps a bounded recent-messages window (AI_HISTORY_WINDOW_SIZE, default 20) verbatim, and — if session.history_summary is non-empty — prepends it as a labeled "already condensed, older than the recent window" system-style message ahead of the recent window.
  • aicore/tasks.py::update_history_summary (Celery task, fire-and-forget, triggered after every assistant reply) does the actual folding:
    • Computes pending = (total_non_system_messages - window_size) - summarized_message_count; skips (cheap DB count only, no LLM call) if pending < AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES (default 10) — so summarization only costs an LLM call once enough new older messages have piled up, not on every single message.
    • Builds a transcript of the newly-eligible older messages, sends it plus the prior summary to the platform's default provider (resolve_provider_name() with no argument) with a dedicated system prompt instructing a concise, factual, context-preserving summary.
    • Truncates to AI_HISTORY_SUMMARY_MAX_CHARS (default 2000) and saves history_summary + advances summarized_message_count.
    • Retries up to 3 times with backoff on any exception (autoretry_for).
  • Never runs for BYOK sessions — there's no user API key available in a background task, and spending a platform key on a BYOK user's behalf without their knowledge would be an undisclosed behavior change. The task checks session.provider == ProviderName.BYOK and skips.

Settings

  • AI_HISTORY_WINDOW_SIZE (default 20)
  • AI_HISTORY_SUMMARIZATION_TRIGGER_MESSAGES (default 10)
  • AI_HISTORY_SUMMARY_MAX_CHARS (default 2000)

Acceptance criteria

  • Conversations longer than the window still have earlier context available to the model via the summary, without resending the full transcript.
  • Summarization is fire-and-forget and doesn't block the chat response.
  • BYOK sessions are never summarized using a platform key.
  • A transient summarization failure doesn't affect the chat response that triggered it (task runs after the reply, independently retried).
  • aicore/models/chat.py — history_summary, summarized_message_count fields.
  • aicore/migrations/0005_chatsession_history_summary.py.
  • aicore/services/chat_service.py — list_history prepends the summary; triggers the task after each reply.
  • aicore/tasks.py — update_history_summary.
  • paystream/app_settings/ai.py — three new settings.

2. Admin AI Usage Dashboard

Goal

Give admins visibility into which models are actually being used, request counts, and error rates — before deciding what (if anything) to further restrict, per-model or per-tier.

Design

  • New ChatUsageEvent model: provider, model_name, status (success/error), error_message (capped at 500 chars), session, user, standard audit fields.
  • aicore/services/usage_stats.py::record_usage_event() — best-effort write from ChatService.stream_reply at each meaningful outcome. Wrapped in a bare except Exception (logged, not raised) since a telemetry write must never break the actual chat response. Deliberately not called for requests rejected before a session/provider is resolved (e.g. access or rate-limit rejections) — those are access/rate-limit concerns, not model-usage ones.
  • get_usage_summary(since=None, user=None) — aggregates by (provider, model_name): total requests, successes, errors, error rate, last-used timestamp. Non-admin users see only their own usage (PlatformAdminService.is_platform_admin gate); platform admins see aggregate usage across all users.
  • get_recent_errors(limit=20, user=None) — most recent failed requests, same user-scoping.
  • usage_window_start(days) — helper for the dashboard's 7 / 30 / 90 / All time window selector.
  • Dashboard template (chatusageevent/dashboard.html): a "By provider/model" table (requests, successes, errors, error rate, last used) and a "Recent errors" table (when, user, provider, model, error), plus a "View raw log" link into the regular changelist.

Follow-on fixes, discovered and applied after initial ship

  • Timestamps show an explicit timezone abbreviation instead of a bare naive-looking datetime — {{ row.last_used_at|localtime|date:"Y-m-d H:i T" }} ({% load tz %}), e.g. 2026-08-24 13:05 UTC, so the same dashboard reads correctly regardless of the viewing admin's locale/timezone activation.
  • Full error messages, not a truncated tooltip. A global, app-wide script (table-truncate.js) truncates any admin table cell over 32 chars to text + "..." with a native title tooltip — and does this by overwriting td.innerText, which silently destroys any child DOM (it already special-cased <a> tags for changelist action columns, but not buttons). This wiped out the dashboard's copy-to-clipboard button entirely. Fixed by extending the same exclusion pattern:
    if (td.querySelector("a")) return;
    if (td.querySelector("button") || td.classList.contains("usage-error-cell")) return;
    
    Once exempted, the cell's own CSS (white-space: pre-wrap, overflow-wrap: anywhere) renders the full message wrapping naturally — no truncation, no scroll hack needed, and the copy button now actually renders.
  • Copy-to-clipboard on the (now full-length) error message cell, so the exact error text is available for debugging without depending on a hover tooltip.
  • ChatSession/ChatMessage excluded from generic admin discovery — configs/admin_registry.json's aicore entry sets primary_model: "chatusageevent" and lists chatmessage/chatsession in both restricted_models and model_exclusions, so the intended admin surface is the usage-event-centric dashboard, not raw CRUD on chat transcripts.
  • Distinct icons for ChatSession / ChatMessage / ChatUsageEvent in shared/admin/admin_metadata.py::MODEL_ICON_MAP (previously all three fell back to the same generic fas fa-database), and an aicore entry (fas fa-robot) in APP_ICON_MAP (previously missing entirely, so the Applications dashboard card had no icon at all).
  • Admin card titles respect proper title casing — a migration (0010_alter_chatmessage_options_alter_chatsession_options_and_more) sets explicit verbose_name / verbose_name_plural on all three models (they were rendering lowercase by default).
  • aicore displays as "AI Assistant" in the dashboard UI only — Django app label itself is unchanged; DASHBOARD_APP_LABEL_OVERRIDES in shared/ui/dashboard/services/dashboard_stats_service.py maps "aicore" -> "AI Assistant" for display purposes, the same centralized mechanism used for any other app that needs a friendlier dashboard label.

Acceptance criteria

  • Dashboard shows per-provider/model request counts, success/error counts, error rate, and last-used time, for configurable windows.
  • Non-admin users see only their own usage; platform admins see aggregate usage.
  • Recording usage telemetry never breaks the actual chat response on failure.
  • Full error messages are visible and copyable, not just a truncated tooltip.
  • Timestamps show their timezone explicitly.
  • aicore has a proper app icon; ChatSession/ChatMessage/ ChatUsageEvent have distinct model icons; admin card titles are properly cased.
  • Raw chat transcripts (ChatSession/ChatMessage) aren't exposed through generic admin discovery.
  • aicore/models/chat.py — ChatUsageEvent.
  • aicore/migrations/0006_chatusageevent.py, 0007_alter_chatusageevent_provider.py, 0008_delete_chatusageevent.py / 0009_chatusageevent.py (schema revision), 0010_alter_chatmessage_options_alter_chatsession_options_and_more.py.
  • aicore/services/usage_stats.py — new.
  • aicore/admin.py — dashboard view registration.
  • frontend/templates/admin/custom/aicore/chatusageevent/dashboard.html.
  • frontend/static/assets/css/admin/07_pages/changelist.css — .usage-error-cell / .usage-error-copy styles.
  • frontend/static/assets/js/admin/table-truncate.js — interactive-cell exclusion fix.
  • shared/admin/admin_metadata.py — MODEL_ICON_MAP / APP_ICON_MAP entries.
  • configs/admin_registry.json — aicore restriction config.
  • shared/ui/dashboard/services/dashboard_stats_service.py — DASHBOARD_APP_LABEL_OVERRIDES.

3. BYOK Per-Tab Throttling

Goal

Stop a single browser tab from hammering an arbitrary user-supplied OpenAI-compatible endpoint via BYOK, as an additional layer on top of the existing account-level RPM/RPD limits (which already apply to BYOK, per the production-BYOK decision, and are unchanged).

Design

  • Tab id: chat_widget.js lazily generates crypto.randomUUID() (with a fallback for older browsers) on first BYOK use, stores it in sessionStorage under djangoplay:ai:tabId — same per-tab lifetime as the BYOK credential storage (dies on tab close, never shared cross-tab). Not a secret; just an opaque bucket key.
  • Transport: sent as X-AI-Tab-Id, only alongside X-AI-BYOK-Token (never on platform-provider requests, where it's meaningless).
  • Backend: aicore/views/chat.py reads the header and passes tab_id into ChatService.stream_reply. New check_tab_request_limit() in rate_limit.py, keyed aicore:rl:byoktab:{user_id}:{tab_id}:{minute_bucket} — namespaced by user_id too, so a copied/guessed tab id from another account can't be used to grief someone else's bucket. Enforced only when byok_credentials is not None, immediately after the existing account-level check_request_limits call, same error-handling shape (raises AIRateLimitError, caught, yields a StreamEvent(type="error", ...), no usage event recorded).
  • Missing/malformed tab id: never blocks the request — the check is simply skipped. Account-level limits and the token budget still apply regardless, so there's no security regression if an old cached frontend build doesn't send the header yet.

Settings

  • AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE (default 10) — deliberately looser than the account-level AI_CHAT_RATE_LIMIT_PER_MINUTE (default 5), since it's a secondary backstop layered on top of, not replacing, the account limit.

Acceptance criteria

  • A tab id is generated once per tab, stored in sessionStorage, never in localStorage, never sent on non-BYOK requests.
  • Exceeding AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE from one tab returns a clear error without touching another tab's or another user's bucket.
  • A request missing X-AI-Tab-Id still succeeds (subject to the account-level limits) — additive, not a hard requirement.
  • No usage event or DB write happens when this check rejects a request.
  • aicore/services/rate_limit.py — check_tab_request_limit().
  • aicore/services/chat_service.py — tab_id param on stream_reply.
  • aicore/views/chat.py — reads X-AI-Tab-Id.
  • paystream/app_settings/ai.py — new setting.
  • frontend/static/assets/js/components/chat_widget.js — getOrCreateTabId(), header attachment.

4. Per-Tier Token Budgets

Goal

Replace the single shared AI_CHAT_TOKEN_BUDGET_PER_MONTH with independent monthly budgets per provider tier, so unrelated usage doesn't share a ceiling.

Decisions

  1. BYOK gets its own, separate ceiling — explicitly a pure abuse backstop, not a cost control, since BYOK tokens are billed to the user's own provider key, not the platform.
  2. Platform providers (xai/openai/custom) and OpenRouter free-catalog models are also split from each other, not just from BYOK — OpenRouter is free per-token but still shares OpenRouter's own account-level rate limits, so it merits its own ceiling distinct from directly-configured platform providers.

Design

  • aicore/services/chat_service.py::token_budget_tier_for(provider) maps ProviderName.BYOK -> "byok", ProviderName.OPENROUTER -> "openrouter", everything else (xai/openai/custom) -> "platform".
  • rate_limit.py's cache key for the monthly counter includes the tier: aicore:rl:tokmo:{tier}:{user_id}:{month_bucket}. check_token_budget(), record_tokens_used(), and tokens_used_this_month() all take a required tier kwarg.
  • ChatService.stream_reply resolves budget_tier = token_budget_tier_for(session.provider) once (after session creation, since that's when session.provider is known), looks up the matching limit, and uses it for both the pre-flight check_token_budget call and the post-response record_tokens_used call.
  • Purely cache-based, like the pre-existing budget mechanism — no migration required.

Settings

  • AI_CHAT_TOKEN_BUDGET_PER_MONTH (existing, unchanged default 50000) — now explicitly scoped to the "platform" tier (xai/openai/custom).
  • AI_CHAT_TOKEN_BUDGET_OPENROUTER_PER_MONTH (new) — free-catalog tier.
  • AI_CHAT_TOKEN_BUDGET_BYOK_PER_MONTH (new) — BYOK tier, intended to be higher than the other two since it's an abuse backstop, not a spend control.

Acceptance criteria

  • Exhausting the platform-tier budget does not affect a user's OpenRouter or BYOK budget, and vice versa (three fully independent monthly counters per user).
  • Existing behavior for platform-provider sessions is otherwise unchanged (same default number, same error message).
  • No migration required.
  • Not yet done: surfacing per-tier usage on the admin usage dashboard. get_usage_summary() already aggregates by (provider, model_name), which is enough to derive tier-level totals manually today; a dedicated tier rollup would be a small follow-up if wanted directly on the dashboard.
  • aicore/services/rate_limit.py — tiered token-budget keys/functions.
  • aicore/services/chat_service.py — token_budget_tier_for(), tier-aware budget calls.
  • paystream/app_settings/ai.py — two new settings (verify these are actually present — see note below).

5. Open item to verify before shipping

chat_service.py references AI_CHAT_TOKEN_BUDGET_OPENROUTER_PER_MONTH, AI_CHAT_TOKEN_BUDGET_BYOK_PER_MONTH, and AI_BYOK_TAB_RATE_LIMIT_PER_MINUTE via getattr(settings, ..., default). Confirm all three are actually declared in paystream/app_settings/ai.py (not just relied on via the getattr fallback default) so they're tunable through the normal encrypted-env pattern like every other AI_* setting, and confirm the BYOK budget default matches the intended "deliberately higher than the other two tiers" design above.