djangoplay-web / Apps / Production Bring-Your-Own-Key (Session-Only)
DocsDjangoPlay WebAppsProduction Bring-Your-Own-Key (Session-Only)

Production Bring-Your-Own-Key (Session-Only)

Depends on: free-model-catalog-and-selection.md (chat header UI slot, provider registry pattern)

6 min readApplies to v1.2.2Added in 1.2.2
On this page ▾
  1. 1. Goal
  2. 2. The core design problem
  3. Decisions
  4. What this does NOT protect against (disclosure copy must say this
  5. 3. Design
  6. 3.1 aicore/services/byok.py
  7. 3.2 Endpoint: POST /aicore/byok/validate/
  8. 3.3 Chat request path
  9. 3.4 Frontend
  10. 4. Settings (paystream/app_settings/ai.py)
  11. 5. Acceptance criteria
  12. 6. Related Files

1. Goal

Let a logged-in production user supply their own AI provider, model, and API key for the current browser session only, via a "Use your own agent" link in the chat header — without the key ever touching the database, application logs, or other browser tabs of the same session.

2. The core design problem

Three requirements were in tension and had to be resolved explicitly:

  1. "Not shared across other tabs" — Django's session cookie is shared across all tabs of the same browser for the same origin, so storing the key in request.session would not give tab isolation.
  2. "Not stored" — but every chat request after entry still needs the key to reach the server, for every message sent.
  3. Genuinely session-only (survives refresh within the tab, dies on tab close) — exactly what browser sessionStorage provides natively (it is per-tab, unlike localStorage or cookies).

Decisions

  • Header-per-request (not a short-lived server-side encrypted token): the raw key lives only in the browser tab's sessionStorage; each chat request sends it in a custom header (X-AI-BYOK-Token), read per-request server-side, used to build an ephemeral provider, then discarded. No server-side state for the key at all — simplest to reason about, matches "not stored" literally.
  • Existing account-level rate limits and the monthly token budget continue to apply to BYOK requests — explicitly not waived, so the platform can't be used as a free relay for unrelated traffic even though the tokens are billed to the user's own key.
  • Production HTTPS end-to-end (SECURE_SSL_REDIRECT / HSTS) was already in place, which this design's client-side guarantees assume.
  • The modal validates the key with a cheap test call before accepting it, rather than only discovering it's invalid on the first real chat message.

What this does NOT protect against (disclosure copy must say this

plainly, not just reassuringly)

  • A malicious browser extension in that tab can still read sessionStorage.
  • This protects against server-side persistence and cross-tab leakage, not a compromised client machine.

3. Design

3.1 aicore/services/byok.py

  • BYOKCredentials — a frozen dataclass (base_url, model, api_key), constructed fresh from the incoming header on every request, never assigned to a model field, cached, or passed into a logger.
  • decode_byok_header(raw_header_value) — decodes base64(json({"base_url", "model", "api_key"})) from X-AI-BYOK-Token. Returns None if the header is absent (BYOK simply isn't in play). Raises AIByokError for a present-but-malformed header — callers must surface that as a client error, never silently fall back to a platform provider. Hard length cap (8192 chars) checked before attempting to decode at all. Gated by AI_BYOK_ENABLED as an independent kill switch.
  • build_credentials() validates shape: all three fields required, length-capped (AI_BYOK_MAX_BASE_URL_LENGTH / _MAX_MODEL_LENGTH / _MAX_API_KEY_LENGTH), base URL must parse as http(s):// with a host. Shared by both the per-request header path and the one-off validation endpoint.
  • build_ephemeral_provider(credentials) — constructs an OpenAICompatProvider for exactly one request; never cached, never attached to a session.
  • test_credentials(credentials) — the "cheap test call" for the modal. Calls the provider directly with requests (not OpenAICompatProvider.stream_chat) so it can return a specific, field-attributable message instead of a bare HTTP status:
    • Distinguishes timeout vs. connection error vs. unexpected network error.
    • _looks_like_auth_error() inspects the error body's own wording (not just the HTTP status), because some OpenAI-compatible endpoints (e.g. Gemini's compat layer) return 400 for a bad key instead of 401/403.
    • _extract_provider_error() tolerates both the standard OpenAI error shape ({"error": {...}}) and Gemini's shape (a top-level array wrapping one error object).
    • Maps 401/403/auth-looking-400 → "key rejected," 404 → "check base URL is the API root," 429 → "rate-limited," 5xx → "provider server error," with the underlying provider message included where available.

3.2 Endpoint: POST /aicore/byok/validate/

aicore/views/byok.py — login_required, gated by ai_access_for(). Reads the raw JSON body (not the encoded header — this is the one deliberate exception to "the key only ever travels in a header," since it's a single explicit user action at entry time, not part of the ongoing per-message chat flow), builds credentials, runs test_credentials(), returns {ok, detail}. Body is read, used once, discarded — never written to the database or a log line.

3.3 Chat request path

  • ChatStreamView (aicore/views/chat.py) reads the optional X-AI-BYOK-Token header on every message via decode_byok_header() and threads the resulting BYOKCredentials | None through to ChatService.stream_reply.
  • ChatService builds and discards an ephemeral provider per message for BYOK sessions, instead of resolving one from the settings-backed registry. A BYOK ChatSession persists only the non-secret model name.
  • aicore.constants.ProviderName.BYOK is a session-labeling marker only — deliberately excluded from resolve_provider_name's default fallback chain, and providers/registry.py::get_provider("byok") raises defensively if ever hit directly (BYOK credentials are never resolvable from settings by design; ChatService must branch to build_ephemeral_provider instead).

3.4 Frontend

  • "Use your own agent" link in the chat header, next to the model dropdown.
  • Disclosure modal: base URL / model / API key fields, required acknowledgment checkbox, Test → "Use this agent" flow.
  • Credentials held only in sessionStorage — genuinely per-tab, cleared on tab close, never localStorage, never shared across tabs.
  • An active-BYOK banner with a "Stop" button.
  • A synthetic "Integrated AI Model" entry appears in the Phase-1-style model dropdown only after a successful integration, letting the user switch between their own agent and catalog models via the existing model-switch confirmation — without discarding the stored credentials (only "Stop" clears them).
  • The key is sent as X-AI-BYOK-Token on each chat request, never in the request body — deliberate, so it's easier to guarantee exclusion from any request-body logging middleware.

4. Settings (paystream/app_settings/ai.py)

  • AI_BYOK_ENABLED — independent kill switch.
  • AI_BYOK_REQUEST_TIMEOUT — timeout for the ephemeral provider's actual chat requests.
  • AI_BYOK_VALIDATION_TIMEOUT — timeout for the modal's test call (shorter, since it's a synchronous UI wait).
  • AI_BYOK_MAX_BASE_URL_LENGTH, AI_BYOK_MAX_MODEL_LENGTH, AI_BYOK_MAX_API_KEY_LENGTH — sanity length caps on the per-request header fields.

5. Acceptance criteria

  • Disclosure text is shown and must be acknowledged before the modal accepts input.
  • Key is never present in any Django log line, DB row, or Sentry/APM breadcrumb — verified against APIRequestLoggingMiddleware, which only ever persists path/method/status/user-agent, never headers or bodies.
  • Opening the same account in a second tab does not expose or reuse the BYOK key entered in the first tab.
  • Closing the tab ends BYOK availability for that tab.
  • Existing ai_access_for() gating still applies — BYOK doesn't bypass who's allowed to use the chatbot, only which provider serves them.
  • BYOK requests still pass through the account-level rate limits (later joined by a per-tab throttle and its own token-budget tier — see the usage-controls-and-hardening doc).
  • aicore/services/byok.py — new.
  • aicore/views/byok.py — new (/aicore/byok/validate/).
  • aicore/views/chat.py — reads X-AI-BYOK-Token.
  • aicore/services/chat_service.py — byok_credentials param on get_or_create_session / stream_reply.
  • aicore/providers/registry.py — ProviderName.BYOK guard.
  • paystream/app_settings/ai.py — BYOK settings.
  • frontend/templates/base/chat_widget.html, frontend/static/assets/js/components/chat_widget.js — modal, banner, sessionStorage credential handling, header attachment.