Production Bring-Your-Own-Key (Session-Only)
Depends on: free-model-catalog-and-selection.md (chat header UI slot, provider registry pattern)
On this page ▾
1. Goal
Let a logged-in production user supply their own AI provider, model, and API key for the current browser session only, via a "Use your own agent" link in the chat header — without the key ever touching the database, application logs, or other browser tabs of the same session.
2. The core design problem
Three requirements were in tension and had to be resolved explicitly:
- "Not shared across other tabs" — Django's session cookie is shared
across all tabs of the same browser for the same origin, so storing the
key in
request.sessionwould not give tab isolation. - "Not stored" — but every chat request after entry still needs the key to reach the server, for every message sent.
- Genuinely session-only (survives refresh within the tab, dies on tab
close) — exactly what browser
sessionStorageprovides natively (it is per-tab, unlikelocalStorageor cookies).
Decisions
- Header-per-request (not a short-lived server-side encrypted token):
the raw key lives only in the browser tab's
sessionStorage; each chat request sends it in a custom header (X-AI-BYOK-Token), read per-request server-side, used to build an ephemeral provider, then discarded. No server-side state for the key at all — simplest to reason about, matches "not stored" literally. - Existing account-level rate limits and the monthly token budget continue to apply to BYOK requests — explicitly not waived, so the platform can't be used as a free relay for unrelated traffic even though the tokens are billed to the user's own key.
- Production HTTPS end-to-end (
SECURE_SSL_REDIRECT/ HSTS) was already in place, which this design's client-side guarantees assume. - The modal validates the key with a cheap test call before accepting it, rather than only discovering it's invalid on the first real chat message.
What this does NOT protect against (disclosure copy must say this
plainly, not just reassuringly)
- A malicious browser extension in that tab can still read
sessionStorage. - This protects against server-side persistence and cross-tab leakage, not a compromised client machine.
3. Design
3.1 aicore/services/byok.py
BYOKCredentials— a frozen dataclass (base_url,model,api_key), constructed fresh from the incoming header on every request, never assigned to a model field, cached, or passed into a logger.decode_byok_header(raw_header_value)— decodesbase64(json({"base_url", "model", "api_key"}))fromX-AI-BYOK-Token. ReturnsNoneif the header is absent (BYOK simply isn't in play). RaisesAIByokErrorfor a present-but-malformed header — callers must surface that as a client error, never silently fall back to a platform provider. Hard length cap (8192 chars) checked before attempting to decode at all. Gated byAI_BYOK_ENABLEDas an independent kill switch.build_credentials()validates shape: all three fields required, length-capped (AI_BYOK_MAX_BASE_URL_LENGTH/_MAX_MODEL_LENGTH/_MAX_API_KEY_LENGTH), base URL must parse ashttp(s)://with a host. Shared by both the per-request header path and the one-off validation endpoint.build_ephemeral_provider(credentials)— constructs anOpenAICompatProviderfor exactly one request; never cached, never attached to a session.test_credentials(credentials)— the "cheap test call" for the modal. Calls the provider directly withrequests(notOpenAICompatProvider.stream_chat) so it can return a specific, field-attributable message instead of a bare HTTP status:- Distinguishes timeout vs. connection error vs. unexpected network error.
_looks_like_auth_error()inspects the error body's own wording (not just the HTTP status), because some OpenAI-compatible endpoints (e.g. Gemini's compat layer) return 400 for a bad key instead of 401/403._extract_provider_error()tolerates both the standard OpenAI error shape ({"error": {...}}) and Gemini's shape (a top-level array wrapping one error object).- Maps 401/403/auth-looking-400 → "key rejected," 404 → "check base URL is the API root," 429 → "rate-limited," 5xx → "provider server error," with the underlying provider message included where available.
3.2 Endpoint: POST /aicore/byok/validate/
aicore/views/byok.py — login_required, gated by ai_access_for().
Reads the raw JSON body (not the encoded header — this is the one
deliberate exception to "the key only ever travels in a header," since
it's a single explicit user action at entry time, not part of the ongoing
per-message chat flow), builds credentials, runs test_credentials(),
returns {ok, detail}. Body is read, used once, discarded — never written
to the database or a log line.
3.3 Chat request path
ChatStreamView(aicore/views/chat.py) reads the optionalX-AI-BYOK-Tokenheader on every message viadecode_byok_header()and threads the resultingBYOKCredentials | Nonethrough toChatService.stream_reply.ChatServicebuilds and discards an ephemeral provider per message for BYOK sessions, instead of resolving one from the settings-backed registry. A BYOKChatSessionpersists only the non-secret model name.aicore.constants.ProviderName.BYOKis a session-labeling marker only — deliberately excluded fromresolve_provider_name's default fallback chain, andproviders/registry.py::get_provider("byok")raises defensively if ever hit directly (BYOK credentials are never resolvable from settings by design;ChatServicemust branch tobuild_ephemeral_providerinstead).
3.4 Frontend
- "Use your own agent" link in the chat header, next to the model dropdown.
- Disclosure modal: base URL / model / API key fields, required acknowledgment checkbox, Test → "Use this agent" flow.
- Credentials held only in
sessionStorage— genuinely per-tab, cleared on tab close, neverlocalStorage, never shared across tabs. - An active-BYOK banner with a "Stop" button.
- A synthetic "Integrated AI Model" entry appears in the Phase-1-style model dropdown only after a successful integration, letting the user switch between their own agent and catalog models via the existing model-switch confirmation — without discarding the stored credentials (only "Stop" clears them).
- The key is sent as
X-AI-BYOK-Tokenon each chat request, never in the request body — deliberate, so it's easier to guarantee exclusion from any request-body logging middleware.
4. Settings (paystream/app_settings/ai.py)
AI_BYOK_ENABLED— independent kill switch.AI_BYOK_REQUEST_TIMEOUT— timeout for the ephemeral provider's actual chat requests.AI_BYOK_VALIDATION_TIMEOUT— timeout for the modal's test call (shorter, since it's a synchronous UI wait).AI_BYOK_MAX_BASE_URL_LENGTH,AI_BYOK_MAX_MODEL_LENGTH,AI_BYOK_MAX_API_KEY_LENGTH— sanity length caps on the per-request header fields.
5. Acceptance criteria
- Disclosure text is shown and must be acknowledged before the modal accepts input.
- Key is never present in any Django log line, DB row, or Sentry/APM
breadcrumb — verified against
APIRequestLoggingMiddleware, which only ever persists path/method/status/user-agent, never headers or bodies. - Opening the same account in a second tab does not expose or reuse the BYOK key entered in the first tab.
- Closing the tab ends BYOK availability for that tab.
- Existing
ai_access_for()gating still applies — BYOK doesn't bypass who's allowed to use the chatbot, only which provider serves them. - BYOK requests still pass through the account-level rate limits (later joined by a per-tab throttle and its own token-budget tier — see the usage-controls-and-hardening doc).
6. Related Files
aicore/services/byok.py— new.aicore/views/byok.py— new (/aicore/byok/validate/).aicore/views/chat.py— readsX-AI-BYOK-Token.aicore/services/chat_service.py—byok_credentialsparam onget_or_create_session/stream_reply.aicore/providers/registry.py—ProviderName.BYOKguard.paystream/app_settings/ai.py— BYOK settings.frontend/templates/base/chat_widget.html,frontend/static/assets/js/components/chat_widget.js— modal, banner,sessionStoragecredential handling, header attachment.