--- since: 1.2.2 --- # Free OpenRouter Model Catalog & Selection Depends on: nothing (builds on the existing `aicore` app and provider registry) ## 1. Goal Let a user pick which free AI model powers their chat session, populated from OpenRouter's live model catalog, filtered to free + text-only models. ## 2. Decisions - Platform admins can restrict the catalog with an allowlist/blocklist of specific free models, on top of "every free, text-only model OpenRouter currently returns." - The previously-active model is pre-selected the next time the widget opens (read from `localStorage`), rather than always resetting to a platform default. - Model selection is **locked for the lifetime of a chat session**. Switching models mid-conversation is only allowed by starting a new session — never a silent mid-session provider/model mutation. ## 3. Design ### 3.1 Provider - New `openrouter` entry in `aicore.constants.ProviderName`. No new provider class — OpenRouter's `/chat/completions` is OpenAI-compatible, so the existing `OpenAICompatProvider` is reused as-is (`aicore/providers/registry.py`). - `openrouter` is deliberately **not** added to `resolve_provider_name`'s default fallback chain (`xai -> custom -> openai`) — it's only selected when a user explicitly picks a catalog model, never silently becomes the platform default. ### 3.2 Catalog service (`aicore/services/model_catalog.py`) - `get_free_models()` fetches `GET {AI_OPENROUTER_BASE_URL}/models` via `requests`, filters to: - free (`pricing.prompt == 0 and pricing.completion == 0`) - text-only (`"text" in input_modalities` and `output_modalities == ["text"]`) - Applies optional admin curation on top (`_apply_admin_curation`): allowlist wins if non-empty, otherwise blocklist removes specific IDs. - Result is cached server-side (Django cache framework, key `aicore:model_catalog:free`, TTL = `AI_MODEL_CATALOG_CACHE_TTL`, default 30 min) so the catalog isn't re-fetched from OpenRouter on every user's dropdown open. - On upstream failure, falls back to a much longer-lived stale cache (`aicore:model_catalog:free:stale`, TTL 24× the normal TTL) if one exists; otherwise returns an empty list with a short error string — never crashes the chat widget over a catalog fetch failure. - `is_known_free_model(model_id)` validates a client-supplied model id against the current catalog before `ChatService.get_or_create_session` trusts it — prevents a client from smuggling an arbitrary non-free/non-catalog model id through the session-creation parameter. ### 3.3 Endpoint `GET /api/v1/ai/models/` (`login_required`, JSON response): ```json { "models": [ {"id": "meta-llama/llama-3.2-3b-instruct:free", "name": "Llama 3.2 3B (free)", "context_length": 131072} ], "fetched_at": "2026-08-23T10:00:00Z" } ``` ### 3.4 Session / model-change behavior - `ChatService.get_or_create_session` accepts an optional `model_id` (implicitly `provider="openrouter"` when a catalog model is chosen). - No mid-session model swap: if the frontend requests a model change on an active session, the backend never mutates `session.model_name` — it always creates a **new** `ChatSession` (the same mechanism the existing "new chat" action uses), now parameterized with the newly selected model. - No new database fields were required — `ChatSession.provider` / `model_name` already supported arbitrary provider/model strings. ### 3.5 Frontend - `chat_widget.html` has a `` to the previously active model. - The last-picked model persists across visits via `localStorage`. ## 4. Settings (`paystream/app_settings/ai.py`) ```python AI_OPENROUTER_API_KEY = get_decrypted_value("OPENROUTER_API_KEY", default="") or "" AI_OPENROUTER_BASE_URL = ( get_decrypted_value("AI_OPENROUTER_BASE_URL", default="https://openrouter.ai/api/v1") or "https://openrouter.ai/api/v1" ).rstrip("/") AI_OPENROUTER_MODEL = get_decrypted_value("AI_OPENROUTER_MODEL", default="") or "" AI_MODEL_CATALOG_CACHE_TTL = int(get_decrypted_value("AI_MODEL_CATALOG_CACHE_TTL", default="1800") or "1800") AI_MODEL_CATALOG_ALLOWLIST = [...] # comma-separated OpenRouter model IDs AI_MODEL_CATALOG_BLOCKLIST = [...] # comma-separated OpenRouter model IDs ``` Note the env var for the API key is `OPENROUTER_API_KEY` (not `AI_OPENROUTER_API_KEY`) — a deliberate choice following the project's existing env var naming convention, learned the hard way from an earlier production settings-naming bug. `AI_OPENROUTER_MODEL` can stay blank — it's only a fallback default; real model IDs come from the catalog/user selection at session-create time. ## 5. Acceptance criteria - [x] Dropdown shows only free, text-in/text-out OpenRouter models. - [x] Dropdown is populated once per widget open, not re-fetched on every open/close cycle within the same page load. - [x] Selecting a different model while a conversation has messages triggers a confirmation warning about losing memory. - [x] Confirming starts a genuinely new `ChatSession` with the new provider/model; declining leaves the current session untouched. - [x] Existing rate limits and token budget continue to apply to OpenRouter sessions (later partitioned into its own budget tier — see the usage-controls-and-hardening doc). - [x] OpenRouter catalog fetch failures degrade gracefully (cached/stale list or clear error), never crash the chat widget. - [x] No change to `AI_ENABLED` / `AI_ENABLED_FOR_SSO_USERS` access gating — this feature only affects users who already have chat access. - [x] Admin allowlist/blocklist curation works on top of the auto-fetched free-model list. - [x] Previously-active model is pre-selected on next widget open. ## 6. Related Files - `aicore/constants.py` — `ProviderName.OPENROUTER`. - `aicore/providers/registry.py` — `openrouter` credential check and provider construction. - `aicore/services/model_catalog.py` — new. - `aicore/services/chat_service.py` — `model_id` param on `get_or_create_session`, validated via `is_known_free_model`. - `aicore/views/*` — new models endpoint. - `paystream/app_settings/ai.py` — new settings. - `frontend/templates/base/chat_widget.html`, `frontend/static/assets/js/components/chat_widget.js` — dropdown, fetch-once caching, confirm-to-switch flow, `localStorage` persistence.