--- since: 1.2.1 --- # DjangoPlay — Operational Runbooks --- ## Table of Contents * [1. Overview](#1-overview) * [2. Production Runtime](#2-production-runtime) * [3. Service Status](#3-service-status) * [4. Application Deployment](#4-application-deployment) * [5. Django Application Troubleshooting](#5-django-application-troubleshooting) * [6. Nginx Troubleshooting](#6-nginx-troubleshooting) * [7. Gunicorn Troubleshooting](#7-gunicorn-troubleshooting) * [8. Celery Troubleshooting](#8-celery-troubleshooting) * [9. Redis Troubleshooting](#9-redis-troubleshooting) * [10. PostgreSQL Troubleshooting](#10-postgresql-troubleshooting) * [11. Authentication / AuthX Troubleshooting](#11-authentication--authx-troubleshooting) * [12. Static Files and Assets](#12-static-files-and-assets) * [13. Logs and Diagnostics](#13-logs-and-diagnostics) * [14. Deployment Failure Recovery](#14-deployment-failure-recovery) * [15. Incident Response](#15-incident-response) * [16. Operational Principles](#16-operational-principles) --- # 1. Overview This document contains practical operational runbooks for maintaining and troubleshooting a DjangoPlay deployment. It is intentionally focused on **current operational procedures**, rather than historical investigations or resolved development issues. The runbooks cover the main components of the DjangoPlay production stack: ```text Cloudflare │ ▼ Nginx │ ▼ Gunicorn │ ▼ DjangoPlay │ ├── PostgreSQL ├── Redis ├── Celery └── AuthX ``` The production environment is intentionally small, so troubleshooting should begin with service health, logs, connectivity, and recent deployment changes before making configuration changes. --- # 2. Production Runtime The production deployment consists of: | Component | Responsibility | | -------------------- | ----------------------------------- | | Cloudflare | DNS, edge and CDN services | | Nginx | Reverse proxy and TLS termination | | Gunicorn | Django application server | | DjangoPlay | Main Django application | | Celery | Background task execution | | Redis | Cache and Celery broker | | PostgreSQL | Application database | | AuthX | Identity and authentication service | | Cloudflare R2 | Object storage / supported assets | | Cloudflare Turnstile | Bot and abuse protection | The application and Celery services are managed by systemd. --- # 3. Service Status ## 3.1 Check DjangoPlay ```bash sudo systemctl status djangoplay ``` Restart: ```bash sudo systemctl restart djangoplay ``` Check whether it is enabled: ```bash sudo systemctl is-enabled djangoplay ``` --- ## 3.2 Check Celery ```bash sudo systemctl status djangoplay-celery ``` Restart: ```bash sudo systemctl restart djangoplay-celery ``` --- ## 3.3 Check Nginx ```bash sudo systemctl status nginx ``` Validate configuration before restarting: ```bash sudo nginx -t ``` Restart: ```bash sudo systemctl restart nginx ``` --- ## 3.4 Check PostgreSQL ```bash sudo systemctl status postgresql ``` --- ## 3.5 Check Redis ```bash sudo systemctl status redis ``` The exact service name can vary by distribution/package installation. --- # 4. Application Deployment DjangoPlay production deployment is performed through **GitLab CI/CD**. The production deployment pipeline: ```text GitLab │ ▼ Manual production deployment │ ▼ SSH to production server │ ▼ Pull main │ ▼ Install dependencies │ ▼ Django migrations │ ▼ Collect static files │ ▼ Restart DjangoPlay │ ▼ Restart Celery ``` The deployment job performs the equivalent of: ```bash git checkout main git pull origin main pip install -e '.[dev]' pip install djangoplay-cli cd webapp python manage.py migrate --noinput python manage.py collectstatic --noinput sudo systemctl restart djangoplay sudo systemctl restart djangoplay-celery ``` The production deployment environment is: ```text https://app.djangoplay.org ``` --- # 5. Django Application Troubleshooting ## 5.1 Application Not Responding Start with service status: ```bash sudo systemctl status djangoplay ``` Then inspect recent logs: ```bash sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager ``` Check whether Gunicorn is listening on its configured socket/port: ```bash sudo ss -lntp ``` If the service has failed: ```bash sudo systemctl restart djangoplay ``` Then immediately inspect the logs again. --- ## 5.2 Django Configuration Error Run Django checks from the application environment: ```bash cd source .venv/bin/activate python manage.py check ``` For deployment-related configuration problems, also verify: * Environment configuration * Database connectivity * Redis connectivity * Allowed hosts * Static-file configuration * AuthX configuration * Cloudflare-related configuration Do not modify production configuration blindly. First identify the failing setting from the application traceback. --- ## 5.3 Migration Failure Check migration state: ```bash python manage.py showmigrations ``` Run: ```bash python manage.py migrate ``` If Django reports conflicting migration branches, inspect the migration graph before attempting a merge. Do not delete migration files from production as a troubleshooting shortcut. --- # 6. Nginx Troubleshooting When the application is unreachable through the public domain, determine whether the problem is: ```text Cloudflare ↓ Nginx ↓ Gunicorn ↓ Django ``` First validate Nginx: ```bash sudo nginx -t ``` Then inspect service status: ```bash sudo systemctl status nginx ``` Inspect recent logs: ```bash sudo journalctl -u nginx --since "30 minutes ago" --no-pager -l ``` For access/error logs: ```bash sudo tail -f /var/log/nginx/access.log sudo tail -f /var/log/nginx/error.log ``` If Nginx configuration is valid but the upstream is unavailable, continue troubleshooting Gunicorn/Django rather than repeatedly restarting Nginx. --- # 7. Gunicorn Troubleshooting Gunicorn runs the Django application behind Nginx. Check the DjangoPlay service: ```bash sudo systemctl status djangoplay ``` Inspect logs: ```bash sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager -l ``` Common failure categories include: * Python import errors * Missing dependencies * Invalid environment configuration * Django startup errors * Database connection failures * Redis connection failures * Application exceptions during startup After correcting the underlying issue: ```bash sudo systemctl restart djangoplay ``` Then verify: ```bash sudo systemctl status djangoplay ``` --- # 8. Celery Troubleshooting Celery handles asynchronous DjangoPlay workloads. Architecture: ```text DjangoPlay │ ▼ Redis Broker │ ▼ Celery Worker │ ▼ Background Task ``` Check the worker: ```bash sudo systemctl status djangoplay-celery ``` Inspect logs: ```bash sudo journalctl -u djangoplay-celery --since "30 minutes ago" --no-pager ``` Restart: ```bash sudo systemctl restart djangoplay-celery ``` If tasks are not executing, check Redis before changing Celery configuration. --- # 9. Redis Troubleshooting Redis is used by DjangoPlay for infrastructure services including caching and Celery task brokering. Check service status: ```bash sudo systemctl status redis ``` Test connectivity: ```bash redis-cli ping ``` Expected response: ```text PONG ``` If Redis is unavailable: 1. Check Redis service status. 2. Inspect Redis logs. 3. Verify the configured Redis URL. 4. Restart Redis only after identifying the failure where practical. 5. Restart dependent services if required. After Redis recovery, verify Celery as well. --- # 10. PostgreSQL Troubleshooting PostgreSQL is the primary DjangoPlay application database. Check PostgreSQL: ```bash sudo systemctl status postgresql ``` Check connectivity using the application's configured database credentials: ```bash cd source .venv/bin/activate python manage.py check ``` If Django reports database connectivity errors, investigate: * PostgreSQL service status * Database availability * Credentials * Host/port configuration * Database permissions * Connection limits * Recent configuration changes Do not modify database schema manually unless the operation is part of an intentional recovery procedure. For schema changes, use Django migrations. --- # 11. Authentication / AuthX Troubleshooting DjangoPlay delegates identity and authentication functionality to **AuthX**. Authentication problems should therefore be separated into two categories: ```text Browser / DjangoPlay │ ▼ DjangoPlay authentication integration │ ▼ AuthX │ ▼ Identity operation ``` ## 11.1 Login Failure Determine whether: * DjangoPlay is reachable. * The login endpoint loads. * AuthX is reachable. * The configured AuthX endpoint is correct. * The authentication request is reaching AuthX. * The returned authentication response is being processed correctly. Inspect DjangoPlay logs: ```bash sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager ``` Do not treat a generic browser error as proof that AuthX itself is unavailable. --- ## 11.2 Authentication Redirect Problems Check: * Site URL configuration * Authentication callback/redirect configuration * Hostname configuration * HTTPS configuration * AuthX integration settings For production, ensure redirects use the production host rather than local development hosts or ports. --- ## 11.3 AuthX Availability If multiple DjangoPlay authentication operations fail simultaneously, test AuthX independently from application-level failures. The objective is to establish whether the failure is: ```text DjangoPlay │ ├── Integration/configuration problem │ └── AuthX availability/service problem ``` Only after that distinction should configuration or application code be changed. --- # 12. Static Files and Assets After deploying changes affecting static assets: ```bash cd webapp source .venv/bin/activate python manage.py collectstatic --noinput ``` If assets are missing: 1. Verify `collectstatic` completed successfully. 2. Check the generated static directory. 3. Check Nginx static-file configuration. 4. Check Cloudflare caching if the request passes through Cloudflare. 5. Inspect browser/network requests for the failing asset. Avoid clearing caches as the first response. Establish whether the asset actually exists and is being served correctly. --- # 13. Logs and Diagnostics ## 13.1 DjangoPlay Logs ```bash sudo journalctl -u djangoplay --no-pager ``` Recent logs: ```bash sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager ``` Follow live logs: ```bash sudo journalctl -u djangoplay -f ``` --- ## 13.2 Celery Logs ```bash sudo journalctl -u djangoplay-celery --no-pager ``` Follow live logs: ```bash sudo journalctl -u djangoplay-celery -f ``` --- ## 13.3 Nginx Logs ```bash sudo journalctl -u nginx --since "30 minutes ago" --no-pager -l ``` And: ```bash sudo tail -f /var/log/nginx/error.log ``` --- ## 13.4 System-Level Diagnostics Check running services: ```bash systemctl --type=service --state=running ``` Check listening ports: ```bash sudo ss -lntp ``` Check disk usage: ```bash df -h ``` Check memory: ```bash free -h ``` Check system load: ```bash uptime ``` These checks are particularly important on the small production VM. --- # 14. Deployment Failure Recovery If a deployment fails, do not immediately perform unrelated infrastructure changes. Use the following sequence: ### Step 1 — Identify the failed pipeline stage Determine whether failure occurred during: * Git pull * Dependency installation * Migration * Static collection * Service restart ### Step 2 — Check application state ```bash sudo systemctl status djangoplay sudo systemctl status djangoplay-celery ``` ### Step 3 — Inspect logs ```bash sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager -l sudo journalctl -u djangoplay-celery --since "30 minutes ago" --no-pager -l ``` ### Step 4 — Validate Django ```bash python manage.py check ``` ### Step 5 — Validate dependencies Ensure the production virtual environment contains the expected application dependencies. ### Step 6 — Validate migrations ```bash python manage.py showmigrations ``` ### Step 7 — Restart only required services ```bash sudo systemctl restart djangoplay sudo systemctl restart djangoplay-celery ``` ### Step 8 — Verify externally Check the production application through: ```text https://app.djangoplay.org ``` --- # 15. Incident Response For a production incident, use a simple progression: ```text Detect │ ▼ Confirm │ ▼ Classify │ ├── Cloudflare / DNS ├── Nginx ├── Django / Gunicorn ├── PostgreSQL ├── Redis / Celery ├── AuthX └── External Integration │ ▼ Recover │ ▼ Verify │ ▼ Document ``` ## 15.1 First Response Checklist ```text [ ] Confirm the issue is reproducible [ ] Check DjangoPlay service [ ] Check Celery service [ ] Check Nginx [ ] Check PostgreSQL [ ] Check Redis [ ] Inspect recent application logs [ ] Inspect recent deployment changes [ ] Check external dependencies when relevant [ ] Identify the failing boundary [ ] Apply the smallest appropriate recovery action [ ] Verify the application ``` --- ## 15.2 HTTP Error Investigation When an HTTP error occurs, identify which layer generated it. ```text Client │ ▼ Cloudflare │ ▼ Nginx │ ▼ Gunicorn │ ▼ DjangoPlay │ ▼ AuthX / Database / Redis / External API ``` Do not assume that an HTTP status code alone identifies the root cause. Use: * Browser/network response * Cloudflare information where applicable * Nginx access/error logs * Django logs * Application configuration * External service status to locate the failing layer. --- # 16. Operational Principles DjangoPlay operational troubleshooting follows several principles: ### Diagnose before changing Identify the failing component before restarting or modifying configuration. ### Prefer reversible actions Restarting a service is generally preferable to changing persistent configuration during initial diagnosis. ### Use logs as the primary evidence Application and system logs should establish the failure before corrective action is taken. ### Preserve database integrity Do not bypass Django migrations or modify production schema casually. ### Treat external services as boundaries AuthX, Cloudflare, email providers, AI providers, and other external APIs should be diagnosed separately from DjangoPlay itself. ### Keep production simple The current deployment intentionally uses a small infrastructure footprint. Operational simplicity is part of the architecture. ### Keep historical artifacts out of runbooks Resolved investigations, temporary diagnostic values, obsolete blockers, and incident-specific artifacts should not become permanent operational procedures. --- ## Summary This runbook provides the operational starting point for DjangoPlay production incidents. The primary troubleshooting path is: ```text Service Health ↓ Logs ↓ Dependency Connectivity ↓ Application Configuration ↓ Recent Deployment Changes ↓ Targeted Recovery ↓ Verification ``` Detailed architecture and deployment documentation should be consulted for system design and deployment procedures; this document is intended primarily as an **operational troubleshooting reference**.