djangoplay-web / Runbooks / DjangoPlay — Operational Runbooks
DocsDjangoPlay WebRunbooksDjangoPlay — Operational Runbooks

DjangoPlay — Operational Runbooks

This document contains practical operational runbooks for maintaining and troubleshooting a DjangoPlay deployment.

10 min readApplies to v1.2.2
On this page ▾
  1. Table of Contents
  2. 3.1 Check DjangoPlay
  3. 3.2 Check Celery
  4. 3.3 Check Nginx
  5. 3.4 Check PostgreSQL
  6. 3.5 Check Redis
  7. 5.1 Application Not Responding
  8. 5.2 Django Configuration Error
  9. 5.3 Migration Failure
  10. 11.1 Login Failure
  11. 11.2 Authentication Redirect Problems
  12. 11.3 AuthX Availability
  13. 13.1 DjangoPlay Logs
  14. 13.2 Celery Logs
  15. 13.3 Nginx Logs
  16. 13.4 System-Level Diagnostics
  17. Step 1 — Identify the failed pipeline stage
  18. Step 2 — Check application state
  19. Step 3 — Inspect logs
  20. Step 4 — Validate Django
  21. Step 5 — Validate dependencies
  22. Step 6 — Validate migrations
  23. Step 7 — Restart only required services
  24. Step 8 — Verify externally
  25. 15.1 First Response Checklist
  26. 15.2 HTTP Error Investigation
  27. Diagnose before changing
  28. Prefer reversible actions
  29. Use logs as the primary evidence
  30. Preserve database integrity
  31. Treat external services as boundaries
  32. Keep production simple
  33. Keep historical artifacts out of runbooks
  34. Summary

Table of Contents


1. Overview

It is intentionally focused on current operational procedures, rather than historical investigations or resolved development issues.

The runbooks cover the main components of the DjangoPlay production stack:

text
Cloudflare
    │
    ▼
Nginx
    │
    ▼
Gunicorn
    │
    ▼
DjangoPlay
    │
    ├── PostgreSQL
    ├── Redis
    ├── Celery
    └── AuthX

The production environment is intentionally small, so troubleshooting should begin with service health, logs, connectivity, and recent deployment changes before making configuration changes.


2. Production Runtime

The production deployment consists of:

Component Responsibility
Cloudflare DNS, edge and CDN services
Nginx Reverse proxy and TLS termination
Gunicorn Django application server
DjangoPlay Main Django application
Celery Background task execution
Redis Cache and Celery broker
PostgreSQL Application database
AuthX Identity and authentication service
Cloudflare R2 Object storage / supported assets
Cloudflare Turnstile Bot and abuse protection

The application and Celery services are managed by systemd.


3. Service Status

3.1 Check DjangoPlay

bash
sudo systemctl status djangoplay

Restart:

bash
sudo systemctl restart djangoplay

Check whether it is enabled:

bash
sudo systemctl is-enabled djangoplay

3.2 Check Celery

bash
sudo systemctl status djangoplay-celery

Restart:

bash
sudo systemctl restart djangoplay-celery

3.3 Check Nginx

bash
sudo systemctl status nginx

Validate configuration before restarting:

bash
sudo nginx -t

Restart:

bash
sudo systemctl restart nginx

3.4 Check PostgreSQL

bash
sudo systemctl status postgresql

3.5 Check Redis

bash
sudo systemctl status redis

The exact service name can vary by distribution/package installation.


4. Application Deployment

DjangoPlay production deployment is performed through GitLab CI/CD.

The production deployment pipeline:

text
GitLab
   │
   ▼
Manual production deployment
   │
   ▼
SSH to production server
   │
   ▼
Pull main
   │
   ▼
Install dependencies
   │
   ▼
Django migrations
   │
   ▼
Collect static files
   │
   ▼
Restart DjangoPlay
   │
   ▼
Restart Celery

The deployment job performs the equivalent of:

bash
git checkout main
git pull origin main
pip install -e '.[dev]'
pip install djangoplay-cli

cd webapp
python manage.py migrate --noinput
python manage.py collectstatic --noinput

sudo systemctl restart djangoplay
sudo systemctl restart djangoplay-celery

The production deployment environment is:

text
https://app.djangoplay.org

5. Django Application Troubleshooting

5.1 Application Not Responding

Start with service status:

bash
sudo systemctl status djangoplay

Then inspect recent logs:

bash
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager

Check whether Gunicorn is listening on its configured socket/port:

bash
sudo ss -lntp

If the service has failed:

bash
sudo systemctl restart djangoplay

Then immediately inspect the logs again.


5.2 Django Configuration Error

Run Django checks from the application environment:

bash
cd <webapp-directory>
source .venv/bin/activate

python manage.py check

For deployment-related configuration problems, also verify:

  • Environment configuration
  • Database connectivity
  • Redis connectivity
  • Allowed hosts
  • Static-file configuration
  • AuthX configuration
  • Cloudflare-related configuration

Do not modify production configuration blindly. First identify the failing setting from the application traceback.


5.3 Migration Failure

Check migration state:

bash
python manage.py showmigrations

Run:

bash
python manage.py migrate

If Django reports conflicting migration branches, inspect the migration graph before attempting a merge.

Do not delete migration files from production as a troubleshooting shortcut.


6. Nginx Troubleshooting

When the application is unreachable through the public domain, determine whether the problem is:

text
Cloudflare
    ↓
Nginx
    ↓
Gunicorn
    ↓
Django

First validate Nginx:

bash
sudo nginx -t

Then inspect service status:

bash
sudo systemctl status nginx

Inspect recent logs:

bash
sudo journalctl -u nginx --since "30 minutes ago" --no-pager -l

For access/error logs:

bash
sudo tail -f /var/log/nginx/access.log
sudo tail -f /var/log/nginx/error.log

If Nginx configuration is valid but the upstream is unavailable, continue troubleshooting Gunicorn/Django rather than repeatedly restarting Nginx.


7. Gunicorn Troubleshooting

Gunicorn runs the Django application behind Nginx.

Check the DjangoPlay service:

bash
sudo systemctl status djangoplay

Inspect logs:

bash
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager -l

Common failure categories include:

  • Python import errors
  • Missing dependencies
  • Invalid environment configuration
  • Django startup errors
  • Database connection failures
  • Redis connection failures
  • Application exceptions during startup

After correcting the underlying issue:

bash
sudo systemctl restart djangoplay

Then verify:

bash
sudo systemctl status djangoplay

8. Celery Troubleshooting

Celery handles asynchronous DjangoPlay workloads.

Architecture:

text
DjangoPlay
    │
    ▼
Redis Broker
    │
    ▼
Celery Worker
    │
    ▼
Background Task

Check the worker:

bash
sudo systemctl status djangoplay-celery

Inspect logs:

bash
sudo journalctl -u djangoplay-celery --since "30 minutes ago" --no-pager

Restart:

bash
sudo systemctl restart djangoplay-celery

If tasks are not executing, check Redis before changing Celery configuration.


9. Redis Troubleshooting

Redis is used by DjangoPlay for infrastructure services including caching and Celery task brokering.

Check service status:

bash
sudo systemctl status redis

Test connectivity:

bash
redis-cli ping

Expected response:

text
PONG

If Redis is unavailable:

  1. Check Redis service status.
  2. Inspect Redis logs.
  3. Verify the configured Redis URL.
  4. Restart Redis only after identifying the failure where practical.
  5. Restart dependent services if required.

After Redis recovery, verify Celery as well.


10. PostgreSQL Troubleshooting

PostgreSQL is the primary DjangoPlay application database.

Check PostgreSQL:

bash
sudo systemctl status postgresql

Check connectivity using the application's configured database credentials:

bash
cd <webapp-directory>
source .venv/bin/activate

python manage.py check

If Django reports database connectivity errors, investigate:

  • PostgreSQL service status
  • Database availability
  • Credentials
  • Host/port configuration
  • Database permissions
  • Connection limits
  • Recent configuration changes

Do not modify database schema manually unless the operation is part of an intentional recovery procedure.

For schema changes, use Django migrations.


11. Authentication / AuthX Troubleshooting

DjangoPlay delegates identity and authentication functionality to AuthX.

Authentication problems should therefore be separated into two categories:

text
Browser / DjangoPlay
        │
        ▼
DjangoPlay authentication integration
        │
        ▼
AuthX
        │
        ▼
Identity operation

11.1 Login Failure

Determine whether:

  • DjangoPlay is reachable.
  • The login endpoint loads.
  • AuthX is reachable.
  • The configured AuthX endpoint is correct.
  • The authentication request is reaching AuthX.
  • The returned authentication response is being processed correctly.

Inspect DjangoPlay logs:

bash
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager

Do not treat a generic browser error as proof that AuthX itself is unavailable.


11.2 Authentication Redirect Problems

Check:

  • Site URL configuration
  • Authentication callback/redirect configuration
  • Hostname configuration
  • HTTPS configuration
  • AuthX integration settings

For production, ensure redirects use the production host rather than local development hosts or ports.


11.3 AuthX Availability

If multiple DjangoPlay authentication operations fail simultaneously, test AuthX independently from application-level failures.

The objective is to establish whether the failure is:

text
DjangoPlay
   │
   ├── Integration/configuration problem
   │
   └── AuthX availability/service problem

Only after that distinction should configuration or application code be changed.


12. Static Files and Assets

After deploying changes affecting static assets:

bash
cd webapp
source .venv/bin/activate

python manage.py collectstatic --noinput

If assets are missing:

  1. Verify collectstatic completed successfully.
  2. Check the generated static directory.
  3. Check Nginx static-file configuration.
  4. Check Cloudflare caching if the request passes through Cloudflare.
  5. Inspect browser/network requests for the failing asset.

Avoid clearing caches as the first response. Establish whether the asset actually exists and is being served correctly.


13. Logs and Diagnostics

13.1 DjangoPlay Logs

bash
sudo journalctl -u djangoplay --no-pager

Recent logs:

bash
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager

Follow live logs:

bash
sudo journalctl -u djangoplay -f

13.2 Celery Logs

bash
sudo journalctl -u djangoplay-celery --no-pager

Follow live logs:

bash
sudo journalctl -u djangoplay-celery -f

13.3 Nginx Logs

bash
sudo journalctl -u nginx --since "30 minutes ago" --no-pager -l

And:

bash
sudo tail -f /var/log/nginx/error.log

13.4 System-Level Diagnostics

Check running services:

bash
systemctl --type=service --state=running

Check listening ports:

bash
sudo ss -lntp

Check disk usage:

bash
df -h

Check memory:

bash
free -h

Check system load:

bash
uptime

These checks are particularly important on the small production VM.


14. Deployment Failure Recovery

If a deployment fails, do not immediately perform unrelated infrastructure changes.

Use the following sequence:

Step 1 — Identify the failed pipeline stage

Determine whether failure occurred during:

  • Git pull
  • Dependency installation
  • Migration
  • Static collection
  • Service restart

Step 2 — Check application state

bash
sudo systemctl status djangoplay
sudo systemctl status djangoplay-celery

Step 3 — Inspect logs

bash
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager -l
sudo journalctl -u djangoplay-celery --since "30 minutes ago" --no-pager -l

Step 4 — Validate Django

bash
python manage.py check

Step 5 — Validate dependencies

Ensure the production virtual environment contains the expected application dependencies.

Step 6 — Validate migrations

bash
python manage.py showmigrations

Step 7 — Restart only required services

bash
sudo systemctl restart djangoplay
sudo systemctl restart djangoplay-celery

Step 8 — Verify externally

Check the production application through:

text
https://app.djangoplay.org

15. Incident Response

For a production incident, use a simple progression:

text
Detect
  │
  ▼
Confirm
  │
  ▼
Classify
  │
  ├── Cloudflare / DNS
  ├── Nginx
  ├── Django / Gunicorn
  ├── PostgreSQL
  ├── Redis / Celery
  ├── AuthX
  └── External Integration
  │
  ▼
Recover
  │
  ▼
Verify
  │
  ▼
Document

15.1 First Response Checklist

text
[ ] Confirm the issue is reproducible
[ ] Check DjangoPlay service
[ ] Check Celery service
[ ] Check Nginx
[ ] Check PostgreSQL
[ ] Check Redis
[ ] Inspect recent application logs
[ ] Inspect recent deployment changes
[ ] Check external dependencies when relevant
[ ] Identify the failing boundary
[ ] Apply the smallest appropriate recovery action
[ ] Verify the application

15.2 HTTP Error Investigation

When an HTTP error occurs, identify which layer generated it.

text
Client
  │
  ▼
Cloudflare
  │
  ▼
Nginx
  │
  ▼
Gunicorn
  │
  ▼
DjangoPlay
  │
  ▼
AuthX / Database / Redis / External API

Do not assume that an HTTP status code alone identifies the root cause.

Use:

  • Browser/network response
  • Cloudflare information where applicable
  • Nginx access/error logs
  • Django logs
  • Application configuration
  • External service status

to locate the failing layer.


16. Operational Principles

DjangoPlay operational troubleshooting follows several principles:

Diagnose before changing

Identify the failing component before restarting or modifying configuration.

Prefer reversible actions

Restarting a service is generally preferable to changing persistent configuration during initial diagnosis.

Use logs as the primary evidence

Application and system logs should establish the failure before corrective action is taken.

Preserve database integrity

Do not bypass Django migrations or modify production schema casually.

Treat external services as boundaries

AuthX, Cloudflare, email providers, AI providers, and other external APIs should be diagnosed separately from DjangoPlay itself.

Keep production simple

The current deployment intentionally uses a small infrastructure footprint. Operational simplicity is part of the architecture.

Keep historical artifacts out of runbooks

Resolved investigations, temporary diagnostic values, obsolete blockers, and incident-specific artifacts should not become permanent operational procedures.


Summary

This runbook provides the operational starting point for DjangoPlay production incidents.

The primary troubleshooting path is:

text
Service Health
      ↓
Logs
      ↓
Dependency Connectivity
      ↓
Application Configuration
      ↓
Recent Deployment Changes
      ↓
Targeted Recovery
      ↓
Verification

Detailed architecture and deployment documentation should be consulted for system design and deployment procedures; this document is intended primarily as an operational troubleshooting reference.