DjangoPlay — Operational Runbooks
This document contains practical operational runbooks for maintaining and troubleshooting a DjangoPlay deployment.
On this page ▾
- Table of Contents
- 3.1 Check DjangoPlay
- 3.2 Check Celery
- 3.3 Check Nginx
- 3.4 Check PostgreSQL
- 3.5 Check Redis
- 5.1 Application Not Responding
- 5.2 Django Configuration Error
- 5.3 Migration Failure
- 11.1 Login Failure
- 11.2 Authentication Redirect Problems
- 11.3 AuthX Availability
- 13.1 DjangoPlay Logs
- 13.2 Celery Logs
- 13.3 Nginx Logs
- 13.4 System-Level Diagnostics
- Step 1 — Identify the failed pipeline stage
- Step 2 — Check application state
- Step 3 — Inspect logs
- Step 4 — Validate Django
- Step 5 — Validate dependencies
- Step 6 — Validate migrations
- Step 7 — Restart only required services
- Step 8 — Verify externally
- 15.1 First Response Checklist
- 15.2 HTTP Error Investigation
- Diagnose before changing
- Prefer reversible actions
- Use logs as the primary evidence
- Preserve database integrity
- Treat external services as boundaries
- Keep production simple
- Keep historical artifacts out of runbooks
- Summary
Table of Contents
- 1. Overview
- 2. Production Runtime
- 3. Service Status
- 4. Application Deployment
- 5. Django Application Troubleshooting
- 6. Nginx Troubleshooting
- 7. Gunicorn Troubleshooting
- 8. Celery Troubleshooting
- 9. Redis Troubleshooting
- 10. PostgreSQL Troubleshooting
- 11. Authentication / AuthX Troubleshooting
- 12. Static Files and Assets
- 13. Logs and Diagnostics
- 14. Deployment Failure Recovery
- 15. Incident Response
- 16. Operational Principles
1. Overview
It is intentionally focused on current operational procedures, rather than historical investigations or resolved development issues.
The runbooks cover the main components of the DjangoPlay production stack:
Cloudflare
│
▼
Nginx
│
▼
Gunicorn
│
▼
DjangoPlay
│
├── PostgreSQL
├── Redis
├── Celery
└── AuthXThe production environment is intentionally small, so troubleshooting should begin with service health, logs, connectivity, and recent deployment changes before making configuration changes.
2. Production Runtime
The production deployment consists of:
| Component | Responsibility |
|---|---|
| Cloudflare | DNS, edge and CDN services |
| Nginx | Reverse proxy and TLS termination |
| Gunicorn | Django application server |
| DjangoPlay | Main Django application |
| Celery | Background task execution |
| Redis | Cache and Celery broker |
| PostgreSQL | Application database |
| AuthX | Identity and authentication service |
| Cloudflare R2 | Object storage / supported assets |
| Cloudflare Turnstile | Bot and abuse protection |
The application and Celery services are managed by systemd.
3. Service Status
3.1 Check DjangoPlay
sudo systemctl status djangoplayRestart:
sudo systemctl restart djangoplayCheck whether it is enabled:
sudo systemctl is-enabled djangoplay3.2 Check Celery
sudo systemctl status djangoplay-celeryRestart:
sudo systemctl restart djangoplay-celery3.3 Check Nginx
sudo systemctl status nginxValidate configuration before restarting:
sudo nginx -tRestart:
sudo systemctl restart nginx3.4 Check PostgreSQL
sudo systemctl status postgresql3.5 Check Redis
sudo systemctl status redisThe exact service name can vary by distribution/package installation.
4. Application Deployment
DjangoPlay production deployment is performed through GitLab CI/CD.
The production deployment pipeline:
GitLab
│
▼
Manual production deployment
│
▼
SSH to production server
│
▼
Pull main
│
▼
Install dependencies
│
▼
Django migrations
│
▼
Collect static files
│
▼
Restart DjangoPlay
│
▼
Restart CeleryThe deployment job performs the equivalent of:
git checkout main
git pull origin main
pip install -e '.[dev]'
pip install djangoplay-cli
cd webapp
python manage.py migrate --noinput
python manage.py collectstatic --noinput
sudo systemctl restart djangoplay
sudo systemctl restart djangoplay-celeryThe production deployment environment is:
https://app.djangoplay.org5. Django Application Troubleshooting
5.1 Application Not Responding
Start with service status:
sudo systemctl status djangoplayThen inspect recent logs:
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pagerCheck whether Gunicorn is listening on its configured socket/port:
sudo ss -lntpIf the service has failed:
sudo systemctl restart djangoplayThen immediately inspect the logs again.
5.2 Django Configuration Error
Run Django checks from the application environment:
cd <webapp-directory>
source .venv/bin/activate
python manage.py checkFor deployment-related configuration problems, also verify:
- Environment configuration
- Database connectivity
- Redis connectivity
- Allowed hosts
- Static-file configuration
- AuthX configuration
- Cloudflare-related configuration
Do not modify production configuration blindly. First identify the failing setting from the application traceback.
5.3 Migration Failure
Check migration state:
python manage.py showmigrationsRun:
python manage.py migrateIf Django reports conflicting migration branches, inspect the migration graph before attempting a merge.
Do not delete migration files from production as a troubleshooting shortcut.
6. Nginx Troubleshooting
When the application is unreachable through the public domain, determine whether the problem is:
Cloudflare
↓
Nginx
↓
Gunicorn
↓
DjangoFirst validate Nginx:
sudo nginx -tThen inspect service status:
sudo systemctl status nginxInspect recent logs:
sudo journalctl -u nginx --since "30 minutes ago" --no-pager -lFor access/error logs:
sudo tail -f /var/log/nginx/access.log
sudo tail -f /var/log/nginx/error.logIf Nginx configuration is valid but the upstream is unavailable, continue troubleshooting Gunicorn/Django rather than repeatedly restarting Nginx.
7. Gunicorn Troubleshooting
Gunicorn runs the Django application behind Nginx.
Check the DjangoPlay service:
sudo systemctl status djangoplayInspect logs:
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager -lCommon failure categories include:
- Python import errors
- Missing dependencies
- Invalid environment configuration
- Django startup errors
- Database connection failures
- Redis connection failures
- Application exceptions during startup
After correcting the underlying issue:
sudo systemctl restart djangoplayThen verify:
sudo systemctl status djangoplay8. Celery Troubleshooting
Celery handles asynchronous DjangoPlay workloads.
Architecture:
DjangoPlay
│
▼
Redis Broker
│
▼
Celery Worker
│
▼
Background TaskCheck the worker:
sudo systemctl status djangoplay-celeryInspect logs:
sudo journalctl -u djangoplay-celery --since "30 minutes ago" --no-pagerRestart:
sudo systemctl restart djangoplay-celeryIf tasks are not executing, check Redis before changing Celery configuration.
9. Redis Troubleshooting
Redis is used by DjangoPlay for infrastructure services including caching and Celery task brokering.
Check service status:
sudo systemctl status redisTest connectivity:
redis-cli pingExpected response:
PONGIf Redis is unavailable:
- Check Redis service status.
- Inspect Redis logs.
- Verify the configured Redis URL.
- Restart Redis only after identifying the failure where practical.
- Restart dependent services if required.
After Redis recovery, verify Celery as well.
10. PostgreSQL Troubleshooting
PostgreSQL is the primary DjangoPlay application database.
Check PostgreSQL:
sudo systemctl status postgresqlCheck connectivity using the application's configured database credentials:
cd <webapp-directory>
source .venv/bin/activate
python manage.py checkIf Django reports database connectivity errors, investigate:
- PostgreSQL service status
- Database availability
- Credentials
- Host/port configuration
- Database permissions
- Connection limits
- Recent configuration changes
Do not modify database schema manually unless the operation is part of an intentional recovery procedure.
For schema changes, use Django migrations.
11. Authentication / AuthX Troubleshooting
DjangoPlay delegates identity and authentication functionality to AuthX.
Authentication problems should therefore be separated into two categories:
Browser / DjangoPlay
│
▼
DjangoPlay authentication integration
│
▼
AuthX
│
▼
Identity operation11.1 Login Failure
Determine whether:
- DjangoPlay is reachable.
- The login endpoint loads.
- AuthX is reachable.
- The configured AuthX endpoint is correct.
- The authentication request is reaching AuthX.
- The returned authentication response is being processed correctly.
Inspect DjangoPlay logs:
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pagerDo not treat a generic browser error as proof that AuthX itself is unavailable.
11.2 Authentication Redirect Problems
Check:
- Site URL configuration
- Authentication callback/redirect configuration
- Hostname configuration
- HTTPS configuration
- AuthX integration settings
For production, ensure redirects use the production host rather than local development hosts or ports.
11.3 AuthX Availability
If multiple DjangoPlay authentication operations fail simultaneously, test AuthX independently from application-level failures.
The objective is to establish whether the failure is:
DjangoPlay
│
├── Integration/configuration problem
│
└── AuthX availability/service problemOnly after that distinction should configuration or application code be changed.
12. Static Files and Assets
After deploying changes affecting static assets:
cd webapp
source .venv/bin/activate
python manage.py collectstatic --noinputIf assets are missing:
- Verify
collectstaticcompleted successfully. - Check the generated static directory.
- Check Nginx static-file configuration.
- Check Cloudflare caching if the request passes through Cloudflare.
- Inspect browser/network requests for the failing asset.
Avoid clearing caches as the first response. Establish whether the asset actually exists and is being served correctly.
13. Logs and Diagnostics
13.1 DjangoPlay Logs
sudo journalctl -u djangoplay --no-pagerRecent logs:
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pagerFollow live logs:
sudo journalctl -u djangoplay -f13.2 Celery Logs
sudo journalctl -u djangoplay-celery --no-pagerFollow live logs:
sudo journalctl -u djangoplay-celery -f13.3 Nginx Logs
sudo journalctl -u nginx --since "30 minutes ago" --no-pager -lAnd:
sudo tail -f /var/log/nginx/error.log13.4 System-Level Diagnostics
Check running services:
systemctl --type=service --state=runningCheck listening ports:
sudo ss -lntpCheck disk usage:
df -hCheck memory:
free -hCheck system load:
uptimeThese checks are particularly important on the small production VM.
14. Deployment Failure Recovery
If a deployment fails, do not immediately perform unrelated infrastructure changes.
Use the following sequence:
Step 1 — Identify the failed pipeline stage
Determine whether failure occurred during:
- Git pull
- Dependency installation
- Migration
- Static collection
- Service restart
Step 2 — Check application state
sudo systemctl status djangoplay
sudo systemctl status djangoplay-celeryStep 3 — Inspect logs
sudo journalctl -u djangoplay --since "30 minutes ago" --no-pager -l
sudo journalctl -u djangoplay-celery --since "30 minutes ago" --no-pager -lStep 4 — Validate Django
python manage.py checkStep 5 — Validate dependencies
Ensure the production virtual environment contains the expected application dependencies.
Step 6 — Validate migrations
python manage.py showmigrationsStep 7 — Restart only required services
sudo systemctl restart djangoplay
sudo systemctl restart djangoplay-celeryStep 8 — Verify externally
Check the production application through:
https://app.djangoplay.org15. Incident Response
For a production incident, use a simple progression:
Detect
│
▼
Confirm
│
▼
Classify
│
├── Cloudflare / DNS
├── Nginx
├── Django / Gunicorn
├── PostgreSQL
├── Redis / Celery
├── AuthX
└── External Integration
│
▼
Recover
│
▼
Verify
│
▼
Document15.1 First Response Checklist
[ ] Confirm the issue is reproducible
[ ] Check DjangoPlay service
[ ] Check Celery service
[ ] Check Nginx
[ ] Check PostgreSQL
[ ] Check Redis
[ ] Inspect recent application logs
[ ] Inspect recent deployment changes
[ ] Check external dependencies when relevant
[ ] Identify the failing boundary
[ ] Apply the smallest appropriate recovery action
[ ] Verify the application15.2 HTTP Error Investigation
When an HTTP error occurs, identify which layer generated it.
Client
│
▼
Cloudflare
│
▼
Nginx
│
▼
Gunicorn
│
▼
DjangoPlay
│
▼
AuthX / Database / Redis / External APIDo not assume that an HTTP status code alone identifies the root cause.
Use:
- Browser/network response
- Cloudflare information where applicable
- Nginx access/error logs
- Django logs
- Application configuration
- External service status
to locate the failing layer.
16. Operational Principles
DjangoPlay operational troubleshooting follows several principles:
Diagnose before changing
Identify the failing component before restarting or modifying configuration.
Prefer reversible actions
Restarting a service is generally preferable to changing persistent configuration during initial diagnosis.
Use logs as the primary evidence
Application and system logs should establish the failure before corrective action is taken.
Preserve database integrity
Do not bypass Django migrations or modify production schema casually.
Treat external services as boundaries
AuthX, Cloudflare, email providers, AI providers, and other external APIs should be diagnosed separately from DjangoPlay itself.
Keep production simple
The current deployment intentionally uses a small infrastructure footprint. Operational simplicity is part of the architecture.
Keep historical artifacts out of runbooks
Resolved investigations, temporary diagnostic values, obsolete blockers, and incident-specific artifacts should not become permanent operational procedures.
Summary
This runbook provides the operational starting point for DjangoPlay production incidents.
The primary troubleshooting path is:
Service Health
↓
Logs
↓
Dependency Connectivity
↓
Application Configuration
↓
Recent Deployment Changes
↓
Targeted Recovery
↓
VerificationDetailed architecture and deployment documentation should be consulted for system design and deployment procedures; this document is intended primarily as an operational troubleshooting reference.