---
since: 1.2.1
---
# DjangoPlay — Scaling Architecture
**Date:** 2026-08-29
---
## 1. Overview
DjangoPlay is designed as a **modular monolith with independently scalable
runtime workloads**.
The application is not currently designed around a microservices topology.
Instead, scaling can be introduced incrementally by separating the workloads
that have different resource and traffic characteristics:
* Web/API request processing
* Background task processing
* Redis cache and task brokering
* PostgreSQL persistence
* Static/CDN asset delivery
* External service integrations
This allows DjangoPlay to scale without prematurely introducing distributed
application services.
The primary scaling strategy is therefore:
```text
Vertical Scaling
│
▼
Workload Separation
│
▼
Horizontal Scaling
│
▼
Database / Cache Optimization
│
▼
Selective Service Extraction
````
---
## 2. Current Runtime Architecture
The current DjangoPlay runtime consists of separate application and supporting
processes.
```mermaid
flowchart TD
CLIENT["Clients
Browser / API"]
NGINX["Nginx
Reverse Proxy"]
DJANGO["DjangoPlay
Web + DRF API"]
POSTGRES["PostgreSQL
Application Database"]
REDIS["Redis
Cache + Celery Broker"]
CELERY["Celery Worker
Background Tasks"]
AUTHX["AuthX
Identity Service"]
EXTERNAL["External Services
Email / R2 / AI / APIs"]
CLIENT --> NGINX
NGINX --> DJANGO
DJANGO --> POSTGRES
DJANGO --> REDIS
REDIS --> CELERY
CELERY --> POSTGRES
CELERY --> AUTHX
CELERY --> EXTERNAL
DJANGO --> AUTHX
DJANGO --> EXTERNAL
classDef client fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef application fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef infrastructure fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef external fill:#eeeeff,stroke:#5a55c9,stroke-width:2px,color:#403b91
class CLIENT client
class NGINX,DJANGO application
class POSTGRES,REDIS,CELERY infrastructure
class AUTHX,EXTERNAL external
```
The important architectural property is that **web traffic and asynchronous
work are already separate workloads**.
```text
DjangoPlay
│
┌─────────┴─────────┐
│ │
▼ ▼
Web / API Background
Processing Processing
│ │
▼ ▼
Gunicorn Celery
│ │
└─────────┬─────────┘
▼
Redis
│
▼
PostgreSQL
```
This separation provides the first scaling boundary without changing the
application architecture.
---
## 3. Scaling Dimensions
DjangoPlay can scale along several independent dimensions.
| Dimension | Primary Scaling Method |
| --------------------- | --------------------------------------------------------------------------------------------------------- |
| HTTP/API traffic | More Gunicorn workers and application instances |
| Background jobs | More Celery worker processes/instances |
| Database workload | Query optimization, indexes, connection management, larger PostgreSQL resources, replicas where justified |
| Cache/broker workload | Redis resource scaling or dedicated Redis infrastructure |
| Static assets | CDN / Cloudflare R2 |
| External integrations | Async processing, provider limits, retries, isolation |
| Application memory | Process sizing and workload separation |
| CPU-bound workloads | Worker/application process scaling |
| I/O-heavy workloads | Async/background execution and workload separation |
Not every workload should be scaled in the same way.
---
# 4. Vertical Scaling
Vertical scaling is the first and simplest scaling mechanism.
It means increasing resources available to the existing runtime:
* CPU
* RAM
* Disk performance
* Network capacity
* PostgreSQL resources
* Redis resources
* Gunicorn worker capacity
* Celery worker capacity
A typical progression is:
```text
Small Instance
│
▼
More CPU / RAM
│
▼
More Gunicorn Workers
│
▼
More Celery Workers
│
▼
Dedicated Resources
```
Vertical scaling is particularly appropriate while DjangoPlay remains a
single-node deployment.
### Advantages
* Minimal architectural change
* Simple operations
* No distributed application state
* No load-balancing complexity
* Low operational overhead
### Limitation
A single machine eventually becomes a capacity and availability boundary.
When that point is reached, horizontal scaling becomes the next step.
---
# 5. Web / API Scaling
DjangoPlay web and DRF traffic is handled by the Django application behind
Nginx and Gunicorn.
```text
Clients
│
▼
Nginx
│
▼
Gunicorn
│
┌────────┼────────┐
▼ ▼ ▼
Django Django Django
Worker Worker Worker
│ │ │
└────────┼────────┘
▼
PostgreSQL
```
On a single host, the first scaling step is increasing the number of
Gunicorn workers within the available CPU and memory limits.
As traffic grows beyond a single application process group, multiple
application instances can be introduced.
```text
Load Balancer
│
┌───────────┼───────────┐
▼ ▼ ▼
App #1 App #2 App #3
│ │ │
└───────────┼───────────┘
▼
PostgreSQL
```
The application layer should remain as stateless as practical so that
requests can be distributed across instances.
---
# 6. Celery Worker Scaling
Background processing is an independent scaling dimension.
DjangoPlay already separates asynchronous work through Celery.
```text
DjangoPlay
│
│ enqueue
▼
Redis
│
▼
Celery Workers
```
Workers can be increased independently of web/API capacity.
```text
Redis Broker
│
┌───────────┼───────────┐
▼ ▼ ▼
Worker #1 Worker #2 Worker #3
```
This is preferable to increasing web-server capacity merely because
background processing has increased.
### Worker Scaling Strategies
Celery capacity can be increased by:
* Increasing worker concurrency
* Running additional worker processes
* Running workers on separate hosts
* Separating workloads into queues
* Allocating different worker capacity to different task classes
Conceptually:
```text
Redis
│
┌────────┼────────┐
│ │ │
▼ ▼ ▼
Default Email Heavy
Queue Queue Queue
│ │ │
▼ ▼ ▼
Workers Workers Workers
```
Queue separation should be introduced when workload characteristics justify
it rather than as an architectural requirement from the beginning.
---
# 7. Web and Worker Isolation
One of the most important scaling characteristics of DjangoPlay is the
ability to scale request processing independently from background processing.
```text
DjangoPlay Runtime
│
┌─────────┴─────────┐
│ │
▼ ▼
Web/API Celery
Workload Workload
│ │
Gunicorn Workers
│ │
▼ ▼
User-facing Background
latency processing
```
For example:
* A large email workload should not require additional web workers.
* A large API traffic increase should not require additional Celery workers.
* A CPU-heavy background task should not consume the resources required by
request processing.
This separation provides an effective scaling boundary while preserving the
modular-monolith architecture.
---
# 8. Redis Scaling
Redis currently serves multiple runtime responsibilities:
* Cache
* Session/runtime support where configured
* Celery broker
```text
Redis
/ \
/ \
Cache Broker
│
▼
Celery
```
Redis should therefore be monitored independently from PostgreSQL.
Scaling options include:
* Increasing Redis memory
* Increasing host resources
* Separating cache and broker workloads when justified
* Moving Redis to dedicated infrastructure
* Introducing highly available Redis infrastructure when availability
requirements justify it
Redis should not become the primary persistence layer for application data.
PostgreSQL remains the authoritative persistent application datastore.
---
# 9. PostgreSQL Scaling
PostgreSQL is the primary persistence boundary for DjangoPlay application
data.
```text
DjangoPlay
│
▼
PostgreSQL
```
Database scaling should begin with optimization before introducing distributed
database infrastructure.
### First-level optimizations
* Query optimization
* Appropriate indexes
* Avoiding unnecessary queries
* Reducing N+1 query patterns
* Pagination
* Efficient queryset construction
* Appropriate transaction boundaries
* Connection management
### Vertical database scaling
Increase:
* CPU
* RAM
* Storage performance
* I/O capacity
### Advanced scaling
When workload justifies it:
```text
PostgreSQL
│
┌─────────┴─────────┐
▼ ▼
Primary Replica(s)
Read / Write Read
```
Read replicas should only be introduced when the application's read/write
patterns and consistency requirements make them useful.
They are not automatically beneficial for every DjangoPlay deployment.
---
# 10. Database Connection Scaling
Increasing application instances also increases the number of database
connections.
For example:
```text
1 App Instance
│
└── N database connections
10 App Instances
│
└── 10 × N database connections
```
Therefore horizontal application scaling must be accompanied by deliberate
database connection management.
Before adding application instances, evaluate:
* Gunicorn worker count
* Database connection count
* PostgreSQL `max_connections`
* Connection lifetime
* Query duration
* Connection pooling requirements
Database connection exhaustion can become a bottleneck before CPU or memory
does.
---
# 11. Caching Strategy
Caching should reduce repeated work and database pressure.
The primary DjangoPlay cache infrastructure is Redis.
```text
Request
│
▼
Application
│
▼
Cache Lookup
│
├── Hit ───────► Return Cached Result
│
└── Miss
│
▼
PostgreSQL
│
▼
Store Result
│
▼
Redis
```
Caching opportunities should be introduced based on measured application
behavior.
Potential targets include:
* Expensive read operations
* Stable reference data
* Frequently accessed computed results
* Rate-limit state
* Session/runtime data where applicable
Caching should not be used to hide inefficient database access indefinitely.
---
# 12. Static Assets and CDN Scaling
Frontend assets should not consume application-server capacity when they can
be delivered through the CDN path.
DjangoPlay supports Cloudflare R2 for the frontend asset workflow.
```text
Frontend Build
│
▼
Cloudflare R2
│
▼
CDN
│
▼
Browser
```
This moves static asset delivery away from Django/Gunicorn.
The local development configuration can serve frontend assets directly, while
the CDN workflow can publish versioned assets to R2.
The CDN therefore provides a natural scaling boundary:
```text
Application Server
│
└── Dynamic Requests
CDN / R2
│
└── Static Frontend Assets
```
---
# 13. External Integration Scaling
DjangoPlay communicates with external services including:
* AuthX
* SMTP/email providers
* Cloudflare R2
* AI providers
* Other external APIs
These services have their own throughput, latency, and rate limits.
The application should therefore avoid making every external operation a
synchronous request-path dependency when asynchronous processing is
appropriate.
```text
DjangoPlay Request
│
▼
Application Service
│
▼
Celery
│
▼
External Provider
```
This is particularly useful for:
* Email
* File processing
* Synchronization
* External API calls
* Long-running integration workflows
External provider capacity should be treated as an independent scaling
constraint.
---
# 14. AuthX Scaling
AuthX is a separate identity service and therefore represents an independent
runtime boundary.
```text
DjangoPlay
│
▼
AuthX
│
▼
AuthX PostgreSQL
```
DjangoPlay application scaling does not automatically scale AuthX.
As authentication traffic grows, AuthX can be scaled independently according
to its own runtime architecture.
The two systems also maintain separate PostgreSQL responsibilities.
```text
DjangoPlay
│
▼
DjangoPlay PostgreSQL
AuthX
│
▼
AuthX PostgreSQL
```
This separation prevents DjangoPlay application growth from automatically
forcing the identity database into the same scaling model.
---
# 15. GenericIssueTracker Scaling
GenericIssueTracker is integrated into DjangoPlay as a reusable application
component rather than as an independently deployed microservice.
```text
DjangoPlay
│
▼
IssueTracker Integration
│
▼
GenericIssueTracker
│
▼
DjangoPlay PostgreSQL
```
Its workload therefore initially scales with the DjangoPlay application and
database.
If IssueTracker traffic becomes significant, the first scaling mechanisms are:
* Database query optimization
* Appropriate indexing
* Pagination
* Queryset optimization
* Caching where appropriate
* Background processing for expensive operations
* Independent Celery worker capacity
The reusable package does not need to become a separate service merely because
its workload increases.
A future service extraction remains possible if the operational requirements
justify the additional distributed-system complexity.
---
# 16. Application-Level Scaling
DjangoPlay's modular architecture allows individual workloads to be optimized
without splitting the entire application.
```text
DjangoPlay
│
┌─────────────┼─────────────┐
│ │ │
▼ ▼ ▼
Web / API Background Integration
Workload Workload Workload
│ │ │
Gunicorn Celery External APIs
```
This is an important distinction:
**Scaling a workload does not require extracting that workload into a
microservice.**
The preferred order is:
```text
Optimize
│
▼
Scale Process
│
▼
Separate Workload
│
▼
Scale Infrastructure
│
▼
Extract Service Only If Justified
```
---
# 17. Horizontal Scaling
Horizontal scaling means running multiple instances of a workload.
A future horizontally scaled DjangoPlay deployment can look like:
```mermaid
flowchart TD
CLIENT["Internet Clients"]
LB["Load Balancer"]
APP1["DjangoPlay
Instance 1"]
APP2["DjangoPlay
Instance 2"]
APP3["DjangoPlay
Instance 3"]
REDIS["Shared Redis"]
CELERY1["Celery Worker 1"]
CELERY2["Celery Worker 2"]
DB["PostgreSQL"]
AUTHX["AuthX"]
R2["Cloudflare R2 / CDN"]
CLIENT --> LB
LB --> APP1
LB --> APP2
LB --> APP3
APP1 --> DB
APP2 --> DB
APP3 --> DB
APP1 --> REDIS
APP2 --> REDIS
APP3 --> REDIS
REDIS --> CELERY1
REDIS --> CELERY2
CELERY1 --> DB
CELERY2 --> DB
APP1 --> AUTHX
APP2 --> AUTHX
APP3 --> AUTHX
APP1 --> R2
APP2 --> R2
APP3 --> R2
classDef client fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a
classDef app fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b
classDef infra fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b
classDef external fill:#eeeeff,stroke:#5a55c9,stroke-width:2px,color:#403b91
class CLIENT client
class LB,APP1,APP2,APP3,CELERY1,CELERY2 app
class REDIS,DB infra
class AUTHX,R2 external
```
The critical requirement is that application instances share the required
stateful infrastructure while remaining independently replaceable.
---
# 18. Stateless Application Requirements
Horizontal application scaling works best when application instances do not
depend on local process state.
The following should not become hidden instance-specific state:
* User sessions
* Celery task state
* Persistent application data
* Uploaded assets
* Authentication state that must survive instance replacement
Shared infrastructure should be used where state must survive between
application processes.
```text
Application Instance
│
├── Stateless request processing
│
├── PostgreSQL → persistent data
│
├── Redis → shared runtime state / broker
│
└── R2 / CDN → shared frontend assets
```
This allows an application instance to be removed or replaced without losing
application state.
---
# 19. Scaling Background and Request Workloads Independently
A mature DjangoPlay deployment should be able to scale the two major compute
workloads independently.
```text
Traffic
│
┌─────────┴─────────┐
│ │
▼ ▼
HTTP/API Traffic Background Jobs
│ │
▼ ▼
Django/Gunicorn Celery
│ │
│ │
└─────────┬─────────┘
▼
Shared Data
```
For example:
```text
High API traffic
→ Increase Django/Gunicorn capacity
High email volume
→ Increase Celery capacity
Heavy report generation
→ Dedicated Celery capacity / queue
High read volume
→ Optimize/cache PostgreSQL workload
```
This avoids scaling unrelated components.
---
# 20. Observability Before Scaling
Scaling decisions should be driven by measured bottlenecks.
Before increasing infrastructure capacity, determine whether the limiting
resource is:
* CPU
* Memory
* Database CPU
* Database I/O
* Database connections
* Redis memory
* Redis latency
* Celery queue depth
* External provider latency
* External provider rate limits
* Application response time
A useful diagnostic model is:
```text
Observed Slowdown
│
▼
Measure
│
▼
Identify Bottleneck
│
▼
Optimize
│
▼
Scale Appropriate Layer
```
Blindly increasing server size can mask the real bottleneck without solving
the underlying problem.
---
# 21. Availability and Scaling
Scaling and high availability are related but not identical.
For example:
```text
One Large Server
```
provides more capacity than:
```text
One Small Server
```
but neither removes the single-server failure boundary.
High availability requires redundancy.
A future highly available architecture can therefore use:
```text
Load Balancer
/ | \
/ | \
App #1 App #2 App #3
\ | /
\ | /
Shared Data
/ \
PostgreSQL Redis
HA/Replica HA/Cluster
```
The level of redundancy should match actual availability requirements.
DjangoPlay should not introduce distributed infrastructure solely for
theoretical scalability.
---
# 22. Scaling Roadmap
A practical DjangoPlay scaling path is:
### Stage 1 — Single Node
```text
Nginx
│
▼
Gunicorn / Django
│
├── PostgreSQL
├── Redis
└── Celery
```
Appropriate for development and smaller production workloads.
### Stage 2 — Vertical Scaling
```text
Larger Host
│
├── More Gunicorn capacity
├── More Celery capacity
├── More PostgreSQL resources
└── More Redis resources
```
No major architecture change.
### Stage 3 — Workload Separation
```text
Web Host
└── Nginx + Gunicorn
Worker Host
└── Celery
Data Infrastructure
├── PostgreSQL
└── Redis
```
Web and background workloads can now scale independently.
### Stage 4 — Horizontal Application Scaling
```text
Load Balancer
│
┌────┼────┐
▼ ▼ ▼
App App App
│ │ │
└────┼────┘
│
PostgreSQL
```
### Stage 5 — Advanced Data Scaling
Only when measured workload requires it:
```text
PostgreSQL
│
├── Primary
└── Read Replica(s)
```
Additional database infrastructure should be introduced only after
application-level optimization and connection management have been addressed.
### Stage 6 — Selective Service Extraction
Only if a specific domain has an independent operational requirement:
```text
DjangoPlay Modular Monolith
│
├── Core Application
│
├── GenericIssueTracker
│
└── Candidate Service
│
▼
Independent Runtime
```
Service extraction is therefore the final scaling option, not the starting
architecture.
---
# 23. What Should Not Be Scaled Prematurely
DjangoPlay should avoid introducing unnecessary distributed infrastructure
before measurable demand exists.
Do not introduce by default:
* Kubernetes
* Microservices
* Multiple databases per Django application
* Database sharding
* Distributed caches
* Complex service meshes
* Multiple message brokers
* Independent deployment pipelines for every Django app
These technologies solve specific scale or operational problems, but they also
introduce:
* Network failure modes
* Deployment complexity
* Distributed tracing requirements
* Data consistency challenges
* Operational overhead
* Higher infrastructure cost
DjangoPlay's modular architecture provides room to introduce these mechanisms
later when there is a concrete requirement.
---
# 24. Scaling Decision Matrix
| Observed Bottleneck | First Response | Advanced Response |
| ----------------------------- | ------------------------------ | --------------------------------------------- |
| High HTTP traffic | Increase Gunicorn capacity | Multiple DjangoPlay instances + load balancer |
| High Celery queue depth | Increase worker concurrency | More workers / dedicated queues |
| PostgreSQL CPU | Query/index optimization | Larger DB / read replicas |
| PostgreSQL connections | Reduce unnecessary connections | Connection pooling / architecture changes |
| Redis memory | Review cache usage | Larger/dedicated Redis |
| Static asset traffic | CDN/R2 | Expanded CDN architecture |
| Email workload | Celery | Dedicated email worker queue |
| Heavy background processing | Celery workers | Dedicated worker hosts/queues |
| External API latency | Async processing / caching | Dedicated integration workers |
| Application memory | Optimize process usage | Larger instances / more instances |
| Single-node availability risk | Backups / recovery | Redundant application and data infrastructure |
---
# 25. Scaling Principles
DjangoPlay follows these scaling principles:
1. **Scale the bottleneck, not the entire platform.**
2. **Optimize before adding infrastructure.**
3. **Scale web and background workloads independently.**
4. **Keep application instances as stateless as practical.**
5. **Keep PostgreSQL as the authoritative persistent datastore.**
6. **Use Redis for cache and task-broker responsibilities rather than primary
persistence.**
7. **Use CDN infrastructure for static asset delivery.**
8. **Use asynchronous processing for workloads that do not need to block
requests.**
9. **Treat external provider limits as independent scaling constraints.**
10. **Use database replicas only when measured read workload justifies them.**
11. **Prefer workload separation before microservice extraction.**
12. **Do not introduce distributed infrastructure without a concrete
operational or capacity requirement.**
13. **Preserve modular application boundaries as the system scales.**
14. **Use observability and measurements to drive scaling decisions.**
15. **Treat high availability separately from raw capacity.**
---
# 26. Architectural Principle
DjangoPlay's scaling strategy is based on **progressive separation rather than
premature distribution**.
```text
DjangoPlay
│
▼
Modular Monolith
│
┌──────────┼──────────┐
│ │ │
▼ ▼ ▼
Web/API Celery Integrations
│ │ │
▼ ▼ ▼
Gunicorn Workers External Services
│ │
└────┬─────┘
▼
Shared Data
┌────┴────┐
▼ ▼
PostgreSQL Redis
```
As demand increases:
```text
Optimize
↓
Vertical Scale
↓
Separate Workloads
↓
Horizontal Scale
↓
Scale Data Infrastructure
↓
Extract Only Where Justified
```
This approach allows DjangoPlay to grow from a small deployment into a
distributed production architecture without requiring a complete rewrite of
the application.
The objective is not to maximize the number of infrastructure components.
The objective is to **scale the required workload while preserving DjangoPlay's
modular boundaries and keeping operational complexity proportional to actual
demand**.