--- since: 1.2.1 --- # DjangoPlay — Scaling Architecture **Date:** 2026-08-29 --- ## 1. Overview DjangoPlay is designed as a **modular monolith with independently scalable runtime workloads**. The application is not currently designed around a microservices topology. Instead, scaling can be introduced incrementally by separating the workloads that have different resource and traffic characteristics: * Web/API request processing * Background task processing * Redis cache and task brokering * PostgreSQL persistence * Static/CDN asset delivery * External service integrations This allows DjangoPlay to scale without prematurely introducing distributed application services. The primary scaling strategy is therefore: ```text Vertical Scaling │ ▼ Workload Separation │ ▼ Horizontal Scaling │ ▼ Database / Cache Optimization │ ▼ Selective Service Extraction ```` --- ## 2. Current Runtime Architecture The current DjangoPlay runtime consists of separate application and supporting processes. ```mermaid flowchart TD CLIENT["Clients
Browser / API"] NGINX["Nginx
Reverse Proxy"] DJANGO["DjangoPlay
Web + DRF API"] POSTGRES["PostgreSQL
Application Database"] REDIS["Redis
Cache + Celery Broker"] CELERY["Celery Worker
Background Tasks"] AUTHX["AuthX
Identity Service"] EXTERNAL["External Services
Email / R2 / AI / APIs"] CLIENT --> NGINX NGINX --> DJANGO DJANGO --> POSTGRES DJANGO --> REDIS REDIS --> CELERY CELERY --> POSTGRES CELERY --> AUTHX CELERY --> EXTERNAL DJANGO --> AUTHX DJANGO --> EXTERNAL classDef client fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a classDef application fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b classDef infrastructure fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b classDef external fill:#eeeeff,stroke:#5a55c9,stroke-width:2px,color:#403b91 class CLIENT client class NGINX,DJANGO application class POSTGRES,REDIS,CELERY infrastructure class AUTHX,EXTERNAL external ``` The important architectural property is that **web traffic and asynchronous work are already separate workloads**. ```text DjangoPlay │ ┌─────────┴─────────┐ │ │ ▼ ▼ Web / API Background Processing Processing │ │ ▼ ▼ Gunicorn Celery │ │ └─────────┬─────────┘ ▼ Redis │ ▼ PostgreSQL ``` This separation provides the first scaling boundary without changing the application architecture. --- ## 3. Scaling Dimensions DjangoPlay can scale along several independent dimensions. | Dimension | Primary Scaling Method | | --------------------- | --------------------------------------------------------------------------------------------------------- | | HTTP/API traffic | More Gunicorn workers and application instances | | Background jobs | More Celery worker processes/instances | | Database workload | Query optimization, indexes, connection management, larger PostgreSQL resources, replicas where justified | | Cache/broker workload | Redis resource scaling or dedicated Redis infrastructure | | Static assets | CDN / Cloudflare R2 | | External integrations | Async processing, provider limits, retries, isolation | | Application memory | Process sizing and workload separation | | CPU-bound workloads | Worker/application process scaling | | I/O-heavy workloads | Async/background execution and workload separation | Not every workload should be scaled in the same way. --- # 4. Vertical Scaling Vertical scaling is the first and simplest scaling mechanism. It means increasing resources available to the existing runtime: * CPU * RAM * Disk performance * Network capacity * PostgreSQL resources * Redis resources * Gunicorn worker capacity * Celery worker capacity A typical progression is: ```text Small Instance │ ▼ More CPU / RAM │ ▼ More Gunicorn Workers │ ▼ More Celery Workers │ ▼ Dedicated Resources ``` Vertical scaling is particularly appropriate while DjangoPlay remains a single-node deployment. ### Advantages * Minimal architectural change * Simple operations * No distributed application state * No load-balancing complexity * Low operational overhead ### Limitation A single machine eventually becomes a capacity and availability boundary. When that point is reached, horizontal scaling becomes the next step. --- # 5. Web / API Scaling DjangoPlay web and DRF traffic is handled by the Django application behind Nginx and Gunicorn. ```text Clients │ ▼ Nginx │ ▼ Gunicorn │ ┌────────┼────────┐ ▼ ▼ ▼ Django Django Django Worker Worker Worker │ │ │ └────────┼────────┘ ▼ PostgreSQL ``` On a single host, the first scaling step is increasing the number of Gunicorn workers within the available CPU and memory limits. As traffic grows beyond a single application process group, multiple application instances can be introduced. ```text Load Balancer │ ┌───────────┼───────────┐ ▼ ▼ ▼ App #1 App #2 App #3 │ │ │ └───────────┼───────────┘ ▼ PostgreSQL ``` The application layer should remain as stateless as practical so that requests can be distributed across instances. --- # 6. Celery Worker Scaling Background processing is an independent scaling dimension. DjangoPlay already separates asynchronous work through Celery. ```text DjangoPlay │ │ enqueue ▼ Redis │ ▼ Celery Workers ``` Workers can be increased independently of web/API capacity. ```text Redis Broker │ ┌───────────┼───────────┐ ▼ ▼ ▼ Worker #1 Worker #2 Worker #3 ``` This is preferable to increasing web-server capacity merely because background processing has increased. ### Worker Scaling Strategies Celery capacity can be increased by: * Increasing worker concurrency * Running additional worker processes * Running workers on separate hosts * Separating workloads into queues * Allocating different worker capacity to different task classes Conceptually: ```text Redis │ ┌────────┼────────┐ │ │ │ ▼ ▼ ▼ Default Email Heavy Queue Queue Queue │ │ │ ▼ ▼ ▼ Workers Workers Workers ``` Queue separation should be introduced when workload characteristics justify it rather than as an architectural requirement from the beginning. --- # 7. Web and Worker Isolation One of the most important scaling characteristics of DjangoPlay is the ability to scale request processing independently from background processing. ```text DjangoPlay Runtime │ ┌─────────┴─────────┐ │ │ ▼ ▼ Web/API Celery Workload Workload │ │ Gunicorn Workers │ │ ▼ ▼ User-facing Background latency processing ``` For example: * A large email workload should not require additional web workers. * A large API traffic increase should not require additional Celery workers. * A CPU-heavy background task should not consume the resources required by request processing. This separation provides an effective scaling boundary while preserving the modular-monolith architecture. --- # 8. Redis Scaling Redis currently serves multiple runtime responsibilities: * Cache * Session/runtime support where configured * Celery broker ```text Redis / \ / \ Cache Broker │ ▼ Celery ``` Redis should therefore be monitored independently from PostgreSQL. Scaling options include: * Increasing Redis memory * Increasing host resources * Separating cache and broker workloads when justified * Moving Redis to dedicated infrastructure * Introducing highly available Redis infrastructure when availability requirements justify it Redis should not become the primary persistence layer for application data. PostgreSQL remains the authoritative persistent application datastore. --- # 9. PostgreSQL Scaling PostgreSQL is the primary persistence boundary for DjangoPlay application data. ```text DjangoPlay │ ▼ PostgreSQL ``` Database scaling should begin with optimization before introducing distributed database infrastructure. ### First-level optimizations * Query optimization * Appropriate indexes * Avoiding unnecessary queries * Reducing N+1 query patterns * Pagination * Efficient queryset construction * Appropriate transaction boundaries * Connection management ### Vertical database scaling Increase: * CPU * RAM * Storage performance * I/O capacity ### Advanced scaling When workload justifies it: ```text PostgreSQL │ ┌─────────┴─────────┐ ▼ ▼ Primary Replica(s) Read / Write Read ``` Read replicas should only be introduced when the application's read/write patterns and consistency requirements make them useful. They are not automatically beneficial for every DjangoPlay deployment. --- # 10. Database Connection Scaling Increasing application instances also increases the number of database connections. For example: ```text 1 App Instance │ └── N database connections 10 App Instances │ └── 10 × N database connections ``` Therefore horizontal application scaling must be accompanied by deliberate database connection management. Before adding application instances, evaluate: * Gunicorn worker count * Database connection count * PostgreSQL `max_connections` * Connection lifetime * Query duration * Connection pooling requirements Database connection exhaustion can become a bottleneck before CPU or memory does. --- # 11. Caching Strategy Caching should reduce repeated work and database pressure. The primary DjangoPlay cache infrastructure is Redis. ```text Request │ ▼ Application │ ▼ Cache Lookup │ ├── Hit ───────► Return Cached Result │ └── Miss │ ▼ PostgreSQL │ ▼ Store Result │ ▼ Redis ``` Caching opportunities should be introduced based on measured application behavior. Potential targets include: * Expensive read operations * Stable reference data * Frequently accessed computed results * Rate-limit state * Session/runtime data where applicable Caching should not be used to hide inefficient database access indefinitely. --- # 12. Static Assets and CDN Scaling Frontend assets should not consume application-server capacity when they can be delivered through the CDN path. DjangoPlay supports Cloudflare R2 for the frontend asset workflow. ```text Frontend Build │ ▼ Cloudflare R2 │ ▼ CDN │ ▼ Browser ``` This moves static asset delivery away from Django/Gunicorn. The local development configuration can serve frontend assets directly, while the CDN workflow can publish versioned assets to R2. The CDN therefore provides a natural scaling boundary: ```text Application Server │ └── Dynamic Requests CDN / R2 │ └── Static Frontend Assets ``` --- # 13. External Integration Scaling DjangoPlay communicates with external services including: * AuthX * SMTP/email providers * Cloudflare R2 * AI providers * Other external APIs These services have their own throughput, latency, and rate limits. The application should therefore avoid making every external operation a synchronous request-path dependency when asynchronous processing is appropriate. ```text DjangoPlay Request │ ▼ Application Service │ ▼ Celery │ ▼ External Provider ``` This is particularly useful for: * Email * File processing * Synchronization * External API calls * Long-running integration workflows External provider capacity should be treated as an independent scaling constraint. --- # 14. AuthX Scaling AuthX is a separate identity service and therefore represents an independent runtime boundary. ```text DjangoPlay │ ▼ AuthX │ ▼ AuthX PostgreSQL ``` DjangoPlay application scaling does not automatically scale AuthX. As authentication traffic grows, AuthX can be scaled independently according to its own runtime architecture. The two systems also maintain separate PostgreSQL responsibilities. ```text DjangoPlay │ ▼ DjangoPlay PostgreSQL AuthX │ ▼ AuthX PostgreSQL ``` This separation prevents DjangoPlay application growth from automatically forcing the identity database into the same scaling model. --- # 15. GenericIssueTracker Scaling GenericIssueTracker is integrated into DjangoPlay as a reusable application component rather than as an independently deployed microservice. ```text DjangoPlay │ ▼ IssueTracker Integration │ ▼ GenericIssueTracker │ ▼ DjangoPlay PostgreSQL ``` Its workload therefore initially scales with the DjangoPlay application and database. If IssueTracker traffic becomes significant, the first scaling mechanisms are: * Database query optimization * Appropriate indexing * Pagination * Queryset optimization * Caching where appropriate * Background processing for expensive operations * Independent Celery worker capacity The reusable package does not need to become a separate service merely because its workload increases. A future service extraction remains possible if the operational requirements justify the additional distributed-system complexity. --- # 16. Application-Level Scaling DjangoPlay's modular architecture allows individual workloads to be optimized without splitting the entire application. ```text DjangoPlay │ ┌─────────────┼─────────────┐ │ │ │ ▼ ▼ ▼ Web / API Background Integration Workload Workload Workload │ │ │ Gunicorn Celery External APIs ``` This is an important distinction: **Scaling a workload does not require extracting that workload into a microservice.** The preferred order is: ```text Optimize │ ▼ Scale Process │ ▼ Separate Workload │ ▼ Scale Infrastructure │ ▼ Extract Service Only If Justified ``` --- # 17. Horizontal Scaling Horizontal scaling means running multiple instances of a workload. A future horizontally scaled DjangoPlay deployment can look like: ```mermaid flowchart TD CLIENT["Internet Clients"] LB["Load Balancer"] APP1["DjangoPlay
Instance 1"] APP2["DjangoPlay
Instance 2"] APP3["DjangoPlay
Instance 3"] REDIS["Shared Redis"] CELERY1["Celery Worker 1"] CELERY2["Celery Worker 2"] DB["PostgreSQL"] AUTHX["AuthX"] R2["Cloudflare R2 / CDN"] CLIENT --> LB LB --> APP1 LB --> APP2 LB --> APP3 APP1 --> DB APP2 --> DB APP3 --> DB APP1 --> REDIS APP2 --> REDIS APP3 --> REDIS REDIS --> CELERY1 REDIS --> CELERY2 CELERY1 --> DB CELERY2 --> DB APP1 --> AUTHX APP2 --> AUTHX APP3 --> AUTHX APP1 --> R2 APP2 --> R2 APP3 --> R2 classDef client fill:#e6f1fb,stroke:#2474b5,stroke-width:2px,color:#174d7a classDef app fill:#e5f6f1,stroke:#159a7a,stroke-width:2px,color:#075d4b classDef infra fill:#fff2d9,stroke:#d89b18,stroke-width:2px,color:#76520b classDef external fill:#eeeeff,stroke:#5a55c9,stroke-width:2px,color:#403b91 class CLIENT client class LB,APP1,APP2,APP3,CELERY1,CELERY2 app class REDIS,DB infra class AUTHX,R2 external ``` The critical requirement is that application instances share the required stateful infrastructure while remaining independently replaceable. --- # 18. Stateless Application Requirements Horizontal application scaling works best when application instances do not depend on local process state. The following should not become hidden instance-specific state: * User sessions * Celery task state * Persistent application data * Uploaded assets * Authentication state that must survive instance replacement Shared infrastructure should be used where state must survive between application processes. ```text Application Instance │ ├── Stateless request processing │ ├── PostgreSQL → persistent data │ ├── Redis → shared runtime state / broker │ └── R2 / CDN → shared frontend assets ``` This allows an application instance to be removed or replaced without losing application state. --- # 19. Scaling Background and Request Workloads Independently A mature DjangoPlay deployment should be able to scale the two major compute workloads independently. ```text Traffic │ ┌─────────┴─────────┐ │ │ ▼ ▼ HTTP/API Traffic Background Jobs │ │ ▼ ▼ Django/Gunicorn Celery │ │ │ │ └─────────┬─────────┘ ▼ Shared Data ``` For example: ```text High API traffic → Increase Django/Gunicorn capacity High email volume → Increase Celery capacity Heavy report generation → Dedicated Celery capacity / queue High read volume → Optimize/cache PostgreSQL workload ``` This avoids scaling unrelated components. --- # 20. Observability Before Scaling Scaling decisions should be driven by measured bottlenecks. Before increasing infrastructure capacity, determine whether the limiting resource is: * CPU * Memory * Database CPU * Database I/O * Database connections * Redis memory * Redis latency * Celery queue depth * External provider latency * External provider rate limits * Application response time A useful diagnostic model is: ```text Observed Slowdown │ ▼ Measure │ ▼ Identify Bottleneck │ ▼ Optimize │ ▼ Scale Appropriate Layer ``` Blindly increasing server size can mask the real bottleneck without solving the underlying problem. --- # 21. Availability and Scaling Scaling and high availability are related but not identical. For example: ```text One Large Server ``` provides more capacity than: ```text One Small Server ``` but neither removes the single-server failure boundary. High availability requires redundancy. A future highly available architecture can therefore use: ```text Load Balancer / | \ / | \ App #1 App #2 App #3 \ | / \ | / Shared Data / \ PostgreSQL Redis HA/Replica HA/Cluster ``` The level of redundancy should match actual availability requirements. DjangoPlay should not introduce distributed infrastructure solely for theoretical scalability. --- # 22. Scaling Roadmap A practical DjangoPlay scaling path is: ### Stage 1 — Single Node ```text Nginx │ ▼ Gunicorn / Django │ ├── PostgreSQL ├── Redis └── Celery ``` Appropriate for development and smaller production workloads. ### Stage 2 — Vertical Scaling ```text Larger Host │ ├── More Gunicorn capacity ├── More Celery capacity ├── More PostgreSQL resources └── More Redis resources ``` No major architecture change. ### Stage 3 — Workload Separation ```text Web Host └── Nginx + Gunicorn Worker Host └── Celery Data Infrastructure ├── PostgreSQL └── Redis ``` Web and background workloads can now scale independently. ### Stage 4 — Horizontal Application Scaling ```text Load Balancer │ ┌────┼────┐ ▼ ▼ ▼ App App App │ │ │ └────┼────┘ │ PostgreSQL ``` ### Stage 5 — Advanced Data Scaling Only when measured workload requires it: ```text PostgreSQL │ ├── Primary └── Read Replica(s) ``` Additional database infrastructure should be introduced only after application-level optimization and connection management have been addressed. ### Stage 6 — Selective Service Extraction Only if a specific domain has an independent operational requirement: ```text DjangoPlay Modular Monolith │ ├── Core Application │ ├── GenericIssueTracker │ └── Candidate Service │ ▼ Independent Runtime ``` Service extraction is therefore the final scaling option, not the starting architecture. --- # 23. What Should Not Be Scaled Prematurely DjangoPlay should avoid introducing unnecessary distributed infrastructure before measurable demand exists. Do not introduce by default: * Kubernetes * Microservices * Multiple databases per Django application * Database sharding * Distributed caches * Complex service meshes * Multiple message brokers * Independent deployment pipelines for every Django app These technologies solve specific scale or operational problems, but they also introduce: * Network failure modes * Deployment complexity * Distributed tracing requirements * Data consistency challenges * Operational overhead * Higher infrastructure cost DjangoPlay's modular architecture provides room to introduce these mechanisms later when there is a concrete requirement. --- # 24. Scaling Decision Matrix | Observed Bottleneck | First Response | Advanced Response | | ----------------------------- | ------------------------------ | --------------------------------------------- | | High HTTP traffic | Increase Gunicorn capacity | Multiple DjangoPlay instances + load balancer | | High Celery queue depth | Increase worker concurrency | More workers / dedicated queues | | PostgreSQL CPU | Query/index optimization | Larger DB / read replicas | | PostgreSQL connections | Reduce unnecessary connections | Connection pooling / architecture changes | | Redis memory | Review cache usage | Larger/dedicated Redis | | Static asset traffic | CDN/R2 | Expanded CDN architecture | | Email workload | Celery | Dedicated email worker queue | | Heavy background processing | Celery workers | Dedicated worker hosts/queues | | External API latency | Async processing / caching | Dedicated integration workers | | Application memory | Optimize process usage | Larger instances / more instances | | Single-node availability risk | Backups / recovery | Redundant application and data infrastructure | --- # 25. Scaling Principles DjangoPlay follows these scaling principles: 1. **Scale the bottleneck, not the entire platform.** 2. **Optimize before adding infrastructure.** 3. **Scale web and background workloads independently.** 4. **Keep application instances as stateless as practical.** 5. **Keep PostgreSQL as the authoritative persistent datastore.** 6. **Use Redis for cache and task-broker responsibilities rather than primary persistence.** 7. **Use CDN infrastructure for static asset delivery.** 8. **Use asynchronous processing for workloads that do not need to block requests.** 9. **Treat external provider limits as independent scaling constraints.** 10. **Use database replicas only when measured read workload justifies them.** 11. **Prefer workload separation before microservice extraction.** 12. **Do not introduce distributed infrastructure without a concrete operational or capacity requirement.** 13. **Preserve modular application boundaries as the system scales.** 14. **Use observability and measurements to drive scaling decisions.** 15. **Treat high availability separately from raw capacity.** --- # 26. Architectural Principle DjangoPlay's scaling strategy is based on **progressive separation rather than premature distribution**. ```text DjangoPlay │ ▼ Modular Monolith │ ┌──────────┼──────────┐ │ │ │ ▼ ▼ ▼ Web/API Celery Integrations │ │ │ ▼ ▼ ▼ Gunicorn Workers External Services │ │ └────┬─────┘ ▼ Shared Data ┌────┴────┐ ▼ ▼ PostgreSQL Redis ``` As demand increases: ```text Optimize ↓ Vertical Scale ↓ Separate Workloads ↓ Horizontal Scale ↓ Scale Data Infrastructure ↓ Extract Only Where Justified ``` This approach allows DjangoPlay to grow from a small deployment into a distributed production architecture without requiring a complete rewrite of the application. The objective is not to maximize the number of infrastructure components. The objective is to **scale the required workload while preserving DjangoPlay's modular boundaries and keeping operational complexity proportional to actual demand**.