djangoplay-web / Architecture / DjangoPlay — Scaling Architecture
DocsDjangoPlay WebArchitectureDjangoPlay — Scaling Architecture

DjangoPlay — Scaling Architecture

Date: 2026-08-29

16 min readApplies to v1.2.2
On this page ▾
  1. 1. Overview
  2. 2. Current Runtime Architecture
  3. 3. Scaling Dimensions
  4. Advantages
  5. Limitation
  6. Worker Scaling Strategies
  7. First-level optimizations
  8. Vertical database scaling
  9. Advanced scaling
  10. Stage 1 — Single Node
  11. Stage 2 — Vertical Scaling
  12. Stage 3 — Workload Separation
  13. Stage 4 — Horizontal Application Scaling
  14. Stage 5 — Advanced Data Scaling
  15. Stage 6 — Selective Service Extraction

1. Overview

DjangoPlay is designed as a modular monolith with independently scalable runtime workloads.

The application is not currently designed around a microservices topology. Instead, scaling can be introduced incrementally by separating the workloads that have different resource and traffic characteristics:

  • Web/API request processing
  • Background task processing
  • Redis cache and task brokering
  • PostgreSQL persistence
  • Static/CDN asset delivery
  • External service integrations

This allows DjangoPlay to scale without prematurely introducing distributed application services.

The primary scaling strategy is therefore:

text
Vertical Scaling
       │
       ▼
Workload Separation
       │
       ▼
Horizontal Scaling
       │
       ▼
Database / Cache Optimization
       │
       ▼
Selective Service Extraction
`

2. Current Runtime Architecture

The current DjangoPlay runtime consists of separate application and supporting processes.

The important architectural property is that web traffic and asynchronous work are already separate workloads.

text
                    DjangoPlay
                        │
              ┌─────────┴─────────┐
              │                   │
              ▼                   ▼
        Web / API             Background
        Processing             Processing
              │                   │
              ▼                   ▼
          Gunicorn             Celery
              │                   │
              └─────────┬─────────┘
                        ▼
                      Redis
                        │
                        ▼
                   PostgreSQL

This separation provides the first scaling boundary without changing the application architecture.


3. Scaling Dimensions

DjangoPlay can scale along several independent dimensions.

Dimension Primary Scaling Method
HTTP/API traffic More Gunicorn workers and application instances
Background jobs More Celery worker processes/instances
Database workload Query optimization, indexes, connection management, larger PostgreSQL resources, replicas where justified
Cache/broker workload Redis resource scaling or dedicated Redis infrastructure
Static assets CDN / Cloudflare R2
External integrations Async processing, provider limits, retries, isolation
Application memory Process sizing and workload separation
CPU-bound workloads Worker/application process scaling
I/O-heavy workloads Async/background execution and workload separation

Not every workload should be scaled in the same way.


4. Vertical Scaling

Vertical scaling is the first and simplest scaling mechanism.

It means increasing resources available to the existing runtime:

  • CPU
  • RAM
  • Disk performance
  • Network capacity
  • PostgreSQL resources
  • Redis resources
  • Gunicorn worker capacity
  • Celery worker capacity

A typical progression is:

text
Small Instance
     │
     ▼
More CPU / RAM
     │
     ▼
More Gunicorn Workers
     │
     ▼
More Celery Workers
     │
     ▼
Dedicated Resources

Vertical scaling is particularly appropriate while DjangoPlay remains a single-node deployment.

Advantages

  • Minimal architectural change
  • Simple operations
  • No distributed application state
  • No load-balancing complexity
  • Low operational overhead

Limitation

A single machine eventually becomes a capacity and availability boundary.

When that point is reached, horizontal scaling becomes the next step.


5. Web / API Scaling

DjangoPlay web and DRF traffic is handled by the Django application behind Nginx and Gunicorn.

text
                    Clients
                       │
                       ▼
                     Nginx
                       │
                       ▼
                  Gunicorn
                       │
              ┌────────┼────────┐
              ▼        ▼        ▼
           Django   Django   Django
           Worker   Worker   Worker
              │        │        │
              └────────┼────────┘
                       ▼
                  PostgreSQL

On a single host, the first scaling step is increasing the number of Gunicorn workers within the available CPU and memory limits.

As traffic grows beyond a single application process group, multiple application instances can be introduced.

text
                    Load Balancer
                         │
             ┌───────────┼───────────┐
             ▼           ▼           ▼
          App #1       App #2      App #3
             │           │           │
             └───────────┼───────────┘
                         ▼
                    PostgreSQL

The application layer should remain as stateless as practical so that requests can be distributed across instances.


6. Celery Worker Scaling

Background processing is an independent scaling dimension.

DjangoPlay already separates asynchronous work through Celery.

text
DjangoPlay
     │
     │ enqueue
     ▼
   Redis
     │
     ▼
Celery Workers

Workers can be increased independently of web/API capacity.

text
                 Redis Broker
                      │
          ┌───────────┼───────────┐
          ▼           ▼           ▼
      Worker #1   Worker #2   Worker #3

This is preferable to increasing web-server capacity merely because background processing has increased.

Worker Scaling Strategies

Celery capacity can be increased by:

  • Increasing worker concurrency
  • Running additional worker processes
  • Running workers on separate hosts
  • Separating workloads into queues
  • Allocating different worker capacity to different task classes

Conceptually:

text
                 Redis
                   │
          ┌────────┼────────┐
          │        │        │
          ▼        ▼        ▼
       Default   Email    Heavy
       Queue     Queue    Queue
          │        │        │
          ▼        ▼        ▼
       Workers  Workers  Workers

Queue separation should be introduced when workload characteristics justify it rather than as an architectural requirement from the beginning.


7. Web and Worker Isolation

One of the most important scaling characteristics of DjangoPlay is the ability to scale request processing independently from background processing.

text
             DjangoPlay Runtime
                    │
          ┌─────────┴─────────┐
          │                   │
          ▼                   ▼
       Web/API             Celery
       Workload            Workload
          │                   │
      Gunicorn             Workers
          │                   │
          ▼                   ▼
      User-facing        Background
       latency            processing

For example:

  • A large email workload should not require additional web workers.
  • A large API traffic increase should not require additional Celery workers.
  • A CPU-heavy background task should not consume the resources required by request processing.

This separation provides an effective scaling boundary while preserving the modular-monolith architecture.


8. Redis Scaling

Redis currently serves multiple runtime responsibilities:

  • Cache
  • Session/runtime support where configured
  • Celery broker
text
                  Redis
                /       \
               /         \
           Cache         Broker
                           │
                           ▼
                        Celery

Redis should therefore be monitored independently from PostgreSQL.

Scaling options include:

  • Increasing Redis memory
  • Increasing host resources
  • Separating cache and broker workloads when justified
  • Moving Redis to dedicated infrastructure
  • Introducing highly available Redis infrastructure when availability requirements justify it

Redis should not become the primary persistence layer for application data.

PostgreSQL remains the authoritative persistent application datastore.


9. PostgreSQL Scaling

PostgreSQL is the primary persistence boundary for DjangoPlay application data.

text
DjangoPlay
     │
     ▼
PostgreSQL

Database scaling should begin with optimization before introducing distributed database infrastructure.

First-level optimizations

  • Query optimization
  • Appropriate indexes
  • Avoiding unnecessary queries
  • Reducing N+1 query patterns
  • Pagination
  • Efficient queryset construction
  • Appropriate transaction boundaries
  • Connection management

Vertical database scaling

Increase:

  • CPU
  • RAM
  • Storage performance
  • I/O capacity

Advanced scaling

When workload justifies it:

text
                PostgreSQL
                    │
          ┌─────────┴─────────┐
          ▼                   ▼
       Primary             Replica(s)
      Read / Write           Read

Read replicas should only be introduced when the application's read/write patterns and consistency requirements make them useful.

They are not automatically beneficial for every DjangoPlay deployment.


10. Database Connection Scaling

Increasing application instances also increases the number of database connections.

For example:

text
1 App Instance
    │
    └── N database connections


10 App Instances
    │
    └── 10 × N database connections

Therefore horizontal application scaling must be accompanied by deliberate database connection management.

Before adding application instances, evaluate:

  • Gunicorn worker count
  • Database connection count
  • PostgreSQL max_connections
  • Connection lifetime
  • Query duration
  • Connection pooling requirements

Database connection exhaustion can become a bottleneck before CPU or memory does.


11. Caching Strategy

Caching should reduce repeated work and database pressure.

The primary DjangoPlay cache infrastructure is Redis.

text
Request
   │
   ▼
Application
   │
   ▼
Cache Lookup
   │
   ├── Hit ───────► Return Cached Result
   │
   └── Miss
        │
        ▼
    PostgreSQL
        │
        ▼
    Store Result
        │
        ▼
      Redis

Caching opportunities should be introduced based on measured application behavior.

Potential targets include:

  • Expensive read operations
  • Stable reference data
  • Frequently accessed computed results
  • Rate-limit state
  • Session/runtime data where applicable

Caching should not be used to hide inefficient database access indefinitely.


12. Static Assets and CDN Scaling

Frontend assets should not consume application-server capacity when they can be delivered through the CDN path.

DjangoPlay supports Cloudflare R2 for the frontend asset workflow.

text
Frontend Build
      │
      ▼
Cloudflare R2
      │
      ▼
CDN
      │
      ▼
Browser

This moves static asset delivery away from Django/Gunicorn.

The local development configuration can serve frontend assets directly, while the CDN workflow can publish versioned assets to R2.

The CDN therefore provides a natural scaling boundary:

text
Application Server
    │
    └── Dynamic Requests

CDN / R2
    │
    └── Static Frontend Assets

13. External Integration Scaling

DjangoPlay communicates with external services including:

  • AuthX
  • SMTP/email providers
  • Cloudflare R2
  • AI providers
  • Other external APIs

These services have their own throughput, latency, and rate limits.

The application should therefore avoid making every external operation a synchronous request-path dependency when asynchronous processing is appropriate.

text
DjangoPlay Request
       │
       ▼
Application Service
       │
       ▼
Celery
       │
       ▼
External Provider

This is particularly useful for:

  • Email
  • File processing
  • Synchronization
  • External API calls
  • Long-running integration workflows

External provider capacity should be treated as an independent scaling constraint.


14. AuthX Scaling

AuthX is a separate identity service and therefore represents an independent runtime boundary.

text
DjangoPlay
     │
     ▼
AuthX
     │
     ▼
AuthX PostgreSQL

DjangoPlay application scaling does not automatically scale AuthX.

As authentication traffic grows, AuthX can be scaled independently according to its own runtime architecture.

The two systems also maintain separate PostgreSQL responsibilities.

text
DjangoPlay
     │
     ▼
DjangoPlay PostgreSQL


AuthX
     │
     ▼
AuthX PostgreSQL

This separation prevents DjangoPlay application growth from automatically forcing the identity database into the same scaling model.


15. GenericIssueTracker Scaling

GenericIssueTracker is integrated into DjangoPlay as a reusable application component rather than as an independently deployed microservice.

text
DjangoPlay
     │
     ▼
IssueTracker Integration
     │
     ▼
GenericIssueTracker
     │
     ▼
DjangoPlay PostgreSQL

Its workload therefore initially scales with the DjangoPlay application and database.

If IssueTracker traffic becomes significant, the first scaling mechanisms are:

  • Database query optimization
  • Appropriate indexing
  • Pagination
  • Queryset optimization
  • Caching where appropriate
  • Background processing for expensive operations
  • Independent Celery worker capacity

The reusable package does not need to become a separate service merely because its workload increases.

A future service extraction remains possible if the operational requirements justify the additional distributed-system complexity.


16. Application-Level Scaling

DjangoPlay's modular architecture allows individual workloads to be optimized without splitting the entire application.

text
                 DjangoPlay
                     │
       ┌─────────────┼─────────────┐
       │             │             │
       ▼             ▼             ▼
   Web / API      Background     Integration
    Workload       Workload       Workload
       │             │             │
    Gunicorn        Celery       External APIs

This is an important distinction:

Scaling a workload does not require extracting that workload into a microservice.

The preferred order is:

text
Optimize
   │
   ▼
Scale Process
   │
   ▼
Separate Workload
   │
   ▼
Scale Infrastructure
   │
   ▼
Extract Service Only If Justified

17. Horizontal Scaling

Horizontal scaling means running multiple instances of a workload.

A future horizontally scaled DjangoPlay deployment can look like:

The critical requirement is that application instances share the required stateful infrastructure while remaining independently replaceable.


18. Stateless Application Requirements

Horizontal application scaling works best when application instances do not depend on local process state.

The following should not become hidden instance-specific state:

  • User sessions
  • Celery task state
  • Persistent application data
  • Uploaded assets
  • Authentication state that must survive instance replacement

Shared infrastructure should be used where state must survive between application processes.

text
Application Instance
       │
       ├── Stateless request processing
       │
       ├── PostgreSQL → persistent data
       │
       ├── Redis      → shared runtime state / broker
       │
       └── R2 / CDN   → shared frontend assets

This allows an application instance to be removed or replaced without losing application state.


19. Scaling Background and Request Workloads Independently

A mature DjangoPlay deployment should be able to scale the two major compute workloads independently.

text
                    Traffic
                       │
             ┌─────────┴─────────┐
             │                   │
             ▼                   ▼
       HTTP/API Traffic     Background Jobs
             │                   │
             ▼                   ▼
       Django/Gunicorn          Celery
             │                   │
             │                   │
             └─────────┬─────────┘
                       ▼
                   Shared Data

For example:

text
High API traffic
    → Increase Django/Gunicorn capacity

High email volume
    → Increase Celery capacity

Heavy report generation
    → Dedicated Celery capacity / queue

High read volume
    → Optimize/cache PostgreSQL workload

This avoids scaling unrelated components.


20. Observability Before Scaling

Scaling decisions should be driven by measured bottlenecks.

Before increasing infrastructure capacity, determine whether the limiting resource is:

  • CPU
  • Memory
  • Database CPU
  • Database I/O
  • Database connections
  • Redis memory
  • Redis latency
  • Celery queue depth
  • External provider latency
  • External provider rate limits
  • Application response time

A useful diagnostic model is:

text
Observed Slowdown
       │
       ▼
Measure
       │
       ▼
Identify Bottleneck
       │
       ▼
Optimize
       │
       ▼
Scale Appropriate Layer

Blindly increasing server size can mask the real bottleneck without solving the underlying problem.


21. Availability and Scaling

Scaling and high availability are related but not identical.

For example:

text
One Large Server

provides more capacity than:

text
One Small Server

but neither removes the single-server failure boundary.

High availability requires redundancy.

A future highly available architecture can therefore use:

text
                 Load Balancer
                 /     |     \
                /      |      \
             App #1  App #2  App #3
                \      |      /
                 \     |     /
                  Shared Data
                 /           \
            PostgreSQL       Redis
             HA/Replica      HA/Cluster

The level of redundancy should match actual availability requirements.

DjangoPlay should not introduce distributed infrastructure solely for theoretical scalability.


22. Scaling Roadmap

A practical DjangoPlay scaling path is:

Stage 1 — Single Node

text
Nginx
  │
  ▼
Gunicorn / Django
  │
  ├── PostgreSQL
  ├── Redis
  └── Celery

Appropriate for development and smaller production workloads.

Stage 2 — Vertical Scaling

text
Larger Host
    │
    ├── More Gunicorn capacity
    ├── More Celery capacity
    ├── More PostgreSQL resources
    └── More Redis resources

No major architecture change.

Stage 3 — Workload Separation

text
Web Host
    └── Nginx + Gunicorn

Worker Host
    └── Celery

Data Infrastructure
    ├── PostgreSQL
    └── Redis

Web and background workloads can now scale independently.

Stage 4 — Horizontal Application Scaling

text
Load Balancer
      │
 ┌────┼────┐
 ▼    ▼    ▼
App  App  App
 │    │    │
 └────┼────┘
      │
 PostgreSQL

Stage 5 — Advanced Data Scaling

Only when measured workload requires it:

text
PostgreSQL
    │
    ├── Primary
    └── Read Replica(s)

Additional database infrastructure should be introduced only after application-level optimization and connection management have been addressed.

Stage 6 — Selective Service Extraction

Only if a specific domain has an independent operational requirement:

text
DjangoPlay Modular Monolith
          │
          ├── Core Application
          │
          ├── GenericIssueTracker
          │
          └── Candidate Service
                    │
                    ▼
              Independent Runtime

Service extraction is therefore the final scaling option, not the starting architecture.


23. What Should Not Be Scaled Prematurely

DjangoPlay should avoid introducing unnecessary distributed infrastructure before measurable demand exists.

Do not introduce by default:

  • Kubernetes
  • Microservices
  • Multiple databases per Django application
  • Database sharding
  • Distributed caches
  • Complex service meshes
  • Multiple message brokers
  • Independent deployment pipelines for every Django app

These technologies solve specific scale or operational problems, but they also introduce:

  • Network failure modes
  • Deployment complexity
  • Distributed tracing requirements
  • Data consistency challenges
  • Operational overhead
  • Higher infrastructure cost

DjangoPlay's modular architecture provides room to introduce these mechanisms later when there is a concrete requirement.


24. Scaling Decision Matrix

Observed Bottleneck First Response Advanced Response
High HTTP traffic Increase Gunicorn capacity Multiple DjangoPlay instances + load balancer
High Celery queue depth Increase worker concurrency More workers / dedicated queues
PostgreSQL CPU Query/index optimization Larger DB / read replicas
PostgreSQL connections Reduce unnecessary connections Connection pooling / architecture changes
Redis memory Review cache usage Larger/dedicated Redis
Static asset traffic CDN/R2 Expanded CDN architecture
Email workload Celery Dedicated email worker queue
Heavy background processing Celery workers Dedicated worker hosts/queues
External API latency Async processing / caching Dedicated integration workers
Application memory Optimize process usage Larger instances / more instances
Single-node availability risk Backups / recovery Redundant application and data infrastructure

25. Scaling Principles

DjangoPlay follows these scaling principles:

  1. Scale the bottleneck, not the entire platform.
  2. Optimize before adding infrastructure.
  3. Scale web and background workloads independently.
  4. Keep application instances as stateless as practical.
  5. Keep PostgreSQL as the authoritative persistent datastore.
  6. Use Redis for cache and task-broker responsibilities rather than primary persistence.
  7. Use CDN infrastructure for static asset delivery.
  8. Use asynchronous processing for workloads that do not need to block requests.
  9. Treat external provider limits as independent scaling constraints.
  10. Use database replicas only when measured read workload justifies them.
  11. Prefer workload separation before microservice extraction.
  12. Do not introduce distributed infrastructure without a concrete operational or capacity requirement.
  13. Preserve modular application boundaries as the system scales.
  14. Use observability and measurements to drive scaling decisions.
  15. Treat high availability separately from raw capacity.

26. Architectural Principle

DjangoPlay's scaling strategy is based on progressive separation rather than premature distribution.

text
                 DjangoPlay
                     │
                     ▼
                Modular Monolith
                     │
          ┌──────────┼──────────┐
          │          │          │
          ▼          ▼          ▼
        Web/API    Celery    Integrations
          │          │          │
          ▼          ▼          ▼
       Gunicorn   Workers   External Services
          │          │
          └────┬─────┘
               ▼
          Shared Data
          ┌────┴────┐
          ▼         ▼
      PostgreSQL   Redis

As demand increases:

text
Optimize
   ↓
Vertical Scale
   ↓
Separate Workloads
   ↓
Horizontal Scale
   ↓
Scale Data Infrastructure
   ↓
Extract Only Where Justified

This approach allows DjangoPlay to grow from a small deployment into a distributed production architecture without requiring a complete rewrite of the application.

The objective is not to maximize the number of infrastructure components.

The objective is to scale the required workload while preserving DjangoPlay's modular boundaries and keeping operational complexity proportional to actual demand.