djangoplay-web / User guide / DjangoPlay — Development Data & DevTools
DocsDjangoPlay WebUser guideDjangoPlay — Development Data & DevTools

DjangoPlay — Development Data & DevTools

DjangoPlay provides a dedicated development and data-management toolkit under:

21 min readApplies to v1.2.2
On this page ▾
  1. 1. Overview
  2. 5.1 fetch_reference_data
  3. Behavior
  4. 6.1 generate_geonames_json
  5. Important
  6. 7.1 Global Regions
  7. 7.2 Country Information
  8. 7.3 Timezones
  9. 8.1 Cities
  10. 8.2 Postal Codes
  11. 10.1 import_classifications
  12. 11.1 sync_system_constants
  13. 12.1 Employees
  14. 12.2 Members
  15. 12.3 Development Superuser
  16. 13.1 sync_businesses
  17. Important
  18. 14.1 Finance Data Generator
  19. 14.2 Finance Reference Data
  20. 14.3 Finance Identities
  21. 14.4 Billing Schedules
  22. 14.5 Invoices
  23. 19.1 Invalid Regions
  24. 19.2 Locations
  25. 20.1 masterscript
  26. 20.2 MasterScript Options
  27. Phase 0 — Superuser Bootstrap
  28. Phase 0.5 — Reference Data Fetch
  29. Phase 1 — Global Initialization
  30. 26.1 Command timeout
  31. 26.2 Available-memory check
  32. Business Services
  33. Finance Services
  34. Location Services
  35. User Services
  36. Common Services
  37. Reference data
  38. Synthetic data

1. Overview

text
webapp/devtools/
`

The package contains Django management commands and supporting services for:

  • Reference-data acquisition
  • Reference-data transformation and JSON generation
  • Reference-data database imports
  • Synthetic business/entity generation
  • Finance test-data generation
  • Development employee and member generation
  • Development-data orchestration
  • Data cleanup
  • Redis cache refresh
  • Permission setup
  • Database and application maintenance utilities

The tooling is intended primarily for local development, testing, demonstrations, data preparation, and controlled maintenance.

It is not a production data-management framework.


2. DevTools Architecture

The package follows a command → service architecture.

Management commands are intentionally thin where practical, while the more substantial generation and import logic lives under devtools/services/.

text
webapp/devtools/
│
├── management/
│   ├── commands/
│   │   └── Django management commands
│   │
│   └── auxiliary_commands/
│       └── developer / maintenance utilities
│
├── services/
│   ├── business/
│   │   └── synthetic business generation/import
│   │
│   ├── finance/
│   │   └── finance test-data generation
│   │
│   ├── locations/
│   │   └── reference-data generation/import
│   │
│   ├── users/
│   │   └── employee/member generation
│   │
│   ├── validators/
│   │   └── development-data validation
│   │
│   └── common/
│       └── shared generation/import utilities
│
└── tasks.py
    └── Redis/cache-oriented background tasks and serializers

The general execution pattern is:

text
Django manage.py command
        │
        ▼
DevTools Command
        │
        ▼
Generation / Import / Sync Service
        │
        ├── Reference Data
        ├── Domain Models
        ├── External Data Files
        └── Redis / Celery

3. Data Categories

DjangoPlay development data falls into several distinct categories.

Category Examples Origin
Reference data Countries, regions, subregions, cities, postal codes, timezones External/prepared datasets
Classification data ISIC, CPC, HS classifications UN classification source files
Synthetic domain data Businesses/entities Generated locally
Synthetic finance data Finance identities, billing schedules, invoices Generated locally
Development users Employees, members Generated locally
System master data Roles, departments, statuses, teams, etc. Application constants
Cache data Location, industry and related cached representations Generated from database state

These categories should not be treated as interchangeable.

In particular, reference data is not synthetic application data.


4. Reference-Data Pipeline

Reference-data processing consists of multiple stages.

text
External / Prepared Source
          │
          ▼
generate_geonames_json
          │
          ▼
Compiled JSON
          │
          ├───────────────┐
          │               │
          ▼               ▼
Local DATA_DIR       R2 / Published Dataset
                          │
                          ▼
                 fetch_reference_data
                          │
                          ▼
                     Local DATA_DIR
                          │
                          ▼
                 Import Management Commands
                          │
                          ▼
                       Database

There are therefore two different operations:

  1. Generate/compile source data into JSON
  2. Import prepared JSON into Django models

They should not be confused.


5. Fetching Published Reference Data

5.1 fetch_reference_data

Downloads published reference-data files from the configured R2-compatible data source into DATA_DIR.

bash
python manage.py fetch_reference_data --country NZ

Multiple countries can be requested:

bash
python manage.py fetch_reference_data --country FR,US,JP,IN

The command also supports source selection:

bash
python manage.py fetch_reference_data \
    --country NZ \
    --source cities15000

Available sources can be inspected without downloading data:

bash
python manage.py fetch_reference_data --list-sources

Use --force to download files again when they already exist locally:

bash
python manage.py fetch_reference_data \
    --country NZ \
    --force

Behavior

The command:

  1. Loads ~/.dplay/.appdata
  2. Reads DATA_DIR
  3. Reads DATA_SOURCE_URL
  4. Retrieves the remote manifest.json
  5. Resolves the selected compiled-data source
  6. Downloads global datasets
  7. Downloads country-specific datasets
  8. Skips files already present unless --force is used

The default compiled-data source is:

text
cities15000

6. Generating Reference-Data JSON

6.1 generate_geonames_json

This command converts supported source files into JSON datasets consumed by the Django import commands.

Supported data types include:

text
regions
subregions
cities
country
timezones
postal_codes

Generate all supported datasets:

bash
python manage.py generate_geonames_json \
    --all \
    --source geonames

Generate only cities:

bash
python manage.py generate_geonames_json \
    --cities \
    --source geonames

Generate regions:

bash
python manage.py generate_geonames_json \
    --regions \
    --source geonames

Generate postal codes:

bash
python manage.py generate_geonames_json \
    --postal_codes \
    --source geonames

A larger batch size can be supplied for large datasets:

bash
python manage.py generate_geonames_json \
    --cities \
    --source geonames \
    --batch-size 100000

The source and destination can also be overridden:

bash
python manage.py generate_geonames_json \
    --cities \
    --source geonames \
    --input-file /path/to/source-file \
    --output-dir /path/to/output

Important

generate_geonames_json is a data-preparation command.

It does not populate Django models directly.

The generated JSON becomes input for the appropriate import commands.


7. Global Reference-Data Imports

7.1 Global Regions

bash
python manage.py import_global_regions \
    --datasource geonames

A custom source file can be supplied:

bash
python manage.py import_global_regions \
    --region-file /path/to/regions.json \
    --datasource GOI

The command imports global-region information into the location domain.


7.2 Country Information

bash
python manage.py import_country_info \
    --datasource geonames

Custom files can be supplied:

bash
python manage.py import_country_info \
    --datasource geonames \
    --country-file /path/to/countries.json \
    --phone-postal-file /path/to/phone-postal.json

The command populates country information and associates countries with their global regions.


7.3 Timezones

bash
python manage.py import_timezones

A custom JSON file can be supplied:

bash
python manage.py import_timezones \
    --file /path/to/timezones.json

8. Country Location Imports

8.1 Cities

Cities are imported for a specific country.

bash
python manage.py import_cities \
    --datasource geonames \
    --country NZ

The batch size can be adjusted:

bash
python manage.py import_cities \
    --datasource geonames \
    --country NZ \
    --batch-size 50000

The input/output JSON paths can also be overridden:

bash
python manage.py import_cities \
    --datasource geonames \
    --country NZ \
    --input-file /path/to/input.json \
    --output-file /path/to/output.json

8.2 Postal Codes

Import postal codes for one country:

bash
python manage.py import_postal_codes \
    --country NZ

Import for all supported countries:

bash
python manage.py import_postal_codes \
    --all

Create missing Location records while importing:

bash
python manage.py import_postal_codes \
    --country NZ \
    --create-locations

A custom source file can be supplied:

bash
python manage.py import_postal_codes \
    --country NZ \
    --postal-file /path/to/postal-codes.json

Batch size can be adjusted for large datasets:

bash
python manage.py import_postal_codes \
    --country NZ \
    --batch-size 50000

9. Country Administrative Categories

The command:

bash
python manage.py sync_country_administrative_categories \
    --categories-file /path/to/categories.json \
    --country-mapping-file /path/to/country-mapping.json

seeds and assigns country-specific administrative categories.

Both files are required by the command.

A dry run is available:

bash
python manage.py sync_country_administrative_categories \
    --categories-file /path/to/categories.json \
    --country-mapping-file /path/to/country-mapping.json \
    --dry-run

10. Industry Classification Data

10.1 import_classifications

The classification importer supports three classification sources:

Classification Target model Source
ISIC Rev.5 Industry ISIC_SOURCE
CPC Ver.3.0 CPCCode CPC_SOURCE
HS 2022 HSCode HS_SOURCE

Run all configured classification imports:

bash
python manage.py import_classifications

Import only ISIC:

bash
python manage.py import_classifications \
    --only isic

Import CPC:

bash
python manage.py import_classifications \
    --only cpc

Import HS:

bash
python manage.py import_classifications \
    --only hs

Multiple classifications can be selected:

bash
python manage.py import_classifications \
    --only isic cpc

Validate without database writes:

bash
python manage.py import_classifications \
    --dry-run

The importer expects the relevant source paths to be available through the configured environment/data paths.


11. System Master Data

11.1 sync_system_constants

Synchronizes application master data from the constants defined by the application.

The command manages master records including:

  • Member status
  • Employment status
  • Roles
  • Departments
  • Employee types
  • Leave types
  • Teams

Run:

bash
python manage.py sync_system_constants

The operation creates missing records and updates records whose values differ from the application constants.

Team creation depends on the corresponding Department records being available.


12. Development Users

12.1 Employees

Generate employees:

bash
python manage.py create_employees \
    --count 10 \
    --country IN

Additional options include:

text
--batch-size
--seed
--dry-run
--dp
--with-members

For example:

bash
python manage.py create_employees \
    --count 20 \
    --country NZ \
    --seed 1234 \
    --with-members

--dp generates users using @djangoplay.org email addresses.


12.2 Members

Generate members:

bash
python manage.py create_members \
    --count 10 \
    --country IN

Supported options include:

text
--batch-size
--seed
--dry-run
--dp

Example:

bash
python manage.py create_members \
    --count 20 \
    --country NZ \
    --seed 1234

12.3 Development Superuser

The development toolkit also provides:

bash
python manage.py create_superuser

Optional values can be supplied explicitly:

bash
python manage.py create_superuser \
    --email admin@example.com \
    --username admin \
    --first_name Admin \
    --last_name User

The command is designed to be idempotent and is also used by masterscript.


13. Synthetic Business / Entity Data

13.1 sync_businesses

sync_businesses is the main synthetic business/entity generation command.

Generate business JSON:

bash
python manage.py sync_businesses \
    --country IN \
    --generate \
    --count 5

Generate and import:

bash
python manage.py sync_businesses \
    --country IN \
    --generate \
    --import-data \
    --count 5

The generation can be deterministic:

bash
python manage.py sync_businesses \
    --country IN \
    --generate \
    --import-data \
    --count 5 \
    --seed 1234

Import an already-generated JSON dataset:

bash
python manage.py sync_businesses \
    --country IN \
    --import-data

Generate without writing database changes:

bash
python manage.py sync_businesses \
    --country IN \
    --generate \
    --dry-run \
    --count 5

Remove generated JSON:

bash
python manage.py sync_businesses \
    --country IN \
    --cleanup-json

Soft-delete generated SYNC-* entities:

bash
python manage.py sync_businesses \
    --country IN \
    --cleanup-generated

Overwrite existing generated/imported data when required:

bash
python manage.py sync_businesses \
    --country IN \
    --generate \
    --import-data \
    --overwrite

Important

Generated business data is synthetic development data.

It should not be treated as authoritative business/reference data.


14. Finance Development Data

Finance development data is generated through a set of dedicated commands.

The higher-level command is:

bash
python manage.py generate_finance_data

It coordinates the finance-generation stages.


14.1 Finance Data Generator

Example:

bash
python manage.py generate_finance_data \
    --country IN \
    --invoices 50 \
    --billing-schedules 10

Available controls include:

text
--country
--invoices
--billing-schedules
--seed
--dry-run
--skip-reference-data
--skip-identities
--skip-billing-schedules
--skip-invoices

This allows the individual finance-generation stages to be selectively enabled or skipped.

For example, generate finance data without invoices:

bash
python manage.py generate_finance_data \
    --country IN \
    --billing-schedules 10 \
    --skip-invoices

14.2 Finance Reference Data

bash
python manage.py seed_finance_reference_data \
    --country IN

This seeds the finance reference/configuration data required by the finance generation workflow.

An issuing entity can be explicitly selected:

bash
python manage.py seed_finance_reference_data \
    --country IN \
    --issuing-entity-id 123

Dry run:

bash
python manage.py seed_finance_reference_data \
    --country IN \
    --dry-run

14.3 Finance Identities

bash
python manage.py generate_finance_identities \
    --country IN

A deterministic seed can be supplied:

bash
python manage.py generate_finance_identities \
    --country IN \
    --seed 1234

Dry run:

bash
python manage.py generate_finance_identities \
    --country IN \
    --dry-run

14.4 Billing Schedules

bash
python manage.py generate_finance_billing_schedules \
    --country IN \
    --count 10

Deterministic generation:

bash
python manage.py generate_finance_billing_schedules \
    --country IN \
    --count 10 \
    --seed 1234

Dry run:

bash
python manage.py generate_finance_billing_schedules \
    --country IN \
    --count 10 \
    --dry-run

14.5 Invoices

bash
python manage.py generate_finance_invoices \
    --country IN \
    --count 50

Deterministic generation:

bash
python manage.py generate_finance_invoices \
    --country IN \
    --count 50 \
    --seed 1234

Dry run:

bash
python manage.py generate_finance_invoices \
    --country IN \
    --count 50 \
    --dry-run

15. Finance Reference Status and Payment Methods

The command:

bash
python manage.py import_status_paymentmethods

creates or updates finance invoice statuses and payment methods.

Process both:

bash
python manage.py import_status_paymentmethods --all

Process only status data:

bash
python manage.py import_status_paymentmethods --status

Process only payment methods:

bash
python manage.py import_status_paymentmethods --payment_methods

16. Location Timezone Synchronization

After location data has been imported, timezone relationships can be updated.

For one country:

bash
python manage.py update_locations_timezones \
    --country IN

For all countries:

bash
python manage.py update_locations_timezones \
    --all

Batch size can be controlled:

bash
python manage.py update_locations_timezones \
    --country IN \
    --batch-size 1000

The command requires either --country or --all.


17. Redis Development Cache

Development data is not limited to PostgreSQL.

DjangoPlay also maintains Redis-backed representations used by the application.

The development toolkit provides:

bash
python manage.py refresh_redis_cache

This command refreshes the relevant cached data after development-data changes.

masterscript also invokes cache refresh as part of its finalization stage.


18. Permissions

The development toolkit includes a permission-granting command.

bash
python manage.py grant_permissions \
    --email user@example.com \
    --perm app_label.codename \
    --allow True

Permissions can also be selected by application/model/action:

bash
python manage.py grant_permissions \
    --email user@example.com \
    --app entities,locations \
    --model entities.entity \
    --actions add,change,view \
    --allow True

Dry run:

bash
python manage.py grant_permissions \
    --email user@example.com \
    --perm entities.view_entity \
    --allow True \
    --dry-run

A summary can be requested with:

bash
--summary

19. Development Cleanup

Development data can be cleaned using dedicated commands.

19.1 Invalid Regions

bash
python manage.py cleanup_invalid_regions

Preview the operation:

bash
python manage.py cleanup_invalid_regions \
    --dry-run

19.2 Locations

The development toolkit also contains a country-scoped location deletion command:

bash
python manage.py deletelocations \
    --country NZ

Batch size can be adjusted:

bash
python manage.py deletelocations \
    --country NZ \
    --batch-size 1000

This permanently deletes cities, locations, regions and subregions associated with the specified country.

This is a destructive development-data operation.


20. Master Development-Data Orchestration

20.1 masterscript

masterscript is the primary orchestration command for building a populated development environment.

A minimal invocation is:

bash
python manage.py masterscript \
    --iterations 1 \
    --processes 1 \
    --country IN \
    --count 5

The command requires --iterations.


20.2 MasterScript Options

Option Default Purpose
--iterations Required Number of synthetic generation iterations
--processes 1 Process configuration for country execution; capped at 4
--country IN One or more comma-separated country codes
--count 10 Entities/invoices generated per country/pass
--employee-count 20 Employees created during initialization
--member-count 20 Members created during initialization
--batch-size 25000 Batch size passed to city import
--command-timeout 1800s Override command timeout
--skipglobal Off Skip global initialization
--force-global Off Force global initialization despite existing data
--logs Off Write full command trace to masterscript.log
--min-memory-mb 150 Minimum available memory before launching subprocesses

21. MasterScript Execution Phases

The current orchestration is divided into the following phases.

Phase 0 — Superuser Bootstrap

create_superuser is always executed.

It is intentionally separate from the global initialization skip mechanism.

text
create_superuser

Because the command is idempotent, an existing superuser does not need to be manually handled before running masterscript.


Phase 0.5 — Reference Data Fetch

masterscript automatically ensures that required reference data exists locally:

text
fetch_reference_data

This happens even when --skipglobal is used.

The purpose is to ensure that the subsequent location and business-generation commands have their required JSON datasets available.


Phase 1 — Global Initialization

Unless skipped or automatically detected as already complete, the global initialization stage runs:

text
sync_system_constants
        │
        ▼
import_classifications
        │
        ▼
import_global_regions
        │
        ▼
import_country_info
        │
        ▼
sync_country_administrative_categories
        │
        ▼
import_timezones

Global initialization can be skipped explicitly:

bash
python manage.py masterscript \
    --iterations 1 \
    --country IN \
    --skipglobal

Alternatively, masterscript can automatically skip the global phase when its existing-data checks determine that initialization has already been completed.

Use:

text
--force-global

to force the global initialization phase to execute again.

If both are supplied:

text
--skipglobal
--force-global

--skipglobal takes precedence.


22. MasterScript Country Initialization

Country-specific initialization is performed for each requested country.

The current country initialization includes:

text
create_employees
        │
        ▼
create_members
        │
        ▼
import_cities
        │
        ▼
import_postal_codes
        │
        ▼
sync_businesses
        │
        ▼
generate_finance_data

The finance initialization stage generates finance supporting/reference data and billing schedules while deliberately skipping invoice generation.

Invoices are generated later during the iteration phase.


23. MasterScript Iterations

Each requested iteration performs:

text
sync_businesses
        │
        ▼
generate_finance_data
        │
        └── invoices

During the iterative finance stage:

text
--skip-reference-data
--skip-billing-schedules

are used because those stages have already been handled during country initialization.

The finance stage generates invoices for the entities generated during the current iteration.


24. MasterScript Finalization

After country processing:

text
update_locations_timezones
        │
        ▼
refresh_redis_cache

This ensures that location timezone information and Redis-backed application data are synchronized with the newly generated development records.


25. Multiple-Country Execution

Multiple countries can be supplied:

bash
python manage.py masterscript \
    --iterations 1 \
    --processes 4 \
    --country JP,NZ,US,IN \
    --count 10

When multiple countries are requested, masterscript runs each country as a separate operating-system process, sequentially.

Conceptually:

text
MasterScript
     │
     ├── Country JP → child process → complete
     │
     ├── Country NZ → child process → complete
     │
     ├── Country US → child process → complete
     │
     └── Country IN → child process → complete

The countries are not executed concurrently.

This design limits memory accumulation across large country datasets and allows the operating system to reclaim the memory of each completed child process.


26. Memory and Timeout Protection

masterscript contains explicit protections for resource-constrained development environments.

26.1 Command timeout

The normal subprocess timeout is:

text
1800 seconds

Override it with:

bash
--command-timeout 3600

Individual multi-country child processes have a larger ceiling because a child runs the complete country workflow.


26.2 Available-memory check

Before launching subprocesses, masterscript checks Linux /proc/meminfo for MemAvailable.

The default minimum is:

text
150 MB

Override:

bash
--min-memory-mb 300

Disable the check:

bash
--min-memory-mb 0

This is a protective mechanism intended to fail cleanly rather than allowing development-data generation to drive a memory-constrained machine into swap thrashing.


27. MasterScript Logging

By default, masterscript keeps console output intentionally quiet.

Use:

bash
python manage.py masterscript \
    --iterations 1 \
    --country IN \
    --count 10 \
    --logs

With --logs, the complete per-command trace is written to:

text
masterscript.log

at the repository root.

The log includes:

  • Commands executed
  • Command output
  • Phase information
  • Failures
  • Final summaries

The console remains intentionally concise.

This is particularly useful when large imports would otherwise produce thousands of lines of terminal output.


28. Recommended Development Workflow

For a new development environment, the recommended conceptual workflow is:

text
1. Prepare / fetch reference data
            │
            ▼
2. Import global reference data
            │
            ▼
3. Import country location data
            │
            ▼
4. Generate development users
            │
            ▼
5. Generate synthetic businesses
            │
            ▼
6. Generate finance data
            │
            ▼
7. Update location timezones
            │
            ▼
8. Refresh Redis

For normal development, masterscript is the preferred orchestration entry point because it coordinates these stages and performs the required checks.

Individual commands remain useful when:

  • Debugging a specific data pipeline
  • Rebuilding only one dataset
  • Testing a specific domain generator
  • Importing a newly prepared reference dataset
  • Cleaning a development environment
  • Re-running a failed phase independently

29. Deterministic Development Data

Several generation commands support:

text
--seed

A seed allows developers to make generated data deterministic for repeatable development and testing.

For example:

bash
python manage.py sync_businesses \
    --country IN \
    --generate \
    --import-data \
    --count 10 \
    --seed 1234

Finance generators and user-generation commands also support deterministic seeding.

Deterministic generation is useful when comparing application behavior across runs.


30. Dry-Run Support

Several development-data commands provide:

text
--dry-run

Dry-run mode allows the generation/import logic to be exercised without persisting the resulting database changes.

Commands supporting dry-run include, among others:

text
sync_businesses
create_employees
create_members
generate_finance_data
generate_finance_identities
generate_finance_billing_schedules
generate_finance_invoices
seed_finance_reference_data
grant_permissions
sync_country_administrative_categories
cleanup_invalid_regions

Use dry-run mode whenever evaluating a destructive or high-volume operation before writing to the database.


31. Auxiliary Developer Utilities

devtools/management/auxiliary_commands/ contains utilities that are not part of the normal development-data generation pipeline.

These include:

Utility Purpose
audit_serializers Backfill/prepare audit history data for serializers/models
backfill_history Backfill model history
check_url_conflicts Detect URL pattern conflicts
cleanup_audit_events Remove expired audit events
clear_policy_caches Clear policy-engine Redis caches
deletelocations Permanently remove country location data
find_bad_drf_field Detect serializer/model/view field configuration conflicts
format_templates Reformat Django templates
generate_schema Generate JSON schema for an application's models
remove_blank_lines Remove blank lines from Python source
table_stats Inspect database table counts and columns
validate_email_templates Validate email template syntax and renderability

These utilities support development and maintenance but should not be confused with the synthetic-data generation pipeline.


32. DevTools Service Organization

The supporting service layer is organized by domain.

Business Services

text
devtools/services/business/

Includes functionality for:

  • Business generation
  • Business JSON creation
  • Business import
  • Business persistence
  • Business synchronization
  • Business cleanup
  • Business hierarchy
  • Payload construction
  • Reporting
  • Generation context and counters

Finance Services

text
devtools/services/finance/

Includes:

  • Finance generation context
  • Finance reference-data handling
  • Finance identity generation
  • Billing schedule generation
  • Invoice generation
  • Finance generation orchestration
  • Finance constants

Location Services

text
devtools/services/locations/

Includes:

  • Global region import
  • Country information import
  • City import
  • Postal-code import
  • Timezone import
  • Location JSON generation
  • Administrative-category synchronization
  • Location timezone updates
  • Geo-services integration
  • Location import validation
  • Region-name overrides

User Services

text
devtools/services/users/

Includes:

  • Employee generation
  • Member generation

Common Services

text
devtools/services/common/

Provides reusable generation/import functionality including:

  • Faker support
  • Environment/path loading
  • Locale generation
  • Phone-number utilities
  • Postal-code utilities
  • Tax identifiers
  • Ratio/distribution helpers
  • Progress animation

33. Development Data and Application Services

The development-data tooling is intended to exercise the application's real domain structures.

The preferred pattern is:

text
DevTools Generator
       │
       ▼
Domain Service
       │
       ▼
Domain Model
       │
       ▼
PostgreSQL

Rather than constructing arbitrary database states directly, generators should create records that satisfy the application's relationships, validation rules, and workflow expectations.

This is particularly important for:

  • Business/entity relationships
  • Location hierarchy
  • Finance identities
  • Billing schedules
  • Invoices
  • User/employee/member relationships

34. Reference Data vs Synthetic Data

The distinction should remain explicit.

Reference data

Examples:

text
Countries
Regions
Subregions
Cities
Postal codes
Timezones
Industry classifications

Reference data comes from prepared external datasets and is imported into the application.

Synthetic data

Examples:

text
Businesses
Employees
Members
Finance identities
Billing schedules
Invoices

Synthetic data is generated specifically for development and testing.

The two pipelines are complementary:

text
Reference Data
     │
     ▼
Provides the geographic/classification foundation
     │
     ▼
Synthetic Data
     │
     ▼
Exercises application/domain workflows

35. Production Safety

Development-data commands can create, modify, and in some cases permanently delete database records.

They should therefore be treated as development tooling.

In particular, take care with commands such as:

text
deletelocations
cleanup_invalid_regions
sync_businesses --cleanup-generated

and any command executed against a non-development database.

--dry-run should be preferred when the command supports it and the effect of an operation is not yet understood.


36. Command Quick Reference

Command Purpose
masterscript Full development-data orchestration
fetch_reference_data Download published reference datasets
generate_geonames_json Convert source data to compiled JSON
import_global_regions Import global regions
import_country_info Import country information
import_timezones Import timezone data
import_cities Import city data
import_postal_codes Import postal-code/location data
import_classifications Import ISIC/CPC/HS classifications
sync_country_administrative_categories Seed country administrative categories
sync_system_constants Synchronize application master data
create_superuser Create/ensure development administrator
create_employees Generate employees
create_members Generate members
sync_businesses Generate/import synthetic businesses
seed_finance_reference_data Seed finance reference/configuration data
generate_finance_data Orchestrate finance test-data generation
generate_finance_identities Generate finance identities
generate_finance_billing_schedules Generate billing schedules
generate_finance_invoices Generate invoices
import_status_paymentmethods Synchronize invoice statuses/payment methods
update_locations_timezones Update location timezone relationships
refresh_redis_cache Refresh Redis-backed application data
grant_permissions Grant/revoke development permissions
cleanup_invalid_regions Remove invalid region development data
deletelocations Permanently remove country location data

37. Summary

DjangoPlay's devtools package is a structured development-data and maintenance subsystem rather than a single fixture generator.

Its primary responsibilities are:

text
Reference Data
    ├── Generate / compile
    ├── Fetch published datasets
    └── Import into Django

Synthetic Data
    ├── Users / Members / Employees
    ├── Businesses / Entities
    └── Finance / Invoices

Orchestration
    └── masterscript

Maintenance
    ├── Cleanup
    ├── Permissions
    ├── Cache refresh
    ├── Audit utilities
    └── Developer diagnostics

For normal development-data population, use:

bash
python manage.py masterscript \
    --iterations 1 \
    --processes 1 \
    --country IN \
    --count 5

For specialized work, use the individual management commands directly.

The authoritative implementation of each command is located under:

text
webapp/devtools/management/commands/

with the underlying generation and import logic implemented primarily under:

text
webapp/devtools/services/