How it works

A whole business, generated on demand

PrimaForja Supply isn't a static file we exported once — it's the output of a configuration-driven engine that builds a complete, self-consistent distributor from the ground up, table by table, in dependency order.

62
interrelated tables
809M
rows generated
3,300+
configurable values
34
automated quality checks

From configuration to dataset in five stages

Configure

Every count, rate, range, cost, and behavioral rule lives in one place. Set the company's size, footprint, product mix, pricing, and dozens of other levers — or keep the defaults. The engine validates the whole configuration for internal consistency before it writes a single row.

Build the foundation

Reference entities come first — distribution centers, customers, vendors, manufacturers, the product catalog, carriers, and the sales organization — each with realistic, differentiated characteristics that everything downstream depends on.

Generate the transactions

Orders, order lines, shipments, tracking events, invoices, receivables, backorders, returns, and purchase orders are produced with realistic seasonality, weekday patterns, and segment-driven behavior — streamed in chunks so scale never blows up memory.

Derive the analytics layer

Inventory positions, forecasts, price history, promotions, the monthly history tables, and the full executive financial layer — including the line-level margin waterfall — are computed from the transactions that came before.

Validate & deliver

An automated suite of integrity, arithmetic, and financial-identity checks runs before release. Nothing ships unless it passes. The result lands as Parquet, CSV, or a ready-to-load PostgreSQL database.

What makes it different under the hood

🎛️

Configuration-driven

3,300+ tunable values across ~30 business domains. There are no magic numbers buried in code — every behavior traces back to a documented setting you can change.

🔁

Deterministic

The same settings produce the same dataset, every time. Results are comparable across teams, demos are repeatable, and a benchmark you run today reproduces months from now.

🔗

Referentially whole

Every foreign key resolves and every domain reconciles — revenue ties to orders, labor hours tie to throughput, vendor performance flows through to fill rates.

📈

Scales without breaking

Streaming generation keeps peak memory bounded regardless of size, so the default runs to hundreds of millions of rows on ordinary hardware.

🧮

Financially coherent

A line-level pocket-margin waterfall walks every sale from list price to true contribution — the same way a real finance team measures profitability.

Validated every run

34 automated assertions check referential integrity, arithmetic, financial identities, and realism constraints. An edition that fails is never released.

Delivered in the formats you already use

🅿️

Parquet

Columnar, compressed, and typed — ideal for data lakes and distributed compute (Spark, DuckDB, Snowflake, BigQuery, Databricks).

📄

CSV

Universally importable into any tool or database, for when you just need rows.

🐘

PostgreSQL

A complete documented schema with keys, indexes, and constraints, ready to load — the same schema our live dashboards run on.

🗄️

Your database

Need SQL Server, MySQL, Oracle, or Snowflake? We provide the schema in the right dialect for your engine — types, keys, and constraints translated to match. The Parquet/CSV data loads into all of them unchanged.

📈

Analytics views (add-on)

The library of pre-built analytics and materialized views that powers the dashboards is available as an optional add-on, as regular or materialized views for your database.

See it running

Explore the executive dashboards built directly on this dataset, or request access for a guided walkthrough and a copy to work with.

Explore the Dashboards