A complete, entirely synthetic multi-billion-dollar wholesale distributor — orders, inventory, logistics, procurement, finance, workforce, fleet, and a full executive layer — built to power analytics, BI, and machine learning without a shred of real, private data.
This is not one frozen dataset; it is a dataset generator delivered as data. The entire business is produced from more than 3,300 individually adjustable values across 440+ parameter groups and roughly 30 business domains. Every count, rate, range, cost, and behavioral rule can be dialed up or down. The default models a national distributor doing $4.38 billion a year — $12.0 million every day across ~50,000 orders — turn the dials and it becomes a $400M regional player or a $40B national giant, with revenue, order volume, footprint, segment mix, pricing, costs, and geography all moving together and staying internally consistent.
Company size, number of distribution centers, customers, vendors and reps, the time window, segment behavior, discount and promotion policy, fill rates, on-time-delivery targets, return rates, payment behavior, wage bands, fuel and freight costs, benefits, and commission plans are all parameters — validated for consistency before a single row is generated.
What sets this dataset apart is how much genuine business logic is wired into it — and how the pieces interlock the way they do in a real company. A few of the systems built in:
Owned vs. leased sites carry entirely different cost structures — rent and CAM vs. depreciation, property tax, and mortgage interest — plus utilities on region-specific seasonal curves, insurance, security, and maintenance.
A DC's size and throughput drive its headcount through labor standards. Warehouse and office roles, region-specific wage bands, peak overtime, and a benefits/burden load all reconcile to real people.
Reps carry quotas and territories; pay is a regional base plus tiered commissions on pocket margin with an accelerator above quota — tied to the same profitability metric finance uses.
Fuel prices tracked by state and month, fuel surcharges, carrier tier rates by mode, cartonization and accessorials, and a company fleet whose vehicle age drives its maintenance-cost curve.
List, catalog, and contract pricing; segment- and category-specific discounts; volume breaks and elasticity; percent-off, BOGO, and clearance promotions with rebate accruals and measurable lift.
Receivables model segment payment personalities, aging, and write-offs; payables mirror them with vendor terms, early-pay capture, and peak-cash stretching — a full cash-conversion cycle.
ABC/XYZ classes derived from realized demand, lifecycle stages, health profiles per DC, and forecast error that scales realistically — tight for fast movers, wide for the long tail.
Recordable incidents calibrated to industry rates per 200,000 labor hours, equipment maintenance with downtime, and the company consuming its own catalog — all computable because the underlying hours exist.
Vendor reliability → availability → fill rates → returns → margin. DC size → staff → payroll → overhead → EBITDA. Region drives wages and utilities and insurance and delivery. Causally wired.
Most sample datasets stop at operational transactions. This one extends into the analytics a CFO, COO, CRO, and CEO actually review. At its heart is a line-level pocket-margin waterfall (~121 million rows) that walks each sale from list price through every discount, rebate, freight, and returns effect to a true contribution after cost-to-serve — so you can surface customers and segments with healthy gross margin but eroded pocket margin. Around it sits facility economics, working capital, a full revenue → gross → pocket → contribution → EBITDA bridge, headcount and payroll, sales compensation, fleet, safety, and capex. Very few datasets can populate a C-suite dashboard at all — this one is built for it.
| Typical sample datasets | PrimaForja Supply | |
|---|---|---|
| Scale | Thousands to a few million rows | ~800 million rows, 62 tables |
| Breadth | One or two departments | End-to-end: sales → ops → finance → people |
| Cross-table integrity | Loose or absent | Enforced; every domain reconciles |
| Executive / financial depth | Usually none | Full C-level layer + margin waterfall |
| Realism | Random values, visible artifacts | Interlocking business logic |
| Configurability | Fixed | 3,300+ tunable values — any size or mix |
| Reproducibility | Varies | Deterministic & versioned |
| Provenance / licensing | Often unclear | Fully synthetic, no PII; licensed for your use (showcase demos & dashboards, no raw-data redistribution) |
Delivered in Parquet, CSV, and a fully documented PostgreSQL schema. The schema is also available in the dialect for your database — SQL Server, MySQL, Oracle, or Snowflake — with types, keys, and constraints translated to match, and the Parquet/CSV data loads into any of them unchanged. The companion library of analytics and materialized views that powers the dashboards is offered as an optional add-on, and drives Tableau, Power BI, Looker, and Metabase directly. Every edition is validated against automated integrity and financial-identity checks before release, and is deterministic and versioned so every recipient works from identical data. And because it is entirely synthetic — no PII, no real customer, employee, or financial record, derived from nothing real — it carries none of the privacy risk of production data. The dataset is licensed for your own use and is not for public redistribution, but what you build with it — demos, dashboards, visualizations, and reports — is yours to share publicly.
Explore the executive dashboards built on this dataset, or get in touch for access, a sample, a reduced-scale edition, or a custom configuration.
*Default configuration. Every figure on this page is an adjustable default. The dataset is fully fabricated and contains no real personal, customer, employee, or financial data.
PrimaForja Supply is a product of PrimaForja LLC, Atlanta, Georgia.