The PrimaForja Synthetic Distributor Dataset

A production-scale, fully configurable, deeply realistic dataset

A complete, entirely synthetic multi-billion-dollar wholesale distributor — orders, inventory, logistics, procurement, finance, workforce, fleet, and a full executive layer — built to power analytics, BI, and machine learning without a shred of real, private data.

809M
rows of data
62
interrelated tables
3,300+
configurable values
$12M/day
modeled revenue*
2 yrs
of daily activity
C-level
dashboard-ready

Everything is configurable — nothing is set in stone

This is not one frozen dataset; it is a dataset generator delivered as data. The entire business is produced from more than 3,300 individually adjustable values across 440+ parameter groups and roughly 30 business domains. Every count, rate, range, cost, and behavioral rule can be dialed up or down. The default models a national distributor doing $4.38 billion a year — $12.0 million every day across ~50,000 orders — turn the dials and it becomes a $400M regional player or a $40B national giant, with revenue, order volume, footprint, segment mix, pricing, costs, and geography all moving together and staying internally consistent.

EVERY NUMBER ON THIS PAGE IS A DEFAULT YOU CAN CHANGE.

Model the business you want

Company size, number of distribution centers, customers, vendors and reps, the time window, segment behavior, discount and promotion policy, fill rates, on-time-delivery targets, return rates, payment behavior, wage bands, fuel and freight costs, benefits, and commission plans are all parameters — validated for consistency before a single row is generated.

The realism engine

What sets this dataset apart is how much genuine business logic is wired into it — and how the pieces interlock the way they do in a real company. A few of the systems built in:

🏢

Facility economics

Owned vs. leased sites carry entirely different cost structures — rent and CAM vs. depreciation, property tax, and mortgage interest — plus utilities on region-specific seasonal curves, insurance, security, and maintenance.

👷

Workforce & payroll

A DC's size and throughput drive its headcount through labor standards. Warehouse and office roles, region-specific wage bands, peak overtime, and a benefits/burden load all reconcile to real people.

💼

Sales compensation

Reps carry quotas and territories; pay is a regional base plus tiered commissions on pocket margin with an accelerator above quota — tied to the same profitability metric finance uses.

Transportation & fuel

Fuel prices tracked by state and month, fuel surcharges, carrier tier rates by mode, cartonization and accessorials, and a company fleet whose vehicle age drives its maintenance-cost curve.

🏷️

Pricing & promotions

List, catalog, and contract pricing; segment- and category-specific discounts; volume breaks and elasticity; percent-off, BOGO, and clearance promotions with rebate accruals and measurable lift.

💵

Working capital

Receivables model segment payment personalities, aging, and write-offs; payables mirror them with vendor terms, early-pay capture, and peak-cash stretching — a full cash-conversion cycle.

📦

Inventory & forecasting

ABC/XYZ classes derived from realized demand, lifecycle stages, health profiles per DC, and forecast error that scales realistically — tight for fast movers, wide for the long tail.

🦺

Safety & internal ops

Recordable incidents calibrated to industry rates per 200,000 labor hours, equipment maintenance with downtime, and the company consuming its own catalog — all computable because the underlying hours exist.

🔗

It all interlocks

Vendor reliability → availability → fill rates → returns → margin. DC size → staff → payroll → overhead → EBITDA. Region drives wages and utilities and insurance and delivery. Causally wired.

Built for the boardroom — the C-level layer

Most sample datasets stop at operational transactions. This one extends into the analytics a CFO, COO, CRO, and CEO actually review. At its heart is a line-level pocket-margin waterfall (~121 million rows) that walks each sale from list price through every discount, rebate, freight, and returns effect to a true contribution after cost-to-serve — so you can surface customers and segments with healthy gross margin but eroded pocket margin. Around it sits facility economics, working capital, a full revenue → gross → pocket → contribution → EBITDA bridge, headcount and payroll, sales compensation, fleet, safety, and capex. Very few datasets can populate a C-suite dashboard at all — this one is built for it.

How it compares

Typical sample datasetsPrimaForja Supply
ScaleThousands to a few million rows~800 million rows, 62 tables
BreadthOne or two departmentsEnd-to-end: sales → ops → finance → people
Cross-table integrityLoose or absentEnforced; every domain reconciles
Executive / financial depthUsually noneFull C-level layer + margin waterfall
RealismRandom values, visible artifactsInterlocking business logic
ConfigurabilityFixed3,300+ tunable values — any size or mix
ReproducibilityVariesDeterministic & versioned
Provenance / licensingOften unclearFully synthetic, no PII; licensed for your use (showcase demos & dashboards, no raw-data redistribution)

Formats, quality & privacy

Delivered in Parquet, CSV, and a fully documented PostgreSQL schema. The schema is also available in the dialect for your database — SQL Server, MySQL, Oracle, or Snowflake — with types, keys, and constraints translated to match, and the Parquet/CSV data loads into any of them unchanged. The companion library of analytics and materialized views that powers the dashboards is offered as an optional add-on, and drives Tableau, Power BI, Looker, and Metabase directly. Every edition is validated against automated integrity and financial-identity checks before release, and is deterministic and versioned so every recipient works from identical data. And because it is entirely synthetic — no PII, no real customer, employee, or financial record, derived from nothing real — it carries none of the privacy risk of production data. The dataset is licensed for your own use and is not for public redistribution, but what you build with it — demos, dashboards, visualizations, and reports — is yours to share publicly.

See it in action

Explore the executive dashboards built on this dataset, or get in touch for access, a sample, a reduced-scale edition, or a custom configuration.

*Default configuration. Every figure on this page is an adjustable default. The dataset is fully fabricated and contains no real personal, customer, employee, or financial data.

PrimaForja Supply is a product of PrimaForja LLC, Atlanta, Georgia.