Why 'realistic' is the hard part of synthetic data
Anyone can generate a million rows. Point a fake-data library at a schema, let it fill in random names and numbers, and you'll have something that looks like a dataset in minutes. The trouble starts the moment someone who knows the business looks at it.
They notice that the "distributor" ships more units than it stocks. That a customer's payments don't add up to their invoices. That every product sells at exactly the same rate, that November looks like March, that the biggest accounts churn as often as the smallest, and that the margin on the finance dashboard has no relationship to the prices on the orders. The data isn't wrong in any single cell — it's wrong in the relationships between cells. And relationships are where realism lives.
Realism is a property of the whole, not the row
A believable business dataset has to be internally consistent across every department at once. Revenue in the finance tables has to reconcile to the orders in the sales tables. The labor hours in the warehouse tables have to tie to the throughput in the fulfillment tables. Vendor reliability has to flow through to inventory availability, which flows through to fill rates, which flow through to customer satisfaction and returns, which flow back into margin.
The hard part of synthetic data isn't generating values. It's making millions of independently generated values behave as if a single coherent business produced them.
That's why we don't generate tables in isolation and stitch them together afterward. We build the business in dependency order — the reference entities first, then the transactions they drive, then the analytics derived from those transactions — so that every downstream number inherits the causes that should have produced it.
Interlocking systems, not random draws
Consider something as ordinary as a distribution center. In a naive dataset its headcount is just a number someone typed in. In a realistic one, its size drives its throughput, which drives its staffing through labor standards; its region sets the wage bands and the seasonal shape of its utility bills; and whether it's owned or leased changes its entire cost structure — rent and CAM versus depreciation, property tax, and mortgage interest. Get those links right and a fully-loaded cost-per-order analysis holds up. Get them wrong and a CFO spots it in seconds.
Multiply that by every subsystem in the company — pricing and promotions, receivables and payables, forecasting, fleet, safety, compensation tied to actual margin — and you start to see why "realistic" is the expensive word in the phrase "realistic synthetic data."
The tell of a good dataset: it has something to find
A dataset that's too clean is its own kind of fake. Real businesses carry a tail of unprofitable customers, a handful of delinquent accounts, some late vendors, damaged returns, and the occasional price that slipped below cost. We model that imperfection on purpose — not as noise, but as structure — so that an analyst exploring the data actually discovers something, the way they would in a real company.
That's the bar we set for ourselves: a distribution-industry veteran should look at the data and recognize it, and a model or dashboard built on it should be something you can trust, because the numbers reconcile all the way up to the boardroom and all the way down to the individual shipment.
Why it matters for you
If you're demoing a BI product, training a model, or stress-testing a pipeline, the realism of your data is the ceiling on how convincing your work can be. Random rows get you a screenshot. A coherent business gets you a conversation about outcomes.
That's the whole reason PrimaForja Supply exists — and it's what every one of the 3,300+ configuration values, the 62 interrelated tables, and the 34 validation checks is ultimately in service of.
Want to see it for yourself?
Explore the dashboards built on the dataset, or request access to work with it directly.