Blog

Notes on synthetic data

Field notes on building believable synthetic businesses — the realism engine, the dashboards it powers, and what it takes to generate data an expert can't tell from the real thing.

Anonymized isn't anonymous: why synthetic is the safer default

Stripping names off real data doesn't make it safe — re-identification is a solved attack. Why data that never touched a real person is the only clean answer, and the one caveat that matters.

Read the article →

Charting data on different scales without skewing the story

Revenue in millions next to a fill rate in percent: the classic setup for a chart that lies. Dual axes, indexing, log scales, and small multiples — and when each one is the honest choice.

Read the article →

Choosing the right chart — and the right color

A field guide to matching charts to questions, varying the form so a dashboard stays legible, and using color that means something — illustrated with the distributor data.

Read the article →

What you can actually learn: ML tasks the dataset supports

Forecasting, churn, price elasticity, LTV, anomaly detection, and more — the signal is real because the business behind the data is coherent.

Read the article →

Stress-test your BI stack with synthetic data

Load testing, multi-dialect migrations, dashboard QA, and training — the jobs a realistic synthetic dataset does better than real production data.

Read the article →

Why 'realistic' is the hard part of synthetic data

Anyone can generate a million rows. Making them behave as if a single coherent business produced them is where the real work — and the value — is.

Read the article →

Designing a KPI that reconciles

Two numbers that should match but don't will sink a dashboard. The two failure modes — mixed revenue bases and mixed period grains — and how to fix and guard against them.

Read the article →

Reading a pocket-margin waterfall

Gross margin hides the discounts, rebates, and freight that decide whether a deal actually made money. How the waterfall exposes every leak — and why net merchandise is the revenue anchor.

Read the article →

Multi-dialect by design: one dataset, five databases

Postgres, SQL Server, MySQL, Oracle, and Snowflake from one canonical schema — the transpilation decisions that keep them in lockstep, and why views are a separate add-on.

Read the article →

From 200 rows to 40 million: streaming generation at scale

How the pipeline builds hundreds of millions of interrelated rows with bounded memory — and keeps the derived analytics correct by accumulating summaries as the data streams past.

Read the article →

More articles are on the way. Want a topic covered? Let us know.