How downloading works

Every order has one link. Before payment it shows the invoice, after payment it becomes the download page. Keep it — it is your access.

The manifest

Your order page links a manifest.json: every file with its name, size, sha256 and which older version it replaces. The sync script reads it, compares it with your folder and fetches only what is missing.

The script

curl -fsSL https://keydata.shop/client/fetch-dataset.sh -o fetch-dataset.sh
bash fetch-dataset.sh <your manifest url> ./data

It downloads several files at a time, resumes after an interruption instead of starting over, and verifies every checksum. Re-run it any time: files you already have are skipped.

Layout on disk

Everything lands in one folder. Names carry the source, the coin, the day and the version, so nothing collides:

polymarket.btc.5m_20260925.v4.parquet
binance.btc_20260925.v4.parquet
chainlink.btc_20260925.v4.parquet

Aggregates carry .1s.parquet instead: the same names, one row per second.

If a day is ever re-published, the new file arrives under a new version and the manifest says which one it replaces, so the script can clear the old one.

Reading the files

Everything is Apache Parquet, so there is nothing to unpack. Open a file directly:

import pandas as pd
df = pd.read_parquet("polymarket.btc.5m_20260925.v4.parquet")

# or without loading it all, in DuckDB:
#   select count(*) from 'polymarket.btc.5m_20260925.v4.parquet'

polars.read_parquet, pyarrow.parquet, Spark and ClickHouse read the same files. Each file's schema is stored in its footer.

One row per message, with the microsecond receive timestamp. Sorting several files on it merges every venue into one time-ordered stream.

Limits

Downloads stay open for 30 days after payment. There is a transfer cap per order, set well above the size of what you bought — enough to pull it more than once, not enough to mirror the link for strangers.