Datasets¶
U-Chrom keeps no data in its repository. There are three ways to data:
The atlas — public datasets as .chromdata.zarr stores, opened over HTTP without downloading
them (only what an analysis or a view reads is fetched):
import uchrom.datasets as ds
ds.list_atlas() # id, title, cells, modalities, size
cd = ds.atlas("takei2025_cerebellum") # backed: cells, bins, index now; spots on demand
cell = cd.get_cell(cd.cells.index[0])
Original files — when a tutorial shows how a raw format is read (FOF-CT, .pairs, a .cool, raw
seqFISH+ detections), the file is fetched once from its source (4DN, GEO, Zenodo, GitHub, UCSC) into
the data directory and checked:
path = ds.fetch("takei") # Takei 2021 FOF-CT core table, 22 MB, 4DN
ds.list_datasets() # every source U-Chrom knows
python -m uchrom.datasets list
python -m uchrom.datasets fetch takei stevens2017
python -m uchrom.datasets fetch --default # everything the tutorials fetch (~0.9 GB)
python -m uchrom.datasets path takei
The data directory is $UCHROM_DATA, else ~/.cache/uchrom (ds.data_dir()); set UCHROM_DATA to put
large data elsewhere, or to reuse a folder of earlier downloads.
Your own data — read with the readers of uchrom.io and saved as a store
(data model & storage).
Where every dataset comes from — the study, accession, licence, how it is fetched or built (the recipes of the atlas stores and the benchmark inputs), and who uses it — is listed in data sources.