uchrom.datasets

Every dataset U-Chrom knows, as a ChromData (load: an atlas store over HTTP, or a dataset built once from its original files by a registered loader) or as its original files (fetch). See Datasets; where each dataset comes from: Datasets: sources and recipes.

Datasets: every dataset U-Chrom knows, as a ChromData — or as its original files.

load gives a ChromData: a store of the public atlas (opened over HTTP, backed: only what a view or an analysis reads is fetched), or a dataset built once from its original files by a registered loader and cached in the data directory:

import uchrom.datasets as ds
cd = ds.load("takei2025_cerebellum")             # the atlas store, https://uchrom-atlas-r2.u-science.org/...
cd = ds.load("kim2020_scihic")                   # 1,931 cells, cell types, per-cell contact maps (linked)

Original files (4DN, GEO, Zenodo, GitHub, UCSC) — what the readers of uchrom.io read, the inputs of the benchmarks — are downloaded once into the data directory ($UCHROM_DATA, else ~/.cache/uchrom) and checked:

path = ds.fetch("takei")                          # .../4DNFIHF3JCBY.csv (Takei 2021 FOF-CT)
ds.list_datasets()                                # what there is, sizes, descriptions

Command line: python -m uchrom.datasets list | atlas | load NAME | fetch NAME ... | path NAME.

uchrom.datasets.atlas(name: str, *, root: str | None = None, backed: bool = True, **read_kw)[source]

Open a dataset of the atlas by its id (list_atlas()), backed over HTTP by default.

uchrom.datasets.data_dir() → Path[source]

Where datasets are downloaded and built: $UCHROM_DATA, else ~/.cache/uchrom.

The layout below it is the same for every dataset (e.g. 4DNFIHF3JCBY.csv, stevens2017_mesc/..., schicar_mop/raw/...), so a folder of earlier downloads can simply be named by UCHROM_DATA.

uchrom.datasets.fetch(name: str, *, root: Path | None = None, extract: bool = True, verbose: bool = True) → Path[source]

Download a dataset from its original source into the data directory (files already there are kept) and check its md5 sums; returns path().

uchrom.datasets.files(name: str, *, root: Path | None = None) → List[Path][source]

The local paths of a dataset’s files (downloaded or not).

uchrom.datasets.list_atlas(root: str | None = None)[source]

The datasets of the atlas (chromdata.catalog.DEFAULT_ATLAS unless root) as a table.

uchrom.datasets.list_datasets()[source]

The datasets fetch() knows: name, size (MB), whether load() gives it as a ChromData, whether fetch --default (the tutorials’ inputs) includes it, description.

uchrom.datasets.load(name: str, *, root: Path | None = None, backed: bool | None = None, rebuild: bool = False, **read_kw)[source]

A dataset as a ChromData.

  • A dataset with a loader (list_datasets(), column load): built once from its original files (fetched with fetch()) into <data directory>/<dataset>/<dataset>.chromdata.zarr with its linked files next to it, then read from there (in memory unless backed=True). rebuild=True builds it again.

  • Otherwise a store of the atlas (list_atlas()): opened over HTTP, backed unless backed=False.

read_kw go to chromdata.ChromData.read() (columns=, tracks=, …).

uchrom.datasets.path(name: str, *, root: Path | None = None) → Path[source]

Where a dataset is: its file if it has one, else the folder its files share.