Fetching files

ds.fetch(name) puts files on your disk — once: a second call returns the path without downloading. It returns the file (or the folder of a dataset with several files) in the data directory.

An atlas store, for offline work

import uchrom.datasets as ds

path = ds.fetch("stevens2017_mesc")       # <data directory>/atlas/stevens2017_mesc.cdz (26 MB)
cd = ds.load("stevens2017_mesc")          # now reads the local copy

The .cdz is the store in one uncompressed zip: everything the remote store holds, contact maps and RNA included, read the same way (backed). Its size is ds.list_atlas()["download_mb"] — from 26 MB to several GB; the download is checked against it. Delete the file to go back to the remote store.

Original files

The registry (ds.list_datasets(): name, size, whether a loader exists, description) names every file U-Chrom uses, where it comes from and its md5 sum:

ds.list_datasets().query("size_mb < 100")
path = ds.fetch("takei")                  # Takei 2021 FOF-CT core table (4DN), md5-checked
fasta = ds.fetch("hg19_chr21")            # a reference sequence (UCSC), e.g. GC phasing of compartments
ds.path("takei")                          # where it is (or will be), without downloading

Use them to read a raw format yourself (the import tutorials do), as reference files (genome sequences, gene annotations, published loop lists), or as the inputs of the benchmarks. A few entries are built rather than downloaded whole: slices of large Hi-C files read over HTTP range requests (rao2014_imr90_chr21, rao2014_imr90_chr1, …), the head of a large archive (takei2025_fov0).

From the command line

python -m uchrom.datasets list                      # the registry
python -m uchrom.datasets atlas                     # the atlas (with what is downloaded)
python -m uchrom.datasets fetch takei hg19_chr21    # original files
python -m uchrom.datasets fetch stevens2017_mesc    # an atlas store (.cdz)
python -m uchrom.datasets fetch --default           # every file the tutorials use (~0.9 GB)
python -m uchrom.datasets path takei

--root DIR (or UCHROM_DATA) puts the files elsewhere, e.g. on a cluster’s scratch space.

Adding a dataset

A new dataset is registered in packages/uchrom/uchrom/datasets/_sources.py (files, URLs, md5 sums, size, description; a loader in _load_*.py when it should come as a ChromData) and documented in apps/atlas/recipes/README.md (data sources) in the same change — no silent downloads.