Datasets and the atlas¶
U-Chrom keeps no data in its repository. Every dataset it knows reaches you in one call, as a ChromData
or as its original files, from three places:
Where |
What it is |
Browse |
Get it |
|---|---|---|---|
The atlas |
public 3-D genome datasets, each one |
the atlas page, the web browser, |
|
Datasets with a loader |
published datasets built into a store from their original files the first time you load them |
|
|
Original files |
the files as their source publishes them (4DN, GEO, Zenodo, GitHub, UCSC): raw formats, reference genomes, annotations |
|
|
import uchrom.datasets as ds
ds.list_atlas() # the atlas: id, title, study, organism, cells, modalities, sizes
cd = ds.load("stevens2017_mesc") # an atlas store, opened over HTTP (backed: read in parts)
cd = ds.load("takei2021_mesc") # a dataset with a loader: built once, then read from disk
path = ds.fetch("takei") # an original file: Takei 2021 FOF-CT core table (22 MB, 4DN)
Which one? Analyse a dataset → ds.load (atlas id or loader name: the same call). Work offline or
read a whole atlas store many times → ds.fetch(id) once, then ds.load(id) reads the local copy.
Read a raw format yourself, or need a genome / annotation file → ds.fetch(name). Your own data
→ the readers of uchrom.io, saved as a store (data model & storage).
Everything downloaded or built goes to the data directory: $UCHROM_DATA, else ~/.cache/uchrom
(ds.data_dir()). Set UCHROM_DATA to keep large data elsewhere or to share a folder of downloads.
Step |
Page |
API |
|---|---|---|
find a dataset: the atlas page, the web browser, the catalog from Python |
|
|
a dataset as a |
|
|
files on disk: atlas stores for offline work, original files, reference files |
|
Where every dataset comes from — study, accession, licence, size, how it is fetched or built, who uses it — is listed in data sources.
Tutorial¶
Guides¶
API: uchrom.datasets.