Example data¶
Catalog of the datasets used by the tutorials (under tutorials/) and
benchmarks (under benchmarks/). Small files (< 10 MB) are checked
into the repo; larger files are listed here with their source URL and
auto-downloaded by the tutorials on first run into this directory.
Contents at a glance¶
File / target location |
Size |
Shipped in repo |
Auto-download |
Used by |
|---|---|---|---|---|
|
4.7 MB |
yes |
— |
|
|
5.4 MB |
yes |
— |
only |
|
280 KB |
yes |
— |
generic Hi-C matrix for tests ( |
|
3.8 MB |
yes (only this file of the folder) |
built ( |
|
|
58 KB |
yes |
— |
Takei 2021 cell table (RNA counts, nuclear area) for |
|
219 KB |
yes |
— |
Takei 2021 nascent RNA spots for |
|
~750 KB |
missing — listed but never committed (see its section) |
— |
|
|
1.4 MB |
yes |
— |
web browser: single-cell 3-D structure + its own contacts ( |
|
24 KB each |
yes |
— |
imputation tests / tutorial (Huang 2021 mESC Sox2, 5 cells; see |
|
73 MB (h5cd 1.x: 220 MB) |
no |
Zenodo 7693825 (streamed) |
web browser: 59 IF tracks per spot, 47 cells, IF embedding ( |
|
~400 MB |
no |
built from GEO GSE305439 |
web browser: scHiCAR without 3-D, RNA / ATAC / Hi-C embeddings ( |
|
2.85 GB download; 7.9 GB raw CSV; combined store 2.05 GB as format 2.2 ( |
no |
Zenodo 7693825 (opt-in, |
|
|
1.68 GB + 17 MB |
no |
4DN public S3 (opt-in, |
largest wild-type 4DN FOF-CT; |
Takei 2025 other Zenodo archives (cerebellum rep 2, E14, NMuMG) |
5.3–11 GB each |
no |
manual |
|
|
22 MB |
no |
4DN public S3 |
|
|
47 MB |
no |
4DN public S3 ( |
|
|
144 MB |
no |
Zenodo 3735329 |
|
|
2 MB |
no |
GitHub raw |
|
|
43 MB |
no |
GEO GSE80280 ( |
Fig. 3a (NucDynamics reference; |
|
7.8 MB + 20 MB |
no |
GitHub raw ( |
Fig. 3b ( |
|
261 MB + 1 MB |
no |
Zenodo 3928890 ( |
Fig. 3b ( |
|
200 MB |
no |
Zenodo 3928890 ( |
|
|
2.8 MB |
yes |
Zenodo 3928890 ( |
|
|
3.9 MB |
no (GPL-3 data of github.com/alberlab/igm; downloaded on demand into |
alberlab/igm repository, |
IGM protocol of the native engine vs the original IGM ( |
|
4.6 GB |
no |
4DN public S3 (opt-in, |
|
|
12.6 GB |
no |
4DN public S3 (opt-in, |
|
bulk IGM inputs ( |
0.1–10 GB |
no |
derived on Sherlock ( |
|
|
40 MB |
no |
GitHub, pinned commit ( |
Fig. 3a / 3c ( |
|
3.8 MB each |
IMR90: yes; K562: no |
built from GEO GSE63525 |
Fig. 3b / 3c; IMR90 also the GEM-FISH tutorial |
|
7.8 MB + 3.3 MB |
no |
built the same way ( |
Fig. 3c chr20 / chr22 ( |
|
849 MB |
no |
4DN public S3 (opt-in, |
provenance of |
|
78 KB + 31 KB |
yes |
— |
|
|
1.39 GB (+5 GB streamed, 253 MB kept) |
no |
schic2 S3 / GEO GSE94489 / UCSC (opt-in, |
|
|
597 MB |
no |
Zenodo 4708112 (opt-in, |
|
|
13 GB |
no |
GEO GSE146397 per-GSM (opt-in, |
|
|
~35 MB each |
no |
from the zip |
|
|
128 MB |
no |
UW Noble lab |
|
|
92 KB |
no |
UW Noble lab |
|
|
24 MB |
no |
generated |
output of |
|
~6 MB |
no |
generated |
synthetic offline fallback for tutorials 4/5/6/7 |
|
~40 GB |
no |
manual (too large) |
|
|
61 MB |
no |
GEO GSE305439 |
scHiCAR tutorials (upstream input) |
|
1.47 GB |
no |
GEO GSE305439 (opt-in, |
|
|
1.35 GB + 13 MB |
no |
GEO GSE305439 + UCSC (opt-in, |
|
|
332 MB |
no |
Brain Image Library g.21 |
|
|
28 MB – 0.9 GB |
no |
derived: |
Fig 6c browser benchmarks |
|
— |
no |
derived — recipe pending |
scHiCAR tutorials (see below) |
*.csv, *.h5cd, *.chromdata.zarr, *.cdz (except
sim_cell1/sim_cell1.cdz) and the DNAseqFISH+ zip/folder are ignored by
git; they’re either generated or downloaded on demand. The builders write
the format 2.0 .chromdata.zarr store (--format h5cd still writes the
deprecated HDF5 file); older .h5cd builds keep working (they are read,
and python -m uchrom.io.upgrade converts them).
Small, in-repo datasets¶
cell1.pairs — single-cell Hi-C read pairs (Stevens 2017)¶
Format: plain-text
.pairs(no header) with 7 tab-separated columnsread_id, chrom1, pos1, chrom2, pos2, strand1, strand2.Content: 105,700 paired-end reads from one haploid mESC G1 cell (Stevens Cell 1, mm10).
Source: Stevens et al. 2017, Nature 544:59–64, “3D structures of individual mammalian genomes studied by single-cell Hi-C”, doi:10.1038/nature21429. GEO accession GSE80280 (Cell 1 = GSM2219497). Earlier versions of this catalog cited GSE80006, which is Flyamer et al. 2017 (oocyte / zygote snHi-C), not this study. Distributed with the Nuc Dynamics software (github.com/tjs23/nuc_dynamics,
example_chromo_data.tar.gz→Cell_1_contacts.ncc): every distinct contact ofcell1.pairsappears in that NCC file (minus-strand ends differ by ~70 bp:cell1.pairsstores the read start on both strands). The GEO per-cell file (stevens2017_mesc/) has 111,838 contacts.Used by:
tutorials/reconstruction.ipynb(Nuc Dynamics worked example); mentioned in the rootREADME.mdCLI demo.
cell2.pairs.gz — single-cell Hi-C (v1.0 pairs format, gzipped)¶
Format: gzipped pairs v1.0 (hickit-style header:
#sorted,#shape: upper triangle, 20 mm10#chromosomelines, columnsreadID chr1 pos1 chr2 pos2 strand1 strand2 phase0 phase1), 503,696 contacts (all distinct; 343,247 intra-, 160,449 inter-chromosomal).Source: UNVERIFIED — do not cite it as any published cell. Added in commit
cb6ba3e(2025-09-11, from the older NucBox archive, dated 2024-05) with no provenance note. It is not Stevens 2017: those cells are haploid and unphased, and GSE80280 has 31–112 k contacts per cell; this file carries per-end haplotype phases (0 / 1 /., 23 % of contacts phased on both ends), i.e. a diploid, phased mouse cell processed with hickit (Dip-C-style; e.g. Tan et al. 2019 / 2021 or HiRES, Liu et al. 2023 — none checked). A 30-min search (2026-09-27) did not identify the source cell.Used by: only
tests/smoke/fasthigashi_smoke.py, a manual FastHigashi smoke run (not collected bypytest tests) that needs any second mm10 cell. Nothing in the tutorials, benchmarks or figures uses it. Replace or remove it once that smoke run has another input.
takei2025_cerebellum_fixture/ — Takei 2025 cerebellum DNA seqFISH+ slice (~750 KB)¶
Not in the repository. This entry predates the checks: the folder (and its
derive_fixture.py/verify.py) was never committed. For a real Takei 2025 subset usetakei2025_fov0/below; the description is kept for whoever restores the fixture.
Format: three CSVs that together exercise the
read_seqfish_multiomicsloader end-to-end —dna_spots.csv— 1 070 rows fromcerebellum_rep1_pos0.csv(rep 1, FOV 0, 6 cells, chr19 only), all 76 source columns preserved verbatim (59 z-scores + 3 DBSCAN allele variants + μm coordinates +dot_int / n_rad_score / n_per_dist(um)).locus_annotation.csv— the 838 chr19 rows fromLC1-100k-09022022-mm10-25kb-meta.csvtrimmed toname/chrom/start/end.clustering.csv— the 6 matching rows from100k-002-001-cerebellum_mRNA_cluster_nuc_vol_filtered.csv.derive_fixture.py— the slicing script for reproducibility. Not executed in CI.verify.py— five-layer correctness check (HDF5 layout +ChromDataconsistency + value-level row reconciliation against the source CSVs + semantic checks +.h5cdround-trip). Runpython example-data/takei2025_cerebellum_fixture/verify.pyto re-validate; pass--skip-value-reconplus--spot-glob/--locus/ --clusteringpaths to validate the full Zenodo data instead.
Cell-type coverage: the 6 cells span leiden clusters {0, 2, 3, 4, 6, 7} → cell types Granule, Bergmann, Other, MLI1, Purkinje, MLI2+PLI (one cell per type).
Source: Takei et al. 2025, Nature “Spatial multi-omics reveals cell-type-specific nuclear compartments” (doi:10.1038/s41586-025-08838-x). Raw data: Zenodo record 7693825. Locus annotation + clustering CSVs: CaiGroup/dna-seqfish-plus-multi-omics GitHub repo.
Used by:
seqfish_multiomics_cerebellum.ipynb(loader walk-through.h5cdround-trip).
Real-data run: the full distribution (rep 1 + rep 2 tarballs, ~8 GB compressed; tens of GB uncompressed; tens of millions of spots) is not auto-fetched. Manual download:
wget https://zenodo.org/records/7693825/files/cerebellum_rep1.tar.gz wget https://zenodo.org/records/7693825/files/cerebellum_rep2.tar.gz git clone https://github.com/CaiGroup/dna-seqfish-plus-multi-omics.git
Then point the loader at the extracted CSV(s):
cd = ChromData.from_seqfish_multiomics( spot_glob='cerebellum_rep*/cerebellum_rep*_pos*.csv', locus_annotation='dna-seqfish-plus-multi-omics/data/annotation/' 'LC1-100k-09022022-mm10-25kb-meta.csv', cell_clustering='dna-seqfish-plus-multi-omics/data/cerebellum/' 'clustering/100k-002-001-cerebellum_mRNA_cluster_nuc_vol_filtered.csv', )
4DNFIFINA2U9.csv + 4DNFIJ52NVDV.csv — Takei 2021 FOF-CT companion tables (58 KB + 219 KB)¶
Format: 4DN FOF-CT v0.1 CSV (
##-headers; both files label their namespace4dn_FOF-CT_qualityalthough they are the cell and RNA tables).4DNFIFINA2U9.csv— cell data, 201 rows:Cell_ID, Extra_Cell_ROI_ID, keep1, cent_ROI_x, cent_ROI_y, area(um2), area_cyto(um2)+ RNA copy numbers of 45 genes (Eef2 … Zfp352).4DNFIJ52NVDV.csv— nascent RNA spots, 3,335 rows:Spot_ID, X, Y, Z (µm), RNA_name, Gene_ID, Cell_ID, Extra_Cell_ROI_ID, Peak_Intensity; 23 genes.
Pairs with the core table
4DNFIHF3JCBY.csv(same 201Cell_IDs, same micron frame): 4DN experiment4DNEXM45AILX, experiment set4DNESL2AY9CM.Source: Takei et al. 2021, Nature 590:344–350, “Integrated spatial genomics reveals global architecture of single nuclei”. Downloaded from the 4DN open-data bucket:
https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/c5bafa92-d08b-4e20-84de-6e75103c016d/4DNFIFINA2U9.csv https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/4b66ffc5-50e0-4832-99c2-0020c934ff24/4DNFIJ52NVDV.csv
SHA-256
ebeca858…9941/70171dfd…bd2.Used by:
ChromData.from_fofct(core, cell_table=…, rna_table=…)→cd.cells(rna.<gene>columns,nucleus_area_um2, …) andcd.points['rna'];tests/core/test_fofct_companions.py; the web browser (colour cells by a gene, show RNA spots).The same set also has sequential IF (17 chromatin marks) described in the paper; it is not among the 4DN FOF-CT files.
K562_chr21_30kb.cool — K562 chr21 at 30 kb (0.3 MB; formerly misnamed IMR90_chr21_30kb.cool)¶
Provenance correction (Fig. 3 data check, 2026-09): despite its name, this file is K562 Hi-C. Its source, 4DN 4DNFI4QQPDMR (849,258,990 B, md5
84f4e708e2e7b55078b670ed4b8db709), is “in situ Hi-C on non-heat treated K562 cells with MboI” — experiment set 4DNESU95RUNO, Ray et al. 2019 (PMID 31506350, GEO GSE130758), Lis lab — not Rao et al. 2014 IMR90. Re-deriving chr21 from that mcool with the recipe below gives a matrix identical to this file (all 1 557 × 1 557 entries). Its chr21 total is 1.06 M contacts vs 9.2 M in real Rao 2014 IMR90. It was renamed fromIMR90_chr21_30kb.cool(no alias kept); do not pair it with IMR90 imaging. Real Rao 2014 IMR90: GEO GSE63525 (hg19; sliced bybuild_rao2014_slices.py→rao2014_imr90/) or 4DN set 4DNES1ZEJNRU (GRCh38 mcools 4DNFIR1JDZH7, 4.6 GB, and 4DNFIJTOIGOI, 8.3 GB — both range-readable over HTTP).Format: single-resolution cooler. 1 557 bins × 30 kb covering chr21 (hg38).
Content: K562 in situ Hi-C pair counts (Ray et al. 2019, control condition), aggregated from the native 5 kb to 30 kb.
Source: 4DN accession 4DNFI4QQPDMR (849 MB K562
.mcool,download_data.py --ray2019_k562; we extract chr21 at 30 kb into this small.coolfor redistribution).How it was derived:
from cooler import Cooler; import cooler, numpy as np, pandas as pd c5 = Cooler('4DNFI4QQPDMR.mcool::/resolutions/5000') m = c5.matrix(balance=False, as_pixels=False).fetch('chr21') # aggregate 5 kb × 6 → 30 kb by summation f = 6; n5 = m.shape[0]; n30 = (n5 + f - 1) // f agg = np.zeros((n30, n30)) for i in range(n30): for j in range(n30): agg[i,j] = m[i*f:(i+1)*f, j*f:(j+1)*f].sum() bins = pd.DataFrame({ 'chrom': ['chr21']*n30, 'start': np.arange(n30)*30_000, 'end': np.minimum((np.arange(n30)+1)*30_000, c5.chromsizes['chr21']), }) iu = np.triu_indices(n30, k=0) pixels = pd.DataFrame({'bin1_id': iu[0], 'bin2_id': iu[1], 'count': agg[iu].astype(np.int64)}) pixels = pixels[pixels['count'] > 0] cooler.create_cooler('K562_chr21_30kb.cool', bins, pixels, assembly='hg38') # Balance with ICE so reconstruct_gem_fish can pull balanced counts: # python -m cooler balance K562_chr21_30kb.cool
Used by: tests that need any small Hi-C matrix (
tests/strc/test_convention.py::test_call_tads_di_matches_kernel_on_linked_cool,tests/io/test_io_formats.py::test_load_cool_pixels,tests/recon/test_gem_fish.pywith synthetic FISH),benchmarks/cd2_callers_validation.py(DI caller before / after the calling convention) and the DI example indocs/source/guide/structure.md. The GEM-FISH tutorial used it as “IMR90” until 2026-09; it now usesrao2014_imr90/(real IMR90).
Large datasets — auto-downloaded on first run¶
The tutorials locate these files via a find_*() helper that:
Returns any cached copy under
example-data/or~/Downloads/….Otherwise downloads into
example-data/and caches for the next run.Falls back to a small synthetic dataset if the network is unreachable (only applies to the 4DN CSV path; the Zenodo zip has no synthetic fallback).
Every tutorial’s first data-loading cell prints the location it ends up using.
IMR90_chr21_pyhim.ecsv — Bintu 2018 IMR90 chr21 in PyHiM ECSV format (5.7 MB)¶
Format: PyHiM chromatin trace table (Astropy ECSV). Columns:
Spot_ID, Trace_ID, x, y, z, Chrom, Chrom_Start, Chrom_End, ROI #, Mask_id, Barcode #, label.meta['comments']carriesxyz_unit=micron,genome_assembly=hg38.Content: hg38 chr21:18.6–20.6 Mb, 1,277 traces × 66 loci × 30 kb spacing. IMR90 fibroblasts.
Source: Bintu et al. 2018, Science 362:eaau1783, “Super-resolution chromatin tracing reveals domains and cooperative interactions in single cells”. Original CSV at mendeley.com/datasets/3jkp7zhwbr/1.
How it was derived: The original Bintu CSV has columns
Chromosome index, Segment index, Z, X, Y(nm). The conversion scriptbintu_to_pyhim_ecsv.pymaps these to PyHiM schema:Chromosome index→Trace_ID(1..1277)Segment index→Barcode #(1..66)X, Y, Z(nm) →x, y, z(microns)Chrom = chr21,Chrom_Start/Endderived from segment index × 30 kbMask_id = Chromosome index(each trace = one “cell”)ROI # = 0,label = "None"
Used by:
import_pyhim_ecsv.ipynb— demonstratesChromData.from_pyhim_trace()on real chromatin tracing data;benchmarks/cd2_callers_validation.py(ChromData 2.0: structure callers before / after the calling convention).Generated on demand: If missing, the tutorial runs
bintu_to_pyhim_ecsv.pyto convertIMR90_chr21-18-20Mb.csv(which is auto-downloaded if needed).
4DNFIHF3JCBY.csv — Takei 2021 mESC FOF-CT chromatin tracing (22 MB)¶
Format: 4DN FISH Omics Format — Chromatin Tracing (FOF-CT) core table. Headers (
##…) describe the experiment; the data table has columns:Spot_ID, Trace_ID, X, Y, Z, Chrom, Chrom_Start, Chrom_End, Cell_ID, ….Content: mm10, 20 chromosomes × 60 bins × 25 kb, ~400 traces per chromosome across 201 E14 mESC cells.
Source: Takei et al. 2021, Nature 590:344–350, “Integrated spatial genomics reveals global architecture of single nuclei”. 4DN portal: data.4dnucleome.org/4DNFIHF3JCBY.
Download URL (public S3, no credentials needed):
https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/e699334e-fb34-4a0e-8ef6-670b2099831a/4DNFIHF3JCBY.csv
Used by (all use the
find_fofct()helper):loop_calling.ipynb,tad_calling.ipynb,compartment.ipynb,fishnet_domains.ipynb;benchmarks/cd2_callers_validation.py(ChromData 2.0 validation, all 20 chromosomes);tests/core/test_fofct_writer.pyandbenchmarks/fofct_roundtrip_validation.py(FOF-CT writer) (with the cell / RNA tables4DNFIFINA2U9.csv,4DNFIJ52NVDV.csv).
4DNFIFLJGGNR.csv — Takei 2021 mESC FOF-CT, 1 Mb genome-wide loci (47 MB)¶
Format: FOF-CT core table (same columns as
4DNFIHF3JCBY.csv),##XYZ_unit=micron, mm10.Content: E14 mESC, replicate 1: 2,460 loci (25 kb each) at ~1 Mb spacing across chr1–19 and chrX, 201 cells, 705,143 spots. One
Trace_IDper (cell, chromosome) holds the spots of both homologs (not resolved in the deposited table; ~1.8 spots per locus).benchmarks/screcon/truth.pyseparates the homologs with a constrained 2-means (heuristic; checked on the homolog-resolved 25 kb table: median 97.8 % of spots on the right homolog) and keeps chrX (male line) as a single copy.Source: Takei et al. 2021, Nature 590:344–350 (same study as
4DNFIHF3JCBY.csv). 4DN portal: data.4dnucleome.org/4DNFIFLJGGNR, “DNA-Spot/Trace Data core table for the 2460 genomic loci across the mouse genome at 1-Mb resolution in ESCs, replicate 1”; md5da9826551f13f944815fe35a33a8e9c1(verified).Download URL (public S3):
https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/8189709b-f1e7-4f0b-bf68-89eb9fd0afcf/4DNFIFLJGGNR.csv
or
python example-data/download_data.py --takei2021_1mb.Used by:
benchmarks/screcon/simulate.py truth --dataset takei2021_1mb(imaging-truth simulation benchmark for single-cell Hi-C 3-D reconstruction);benchmarks/screcon/matched_data.py locus-map(recovers the mm10 coordinates of the 2,460 locus names, see “Matched-modality datasets”).
liu2025_mop/ — Liu et al. 2025 mouse MOp DNA-MERFISH, 4DN FOF-CT (1.7 GB)¶
Source: Liu S, … Zhuang X. “Cell type-specific 3D-genome organization and transcription regulation in the brain.” Sci Adv (2025). 4DN experiment set 4DNESMTNNB3N, wild-type mouse primary motor cortex.
core table 4DNFID46OABK (1,684,734,602 B):
https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/7c153998-3e92-4dbe-8e35-7a09570e6ed3/4DNFID46OABK.csvcell table 4DNFICX8IVEK (16,564,294 B):
https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/ad72715d-57cc-4d68-8ddb-536b94ed9bb0/4DNFICX8IVEK.csv
Content: 9,582,069 spots, 188,891 traces, 11,045 traced cells (14,733 rows in the cell table, with 240 RNA counts and cell types). The download helper and build script (
download_data.py --liu2025_mop,build_liu2025_mop.py) come with PR #64.Used by:
benchmarks/fofct_roundtrip_validation.py(FOF-CT writer).
IMR90_chr21-18-20Mb.csv — Bintu 2018 IMR90 chr21:18.6–20.6 Mb chromatin tracing (2 MB)¶
Format: CSV with header line then columns
Chromosome index, Segment index, Z, X, Y. One row per detected segment per imaged chromosome. Coordinates in nanometres. Segment spacing is 30 kb.Content: 1 278 imaged chromosomes × 66 segments in IMR90 cells, covering chr21:18,627,714–20,577,518 (hg38).
Source: Bintu et al. 2018, Science 362, eaau1783, “Super- resolution chromatin tracing reveals domains and cooperative interactions in single cells”. The paper is a higher-resolution follow-up to Wang et al. 2016 (Science 353:598) which Abbas et al. 2019 (GEM-FISH) originally used; Bintu 2018 data is directly accessible from the authors’ GitHub repository.
Download URL:
https://raw.githubusercontent.com/BogdanBintu/ChromatinImaging/master/Data/IMR90_chr21-18-20Mb.csvUsed by:
gem_fish_reconstruction.ipynb(part 2 — real FISH data), paired with the real Rao 2014 IMR90 Hi-Crao2014_imr90/IMR90_chr21_5kb_hg19.coolin hg19 (Bintu’s hg19 coordinates: chr21:20,000,032–21,949,831; the tutorial sums the 5 kb Hi-C into 30 kb bins on the Bintu segment grid). Result on this real pair (2026-09-27): Pearson 0.71–0.73 and mean relative error 0.33–0.35 against the FISH median distances (the same FISH also constrains the model). Earlier versions paired it with the K562K562_chr21_30kb.cool(then misnamed IMR90): Pearson 0.77, relative error 0.35. Also Fig. 3b / 3c:benchmarks/fig3/b_first_look.py,c_gemfish_first_look.py.
GSE63525_GM12878_insitu_primary+replicate_combined_30.hic — GM12878 in-situ combined Hi-C (~40 GB)¶
Format: Juicer
.hic— a multi-resolution contact matrix file (read withhicstraw).Content: GM12878 lymphoblastoid cells, in-situ Hi-C, primary + replicate merged and MAPQ ≥ 30 filtered. Contains every standard Juicer resolution from 1 kb to 2.5 Mb, all chromosomes.
Source: Rao et al. 2014, Cell 159:1665–1680, “A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping”.
Download: not auto-fetched (40 GB is too large to ship through
download_data.py). Grab it manually from one of:GEO GSE63525 (look for
GSE63525_GM12878_insitu_primary+replicate_combined_30.hic)
Used by:
benchmark.ipynb(MDS reconstruction benchmark). The tutorial checksexample-data/first and falls back to theUCHROM_GM12878_HICenv var, so you can keep the file on an external drive:export UCHROM_GM12878_HIC=/path/to/…_combined_30.hic.
DNAseqFISH+.zip — Takei 2021 raw seqFISH+ spots (144 MB)¶
Format: zip containing 8 CSVs — 4 replicates at 1-Mb resolution and 4 at 25-kb resolution. Columns:
fov, channel, cellID, regionID (hyb1-60), x, y, z, dot_intensity, chr{N}_intensity × 20, chromID, labelID.Content: same experiment as the FOF-CT above, but before trace assignment — every row is a detected fluorescent spot with a decoded chromosome ID but ambiguous fiber assignment (median 6 candidate spots per
(cell, chromID, region)).labelID ≥ 0marks the upstream pipeline’s trace choice (useful as ground truth when benchmarking aligners).Source: Takei et al. 2021, Zenodo record 3735329, doi:10.5281/zenodo.3735329.
Download URL:
https://zenodo.org/records/3735329/files/DNAseqFISH%2B.zip?download=1
Used by:
jie_aligner.ipynb(spot-to-fiber tracing). The tutorial opens the CSV directly from the zip without extracting.benchmarks/screcon/matched_data.py takei-mescreads the 1 Mb tables of replicates 1 + 2 (201 + 245 = 446 E14 cells) as the mESC imaging reference of the matched-modality comparison;locus-mapjoins replicate 1 with4DNFIFLJGGNR.csv.Coordinate conventions (applied by the tutorial):
x, yin pixels × 103 nm/pixel;zin pixels × 250 nm/pixel.
H1Esc-HFF.R1.tar.gz + H1Esc-HFF.R1.labeled — Kim 2020 sci-Hi-C (128 MB + 92 KB)¶
Format: tarball of per-cell
.matrixfiles; each is a sparse tripletbin1<TAB>bin2<TAB>count<TAB>weight<TAB>chrom1<TAB>chrom2, withbin1/bin2as global bin indices across the whole hg19 genome at 500 kb (offsets are not encoded — derive them by min-bin per chrom across the cells, or compute from canonical hg19 chromsizes). Chromosome strings are prefixedhuman_(e.g.human_chr14). The companion*.labeledis a 2-column TSVmatrix_filename<TAB>cell_typewith values in{H1Esc, HFF}.Content: 1 931 cells (750 H1Esc + 1 181 HFF), pooled from a combinatorial-indexing sci-Hi-C library at 500 kb. Used as a benchmark in Kim et al.’s topic-model paper.
Source: Kim et al. 2020, Nature Communications 11:6386, “Capturing cell type-specific chromatin compartment patterns by applying topic modeling to single-cell Hi-C data” — accompanying website at noble.gs.washington.edu/proj/schic-topic-model.
Download URLs:
https://noble.gs.washington.edu/proj/schic-topic-model/data/matrix_files/H1Esc-HFF.R1.tar.gz https://noble.gs.washington.edu/proj/schic-topic-model/data/matrix_labels/H1Esc-HFF.R1.labeled
Used by:
higashi_embedding.ipynb(FastHigashi cell embedding + ARI vs ground-truth labels). The tutorial picks a balanced subset (e.g. 150 H1Esc + 150 HFF), converts each.matrixto Higashi v2 contact-pair format, runs FastHigashi at rank 64 withdo_conv/do_rwr/do_col=True, and reports ARI / NMI between the k-means clustering ofcd.cellm['higashi']and the cell-type labels. Reproduces ARI ≈ 0.55 on a Mac mini M2 in ~2 min on CPU at 300 cells.
scHiCAR mouse brain + MERFISH MOp (schicar_mop/)¶
Used by schicar_to_merfish_mapping.ipynb,
schicar_merfish_uchrom_framework_walkthrough.ipynb and (paths only,
no data needed) schicar_linked_multiomics_framework.ipynb. The
notebooks read $UCHROM_SCHICAR_ROOT, defaulting to
example-data/schicar_mop/.
Upstream sources (public, auto-downloadable):
scHiCAR — Wei X, Xu Y, Yang D, … Diao Y. “Trimodal single-cell profiling of transcriptome, epigenome and 3D genome in complex tissues with scHiCAR.” Nature Biotechnology (2026), doi:10.1038/s41587-026-03013-7. GEO GSE305439 (
mouse_brain_scHiCAR_1, mouse frontal cortex; SubSeries of GSE305889, BioProject PRJNA1305748). Files underhttps://ftp.ncbi.nlm.nih.gov/geo/series/GSE305nnn/GSE305439/suppl/:File
Size
Content
GSE305439_RNA.matrix.mtx.gz/.barcodes.tsv.gz/.features.tsv.gz61 MB
RNA counts (10x-style MTX)
GSE305439_mouse_brain_scHiCAR_1_RNA_metadata.txt.gz152 KB
5 313 cells:
RNAbarcode, nCount_RNA, nFeature_RNA, celltype, UMAP1, UMAP2, DNAbarcodeGSE305439_DNA.dedup.pairs.gz1.47 GB
deduplicated contact pairs (keyed by DNA barcode)
GSE305439_DNA.ATAC.tsv.gz1.35 GB
ATAC fragments
chrom, start, end, DNA barcode, 2-nt tag, strand(175,496,023 fragments, all from the 5,313 metadata cells)python example-data/download_data.py --schicar_rnafetches the RNA side;--schicar_dna(opt-in, 1.47 GB) the pairs;--schicar_atac(opt-in, 1.36 GB) the ATAC fragments plus the gene annotation below.UCSC mm10 refGene (NCBI RefSeq genes on mm10, UCSC Genome Browser; Navarro Gonzalez et al. 2021, Nucleic Acids Res 49:D1046) —
https://hgdownload.soe.ucsc.edu/goldenPath/mm10/bigZips/genes/mm10.refGene.gtf.gz(13 MB, GTF) →schicar_mop/raw/mm10.refGene.gtf.gz; gene body + 2 kb upstream windows for ATAC gene activity inbuild_schicar_mop.py. Fetched by--schicar_atac. The larger sibling subseries (e.g. GSE267126) are multi-TB and are not used.MERFISH MOp atlas — Zhang M, Eichhorn SW, Zingg B, … Zhuang X. “Spatially resolved cell atlas of the mouse primary motor cortex by MERFISH.” Nature 598:137–143 (2021), doi:10.1038/s41586-021-03705-x. Brain Image Library, doi:10.35077/g.21; processed files under
https://download.brainimagelibrary.org/cf/1c/cf1c1a431ef8d021/processed_data/:counts.h5ad(305 MB, all 12 experiments, volume-normalised cell × 258 genes) andcell_labels.csv(27 MB,sample_id, slice_id, class_label, subclass, label). The tutorials use experimentmouse2_sample1. Fetch with--merfish_mop.
Derived inputs — recipe pending. The notebooks do not read the raw files directly; they read intermediates that were produced outside this repository and whose build scripts are not yet checked in:
Derived file (relative to |
Derived from |
|---|---|
|
GSE305439 RNA + metadata (MOp-matched cells) |
|
GSE305439 pairs → per-cell contact summary features |
|
GSE305439 ATAC → per-cell 1 Mb bins |
|
scHiCAR paper 5 kb loop calls |
|
BIL |
|
all of the above + Tangram mapping |
Until the build scripts land (tracked as an open item; see the tutorials’ first cell), these two notebooks are not reproducible from public data alone and are shipped with their recorded outputs.
sim_cell1/ — NucDynamics single-cell structure + its contacts (3.2 MB, in repo)¶
Content:
sim_cell1.cdz(0.7 MB; a one-file.chromdata.zarrstore, converted from the former 1.3sim_cell1.h5cdof 1.6 MB withpython -m uchrom.io.upgrade) — NucDynamics (uchrom.recon.sc.nucdyn, CPU) oncell1.pairs(Stevens et al. 2017, mESC G1, mm10): 20 chromosomes, 25,654 particles at ~100 kb, one trace per chromosome, coordinates in model units;cell1_contacts.mcool— the same cell’s 105,700 contacts at 100 kb / 200 kb / 500 kb / 1 Mb, linked withcd.link_cool("cell1_contacts.mcool")(path relative to the store).Known issue: the structure was computed before the
uchrom.io.count_contactsaxis fix (paper/fig3/PLAN.md, finding 5), so its restraints joined scrambled particle pairs. The native engine (native/nucdyn, validated against the original on the 8 Stevens cells) can rebuild it (build_sim_cell1.py, now using that engine), but the rebuild is deferred: the Fig. 6 agent benchmark’s ground truth (benchmarks/fig6/agent_bench/, questions on bond length and contacts vs distance) was computed on this file. Until then treat it as a browser / benchmark demo, not as a valid Cell 1 structure. The mcool is built directly fromcell1.pairsand is not affected.Rebuild (native engine):
pip install ./native/nucdyn, thenpython example-data/build_sim_cell1.py [--device auto|cpu|gpu] [--n-models 4] [--seed 1] [--format cdz|zarr|h5cd](then regenerate the Fig. 6 agent-benchmark answers).Check: bins in contact in the input are ~2× closer in the model (median 3D distance of contacted / non-contacted 1 Mb pairs = 0.44–0.59 on chr1, 2, 5, 11, 19) —
tests/browser_web/test_genome_views.py.Used by: web browser (3D + contact + distance matrices).
takei2025_fov0/ — Takei 2025 cerebellum, rep 1 FOV 0, first 47 cells (73 MB, built)¶
Content: DNA seqFISH+ traces (383,674 spots, 1,500 traces, 47 cells: Granule, Bergmann, Purkinje, MLI1, MLI2+PLI, Other) with 59 per-spot immunofluorescence / RNA-FISH channels as
tracks(H3K27ac, H3K4me3, H3K27me3, H3K9me3, LaminB1, …), cell types and the published UMAP (cellm['umap']); per-cell means of the 48 IF channels ascells['if.<mark>']and a histone / IF embeddingif_pca / _tsne / _umapfrom the within-cell mark–mark correlations (build_takei2025_fov0.py,--updateredoes only that step).Source: Takei et al. 2025, Nature (doi:10.1038/s41586-025-08838-x); Zenodo 7693825
cerebellum_rep1.tar.gz(2.8 GB) — only the start of the first FOV is read (streamed); locus map and clustering from CaiGroup/dna-seqfish-plus-multi-omics.Build:
python example-data/build_takei2025_fov0.py [--rows 400000](~1 min, downloads ~0.1–0.3 GB of the tarball, drops the last possibly truncated cell) →takei2025_fov0.chromdata.zarr(73 MB; the 1.x.h5cdof older builds was 220 MB).Used by: web browser (genome track view, distance matrices, stats).
schicar_mop/ — scHiCAR mouse brain, no 3-D coordinates (built from GEO)¶
Content (built by
python example-data/build_schicar_mop.pyfrom the GSE305439 files above):schicar_mop.chromdata.zarr(5.4 MB; the 1.x.h5cdof older builds was 44 MB) —ChromDatawithout spots: 5,313 cells (RNA barcode ids, 22 cell types, QC), 500 highly variable genes asrna.<gene>(log-normalised dispersion, genes in ≥1 % cells), ATAC gene activity of 484 of them asatac.<gene>(fragment midpoints in gene body + 2 kb upstream, UCSC mm10 refGene),n_fragments_atac,n_contacts_hic;cellm:paper_umap(from the GEO metadata),rna_pca / _tsne / _umap(library-size normalised, 30 PCs),atac_lsi / _tsne / _umap(100 kb bins, TF-IDF + LSI, LSI 1 dropped: r = 0.97 with log depth) andhic_pca / _tsne / _umap(scHiCluster on the 1 Mb scool) — all fromuchrom.emb.embed_cells. The ATAC here is weakly peaked (TSS enrichment ≈ 2.9 on the first 5 M fragments), so 100 kb bins beat 5–50 kb ones for cell-type structure (kNN-10 purity 0.33 vs 0.19–0.30);python benchmarks/omics_embeddings.pygives the per-modality numbers.--updateredoes only the ATAC / Hi-C step;contacts_all.mcool(42 MB) — all 120,280,303 contacts, 250 kb … 2 Mb;contacts_<celltype>.cool— 22 pseudo-bulk maps at 500 kb;contacts_cells_1Mb.scool(257 MB) — one map per cell (all 5,313). Links are stored relative to the dataset folder.
Used by: web browser (matrix, embedding and field views for data without 3-D).
takei2025_cerebellum/ — Takei 2025 cerebellum, replicate 1, all FOVs (Fig. 2 benchmarks)¶
Source: Takei et al. 2025, Nature (doi:10.1038/s41586-025-08838-x); Zenodo 7693825 (CC-BY-4.0),
cerebellum_rep1.tar.gz(2,846,258,062 B, md5758ebb15…4d84471d62b4, checked bydownload_data.py). Locus map and clustering from CaiGroup/dna-seqfish-plus-multi-omics (same URLs asbuild_takei2025_fov0.py). The record’s other files:cerebellum_rep2.tar.gz(5.28 GB),E14_rep1/2.tar.gz(11.1 / 11.3 GB),NMuMG.tar.gz(4.16 GB),transcriptomic-data.zip(0.46 GB), a mean-IF CSV andreadme.txt; rep 2 exists but is not used here.Fetch / build:
python example-data/download_data.py --takei2025_cerebellum python example-data/build_takei2025_cerebellum.py # ~3 min on an M5 Pro
The tarball holds 3 FOVs (
cerebellum_rep1_pos{0,1,2}.csv, 2.23 / 2.77 / 2.87 GB; plus macOS._*entries, skipped). Each FOV is loaded in its own subprocess withread_seqfish_multiomics(defaultdbscan_ldp_nbr_alleletraces, DBSCAN noise dropped, 62 spot tracks) →fov/takei2025_cerebellum_rep1_pos<N>.chromdata.zarr, then concatenated intotakei2025_cerebellum_rep1.chromdata.zarr(the format 2.0 default;--format h5cdwrites the deprecated HDF5 container); stats inbuild_summary.json. Per-FOV files from older builds (fov/*.h5cd, 1.x) are reused as they are.spots
traces
cells
h5cd 1.x
loader peak RSS
FOV 0
3,108,941
16,609
545
1.79 GB
10.6 GB
FOV 1
3,821,962
20,997
592
2.19 GB
12.2 GB
FOV 2
3,981,735
21,506
662
2.29 GB
13.1 GB
combined
10,912,638
59,112
1,799
6.29 GB
13.2 GB (concat)
The combined
.chromdata.zarris 2.17 GB (2.0, Zarr + Parquet; written in 55 s from the per-FOV files, peak RSS 26.5 GB); read withoriginal_order=Trueit equals the 1.x combined file value for value (coords, loci, trace / cell ids, 62 tracks, cells, cellm).The raw CSVs stay in
raw/(7.9 GB;benchmarks/fig2/allele_qc.pyreads their three allele columns);--drop-rawdeletes them.Caveat — allele calls: the default
dbscan_ldp_nbr_alleletraces often merge both homologs / add off-trace spots; seebenchmarks/fig2/allele_qc.pyandpaper/fig2/results/allele_qc.json.Used by:
benchmarks/fig2/b_storage.py,c_access.py(whole-cell subsets, 1e4–1e7 spots; 1e8 by replication,replicate_store.py),allele_qc.py;benchmarks/cd2_zarr_validation.py(→.chromdata.zarr, value identity).
liu2025_mop/ — Liu et al. 2025 mouse MOp DNA-MERFISH, 4DN FOF-CT (1.7 GB)¶
Source: Liu S, Wang CY, Zheng P, … Zhuang X. “Cell type-specific 3D-genome organization and transcription regulation in the brain.” Sci Adv (2025), PMID 40009678. 4DN experiment set 4DNESMTNNB3N (DNA-MERFISH of 2,000 loci + RNA-MERFISH of 240 genes, mouse primary motor cortex, wild type), experiment 4DNEXBUVZYLI.
core table 4DNFID46OABK (1,684,734,602 B, md5
e2234c8d…75a4abb):https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/7c153998-3e92-4dbe-8e35-7a09570e6ed3/4DNFID46OABK.csvcell table 4DNFICX8IVEK (16,564,294 B, md5
f66afc5a…e5e5):https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/ad72715d-57cc-4d68-8ddb-536b94ed9bb0/4DNFICX8IVEK.csv
How it was chosen: a portal search for released files of type
FOF-CT - DNA-spot/trace core(321 files, all open on the public S3 bucket) sorted by size. The top 10 (1.27–2.14 GB) are all Zhuang-lab mouse MOp / visual-cortex tables; the larger ones are Mecp2 KO samples, so we took the largest wild-type core table. It is 77× the Takei 2021 table (4DNFIHF3JCBY, 22 MB). Next largest from other studies: Su et al. 2020 IMR90 DNA-MERFISH (4DNFIGJT3AN3572 MB,4DNFIZ4TAGXZ248 MB).Content (
python example-data/build_liu2025_mop.py→liu2025_mop.chromdata.zarr294 MB +build_summary.json;--format h5cd→ 1.57 GB as h5cd 1.x): 9,582,069 spots, 188,891 traces, 11,045 cells, 20 chromosomes, median 50 spots per trace, µm, GRCm38. The cell table has 14,733 rows (3,695 cells without traces; 7 traced cells without a row), with 240 RNA counts and cell-type labels.from_fofctreads the 1.68 GB core table in 14 s at 11.6 GB peak RSS.Used by:
build_liu2025_mop.py(dataset-scale entry for Fig. 2);benchmarks/fofct_roundtrip_validation.py(FOF-CT writer);benchmarks/cd2_zarr_validation.py(→.chromdata.zarr).
Figure 3 datasets — bridges between contacts and geometry¶
Inputs of benchmarks/fig3/ (plan and reference numbers:
paper/fig3/PLAN.md). All opt-in: python example-data/download_data.py --fig3 fetches the downloadable ones (~370 MB), and
python example-data/build_rao2014_slices.py builds the Hi-C slices.
Everything is gitignored. DOIs checked against Crossref (2026-09-27).
stevens2017_mesc/ — Stevens 2017 haploid mESC single-cell Hi-C + published structures (43 MB)¶
Source: Stevens TJ et al. “3D structures of individual mammalian genomes studied by single-cell Hi-C.” Nature 544:59–64 (2017), doi:10.1038/nature21429. GEO GSE80280, Cells 1–8 = GSM2219497–GSM2219504; open access.
Files (per cell,
https://ftp.ncbi.nlm.nih.gov/geo/samples/GSM2219nnn/<GSM>/suppl/):<GSM>_Cell_<n>_contact_pairs.txt.gz(0.3–1 MB; tab-separatedchrA posA chrB posB, no header; 31,507–111,838 contacts) and<GSM>_Cell_<n>_genome_structure_model.pdb.gz(~4.7 MB; the paper’s 10 final NucDynamics models at 100 kb, coordinates in particle radii). Not fetched:Cell-<n>.hdf5.gz(~2.6 GB each) andGSE80280_RAW.tar(20 GB). GEO publishes no checksums; checked by parsing (8 × 10 models).Content: haploid 129/Ola mESC in G1, mm10.
Download:
download_data.py --stevens2017.Used by: Fig. 3a — NucDynamics reference structures and the paper’s precision metric (
benchmarks/fig3/a_sc_consistency.py); the native NucDynamics engine’s validation on all 8 cells (ensemble RMSD vs ED Table 2, RMSD / distance-matrix r vs these models, restraint violations:a_nucdyn_engine_run.py,a_nucdyn_engine_score.py) and its speed benchmarks against the original code (a_nucdyn_engine_speed.py,a_nucdyn_engine_original_timing.py,a_nucdyn_engine_original_run.py; Cell 1).
bintu2018/ — Bintu 2018 chr21:28–30 Mb tracing, IMR90 + K562 (28 MB)¶
Source: Bintu B et al. “Super-resolution chromatin tracing reveals domains and cooperative interactions in single cells.” Science 362:eaau1783 (2018), doi:10.1126/science.aau1783. GitHub BogdanBintu/ChromatinImaging
Data/(no licence file in the repo; data published with the paper).Files:
IMR90_chr21-28-30Mb.csv(7,795,158 B; 4,871 chromosomes) andK562_chr21-28-30Mb.csv(20,080,720 B; 13,996 chromosomes); sizes match the GitHub API. Same CSV layout asIMR90_chr21-18-20Mb.csv(Chromosome index, Segment index, Z, X, Y, nm; 65 × 30 kb segments).Coordinates: hg38 chr21:28,000,071–29,949,939; hg19 chr21:29,372,390–31,322,257 (from the repo’s
Data/README.md).Download:
download_data.py --bintu2018.Used by: Fig. 3b (
benchmarks/fig3/b_first_look.py): imaging contact frequency vs Rao 2014 Hi-C of the same cell line, and the cross-cell-line controls.
su2020_imr90/ — Su 2020 IMR90 chr21 + genome-scale tracing, binned Hi-C (465 MB)¶
Source: Su J-H et al. “Genome-Scale Imaging of the 3D Organization and Transcriptional Activity of Chromatin.” Cell 182:1641–1659 (2020), doi:10.1016/j.cell.2020.07.032. Zenodo 3928890 (doi:10.5281/zenodo.3928890), CC-BY-4.0.
Files (md5 from Zenodo, verified):
chromosome21.tsv(260,629,499 B,170d9d8b…; columnsZ(nm) X(nm) Y(nm),Genomic coordinate,Chromosome copy number, genes / transcription / TSS),Hi-C_contacts_chromosome21.tsv(1,045,862 B,7842ef31…; the authors’ Rao 2014 IMR90 reads summed into the 651 imaged bins), andREADME_August_2020.txt; genome scale (--su2020_genome):genomic-scale.tsv(200,285,524 B, md5a1d79c2b…, verified) andHi-C_contacts_genome-scale.tsv(2,846,153 B, md513e3607e…, verified; checked in, CC-BY-4.0, < 10 MB).Content: IMR90, hg38, 651 × 50 kb loci across chr21 (10.4–46.7 Mb), 7,591 traced chromosome copies. Genome scale (DNA-MERFISH): 1,041 loci of 100 kb (
chr1:2950000-3050000, ~3 Mb spacing) on chr1–22 and chrX, 1,787 cells from 3 experiments (431 / 917 / 439), each row labelled with its homolog (1 / 2) — 3,720,534 rows = 1,787 × 2 × 1,041, 16.2 % of positions missing — plus the distance of each spot to the nuclear lamina; columnsZ(nm) x(nm) y(nm) genomic coordinate, homolog number, cell number, experiment number, distance to lamina (nm).Hi-C_contacts_genome-scale.tsv: 1,041 × 1,041 Rao 2014 IMR90 read counts summed into 500 kb bins centred on the loci. The other Zenodo files (chr2, replicates with transcription / nuclear bodies, α-amanitin; 0.08–0.65 GB each) are not fetched.Download:
download_data.py --su2020_chr21/--su2020_genome.Used by: Fig. 3b (
benchmarks/fig3/b_first_look.py), whole-chromosome imaging vs Hi-C;benchmarks/screcon/(chr21 truth cells);benchmarks/bulk/(population deconvolution benchmark: genome-scale = primary diploid truth, chr21 = single-chromosome truth population;truths.py,forward.py; the genome-scale Hi-C sums are the cross-check of the IMR90 input,real_inputs.py su_crosscheck).
abbas2019_gemfish/ — GEM-FISH repository data (40 MB)¶
Source: Abbas A et al. “Integrating Hi-C and FISH data for modeling of the 3D organization of chromosomes.” Nat Commun 10:2049 (2019), doi:10.1038/s41467-019-10005-6. GitHub ahmedabbas81/GEM-FISH, MIT licence, pinned at commit
e83fdb4(2019-01-02). The FISH is Wang S et al. Science 353:598–602 (2016), doi:10.1126/science.aaf8084; the Hi-C is Rao 2014 IMR90 (GSE63525).Files (md5 verified):
complete_example.zip(21.8 MB; Wang 2016chr20/21/22.xlsx— per-cell TAD-centre coordinates in µm, 30 / 34 / 27 TADs — the hg19 TAD windowstads_chr2{0,1,2}_hg19.txt, chr20 Rao 2014 IMR90 5–100 kb RAWobserved + KR vectors),GEM-FISH_TAD-level-resolution.zip(0.2 MB; chr21 TAD-level input),GEM-FISH_TAD-conformations.zip(4.2 MB; chr21 per-TAD 5 kb Hi-C),validation_tests_final_models.zip(14 MB; the paper’s final 5 kb models of chr20/21/22 in nm, per-TAD Hi-C). No chrX data in the repo.Download:
download_data.py --abbas2019_gemfish.Used by: Fig. 3a / 3c (
benchmarks/fig3/c_gemfish_wang2016.py). The shipped final chr21 model scores a mean relative error of 0.143 against the shipped FISH, reproducing the paper’s Table 1 (0.14).
rao2014_imr90/, rao2014_k562/ — Rao 2014 Hi-C slices (built, 3.8 MB each)¶
Source: Rao SSP et al. “A 3D Map of the Human Genome at Kilobase Resolution Reveals Principles of Chromatin Looping.” Cell 159:1665–1680 (2014), doi:10.1016/j.cell.2014.11.021. GEO GSE63525:
GSE63525_IMR90_combined.hic(13 GB) andGSE63525_K562_combined.hic, hg19, MAPQ > 0 — the files Bintu 2018 used.Recipe:
python example-data/build_rao2014_slices.py [--cell IMR90 K562] [--chroms 21] [--res 5000]reads only the needed blocks over HTTP range requests withhicstraw(~30 s per chromosome, < 2 GB RAM) and writes raw observed counts torao2014_<cell>/<CELL>_chr21_5kb_hg19.cool(chr21: 10.27 M IMR90 / 9.54 M K562 contacts). For Fig. 3c chr20 / chr22--cell IMR90 --chroms 20and--chroms 22(one file each) writeIMR90_chr20_5kb_hg19.cool(20.78 M contacts, 7.8 MB, 59 s) andIMR90_chr22_5kb_hg19.cool(11.68 M contacts, 3.3 MB, 26 s). Other chromosomes, resolutions and GM12878 work the same way (--cell GM12878 --chroms 22 --res 100000for the miniMDS benchmark).Checked: the IMR90 chr21:28–30 Mb block reproduces Bintu 2018 Fig. 1J (median distance vs Hi-C, ρ = −0.96; here −0.975).
In the repo:
rao2014_imr90/IMR90_chr21_5kb_hg19.cool(3,815,649 B, md5f731fc98d5e4ae81936a066e46afa3fc; 9,626 bins, 2,745,237 pixels, 10,266,877 contacts) is checked in (< 10 MB) because the GEM-FISH tutorial needs it; the other slices stay gitignored.Used by: Fig. 3b / 3c (
benchmarks/fig3/), incl. the HIPPS / DIMES baselines (c_hipps_wang2016.py,b_dimes.py; their code is cloned, not vendored: github.com/anyuzx/HIPPS-DIMES, MIT, commitbeee45d); the IMR90 chr21 slice also bygem_fish_reconstruction.ipynbpart 2.
ray2019_k562/4DNFI4QQPDMR.mcool — K562 in situ Hi-C (849 MB, opt-in)¶
Source: Ray J et al. “Chromatin conformation remains stable upon extensive transcriptional changes driven by heat shock.” PNAS 116:19431 (2019), doi:10.1073/pnas.1901244116; 4DN 4DNFI4QQPDMR (set 4DNESU95RUNO, GEO GSE130758), GRCh38, md5
84f4e708e2e7b55078b670ed4b8db709(verified). S3:https://4dn-open-data-public.s3.amazonaws.com/fourfront-webprod/wfoutput/1c856462-1e76-4850-bfa6-defcf1524d42/4DNFI4QQPDMR.mcool.Download:
download_data.py --ray2019_k562.Used by: provenance check only —
K562_chr21_30kb.coolequals its chr21 summed to 30 kb (see that section).
Large Fig. 3 sources — documented, not downloaded (slice instead)¶
Dataset |
Accession / URL |
Size |
Build |
Slice recipe |
|---|---|---|---|---|
Rao 2014 GM12878 in situ (miniMDS benchmark) |
GEO GSE63525 |
51 GB ( |
hg19 |
|
Rao 2014 IMR90 in situ, 4DN |
set 4DNES1ZEJNRU: 4DNFIR1JDZH7 (4,639,622,711 B, md5 |
4.6–8.3 GB |
GRCh38 |
range-read one resolution: |
Bonev 2017 mESC Hi-C (Takei 2021’s comparison) |
GEO GSE96107; 4DN set 4DNESDXUWBD9 4DNFIC21MG3U (12,551,353,248 B, md5 |
11–13 GB |
mm10 |
same fsspec range read, the 20 Takei 2021 loci × 25 kb; smaller unsorted-ES mcool 4DNFIDA2WGV8 (0.9 GB); whole file on Sherlock: |
Stevens 2017 HDF5 + raw |
GSE80280 |
2.6 GB each / 20 GB |
mm10 |
not needed (contacts + |
Matched-modality datasets — scHi-C structures vs chromatin tracing of the same cell type¶
Used by benchmarks/screcon/matched.py (inputs built by benchmarks/screcon/matched_data.py;
jobs benchmarks/screcon/sherlock/{download,matched}_*.sbatch). All opt-in, all gitignored;
facts below were checked on the downloaded files unless marked unverified.
Locus axis — benchmarks/screcon/data/takei2021_1mb_loci.csv (78 KB, in repo, derived)¶
2,460 Takei 2021 ~1 Mb loci (25 kb probes, mm10):
name(chrN-#kfor the 1,267 channel-1 loci, a nearby gene for the 1,193 channel-2 loci),chrom,start,end.The raw Zenodo tables (3735329, 4708112) give locus names only; the coordinates are in the papers’ Table S1, which is not on Zenodo. Recovered instead from public files: the 4DN FOF-CT table
4DNFIFLJGGNR.csv(coordinates, no names) holds exactly the 705,143 spots ofDNAseqFISH+1Mbloci-E14-replicate1.csvinDNAseqFISH+.zip(201 cells; X = x·0.103, Y = y·0.103, Z = z·0.25 µm); joining on (cell, rounded X/Y/Z) matches 683,808 spots and gives every name exactly one coordinate (no ambiguity). Rebuild:python benchmarks/screcon/matched_data.py locus-map --fofct 4DNFIFLJGGNR.csv --zip DNAseqFISH+.zip. The brain Table S7 uses the same 2,460 names (all present).
benchmarks/screcon/data/tan2021_adult_cortex_cells.tsv (31 KB, in repo)¶
510 rows (gsm, name, age, cross, cell_type): the Tan 2021 adult cortex cells that
passed the authors’ filter (GSE146397_metadata.cells_contacts_100k.txt.gz, ≥ 100 k
contacts), matched to their GSM by the GEO sample records (!Sample_description).
nagano2017_hap/ — Nagano 2017 haploid mESC single-cell Hi-C (1.39 GB + 253 MB extracted)¶
Source: Nagano T, Lubling Y et al. “Cell-cycle dynamics of chromosomal organization at single-cell resolution.” Nature 547:61–67 (2017), doi:10.1038/nature23001. GEO GSE94489 holds raw reads + feature tables only; the per-cell contact maps are the authors’ archives linked from github.com/tanaylab/schic2 (S3 bucket
schic2).Files:
schic_hap_serum_adj_files.tar.gz(882,240,359 B; 790 cellsNST.<n>/adj),schic_hap_2i_adj_files.tar.gz(502,155,700 B; 982 cellsNXT.<n>/adj) —adj= tab-separatedfend1 fend2 count(GATC fragment ends, MboII/DpnII);GSE94489_haploids_features_table.txt.gz(123,901 B; 1,772 cells:cond2i_all / 2i_G1 / Serum_G1/S,passed_qc,total_contacts, cell-cyclegroup, …);GSE94489_README.txt(barcodes);GATC.fends(253,322,243 B;fend chr coord, mm9) streamed out ofschic2_mm9_db.tar.gz(4,996,816,093 B; not kept);mm9ToMm10.over.chain.gz(UCSC, 535,855 B) — contacts are lifted to mm10.Content (verified): 1,772 haploid cells (2i 982, serum 790);
passed_qc= 1 for 1,247. Cell line per GEO: haploid ESC20 / H129-1 (ECACC 14040203), 129 background (GEO calls it “hybrid ESC cell line”; strain details unverified beyond the GEO text).Selection (stated rules,
sherlock/_matched_env.sh): serum =Serum_G1/S,passed_qc == 1, group G1 / early-S / late-S/G2 (no mitotic), ≥ 100 k contacts (403 cells); 2i = both 2i conditions, same QC, ≥ 50 k contacts, 200 drawn with seed 1. Unique fend pairs per selected cell after mm10 liftover: serum median 181,883 (100,509 – 694,393), 2i median 77,415 (50,855 – 384,505); the adj row count matches the feature table’stotal_contacts(±1, checked on 3 cells), < 0.01 % of pairs lost in the liftover. The cell-cyclegroupis the authors’ inference from contact profiles; cells in S / G2 carry replicated chromatin and are modelled as haploid anyway (the G1 subset is reported separately).Download:
download_data.py --nagano2017_hap(orsherlock/download_matched.sbatch PART=mesc).
takei2021_brain/ — Takei 2021 Science mouse cortex DNA seqFISH+ (597 MB)¶
Source: Takei Y et al. “Single-cell nuclear architecture across cell types in the mouse brain.” Science 374:586–594 (2021), doi:10.1126/science.abj1966; Zenodo 4708112. 6–7-week-old female C57BL/6J mice (bioRxiv 10.1101/2021.04.26.441547 methods); 3 biological replicates.
Files:
TableS7_brain_DNAseqFISH_1Mb_voxel_coordinates_2762cells.csv(349,019,829 B; 4,752,662 spots; voxel coordinates, 103 × 103 × 250 nm;cluster label,chromID,geneID= locus name, DBSCANlabelID,XistID),TableS8_..._25kb_...csv(246,648,783 B),TableS5_brain_RNA_profiles_2762cells.csv(497,312 B),TableS10-median-radial-score-per-celltype-1Mb-resolution.csv(441,276 B). Not fetched: the IF / DAPI / ncRNA zips (0.6–13.7 GB each).Content (verified): 2,762 cells;
cluster labelsizes 155, 58, 41, 53, 152, 90, 240, 78, 1,895 (= the paper’s Fig. 1I legend). Names (from Table S5 marker means: Pvalb, Vip, Ndnf, Sst, Mfge8 / Aldoc, Csf1r, Cldn5, Olig1 / Plp1, Slc17a7): 1 Pvalb, 2 Vip, 3 Ndnf, 4 Sst, 5 astrocyte, 6 microglia, 7 endothelial, 8 oligodendrocyte lineage, 9 excitatory. Median 1,678 1 Mb spots per cell (≈ 34 % of 2 × 2,460); DBSCAN gives two homolog clusters for 30 % of (cell, chromosome).Download:
download_data.py --takei2021_brain.
tan2021_cortex/ — Tan 2021 Dip-C, adult mouse cortex (510 cells, 13 GB)¶
Source: Tan L et al. “Changes in genome architecture and transcriptional dynamics progress independently of sensory experience during post-natal brain development.” Cell 184:741–758 (2021), doi:10.1016/j.cell.2020.12.032. GEO SuperSeries GSE162511; Dip-C SubSeries GSE146397. Not fetched:
GSE146397_RAW.tar(all samples, 183 GB) — per-GSM files only.Files (per cell,
https://ftp.ncbi.nlm.nih.gov/geo/samples/<GSMnnn>/<GSM>/suppl/<GSM>_<name>.*):contacts.pairs.txt.gz(hickit pairs, ~2.5 MB),impute.pairs.txt.gz(haplotype-imputed, ~3 MB),20k.{1..5}.clean.3dg.txt.gz(dip-c 3DG, 20 kb particles per haplotype,chrom(mat|pat) start x y z, contact-poor particles removed; ~3.5 MB each; the structures the paper used); plus the metadata tablesGSE146397_metadata.*and README.Content (verified from GEO records + README): cortex P56 (251 cells), P309 (131), P347 (128); P56 / P347 = CAST/EiJ ♀ × C57BL/6J ♂ (
cb), P309 = the reciprocal cross (bc); males (README: male processing for all ages except P1; sex per cell not in the GEO records). Structure types frommetadata.cells_contacts_100k: L2–5 pyramidal 160, oligodendrocyte 74, interneuron 56, L6 pyramidal 54, microglia 37, hippocampal pyramidal 31, astrocyte 30, medium spiny neuron 29, OPC 24, other 19. Contacts per cell (contacts.pairs, every 25th cell): 207 k – 571 k, median ~420 k (the ≥ 100 k filter is the authors’). The publishedclean.3dgcover a median 4,792 of the 2 × 2,460 locus copies. Unverified: per-cell sex (GEO has none), whether the cell-type labels came from the paper’s own clustering of these exact structures (the metadata file calls them “structure types”).Download:
download_data.py --tan2021_cortex(orsherlock/download_matched.sbatch PART=brain).
Liu 2025 DNA-MERFISH, mouse cortex (4DN) — catalogued, not downloaded¶
Liu S, … Zhuang X, Sci Adv (2025). 4DN experiment sets (sizes from the 4DN API, 2026-10-01):
4DNESMTNNB3N wild-type
MOp, 4 experiments, core FOF-CT tables 1.45–1.68 GB each (one, 4DNFID46OABK, is the
liu2025_mop/ entry above) plus 0.9–1.8 GB and 7–9 GB companion tables;
4DNESPE924IP Mecp2 +/−
MOp (4 experiments, 1.3–2.1 GB + 6–10 GB tables);
4DNESQU9R2NY Mecp2 +/−
visual cortex (2 experiments, 1.8 GB + 8–9 GB tables). ~2,000 loci genome-wide (different
from the Takei loci) — a second imaging reference would need its own locus axis; too large to
stage for a baseline.
Bulk Hi-C deconvolution benchmark (benchmarks/bulk/)¶
Population deconvolution of bulk Hi-C (uchrom.recon.bulk.deconv.deconvolve(method="igm");
design benchmarks/bulk/design.md section 6-7, commands benchmarks/bulk/README.md).
Imaging truths (catalogued above): Su 2020 genome-scale (--su2020_genome, primary
diploid genome-wide truth), Su 2020 chr21 (--su2020_chr21), Takei 2021 1 Mb
(--takei2021_1mb, homologs split heuristically by benchmarks/screcon/truth.split_homologs).
Real Hi-C (whole files only on Sherlock; laptop checks use HTTP range reads):
Rao 2014 IMR90, 4DN GRCh38 —
rao2014_imr90/4DNFIR1JDZH7.mcool(4,639,622,711 B, md5f91b77fa61acdc58360911c80b007646, 4DN set 4DNES1ZEJNRU; Rao SSP et al. Cell 159:1665 (2014), GEO GSE63525; 4DN data-use policy: open). GRCh38 like the Su 2020 loci, so no liftover. Levels 1 kb–10 Mb (no 20 kb): the IGM 20 kb matrix is summed from the 10 kb level.download_data.py --rao2014_imr90_4dn(Sherlock:benchmarks/bulk/sherlock/build_real.sbatch).Bonev 2017 mESC, 4DN mm10 —
bonev2017_mesc/4DNFIC21MG3U.mcool(12,551,353,248 B, md5dda4ed67a1179f917721fb81b1adebe0, 4DN set 4DNESDXUWBD9; Bonev B et al. Cell 171:557 (2017), doi:10.1016/j.cell.2017.09.043, GEO GSE96107).download_data.py --bonev2017_mesc_4dn(Sherlock).Rao 2014 IMR90, GEO hg19
GSE63525_IMR90_combined.hic(13 GB) — not downloaded; read over HTTP range requests bybenchmarks/bulk/kr_check.py(its Juicer KR vectors are the reference for U-Chrom’s KR balancing).
Derived inputs (Sherlock, $SCRATCH/uchrom-nd/runs/bulk/real/<cell>/, not stored;
recipe benchmarks/bulk/real_inputs.py preprocess, IGM SI section 6 protocol in
uchrom.recon.bulk.deconv.preprocess): fine_20kb.npz (raw 20 kb counts, chr1–22 + X),
igm_200kb.npz (whole genome, 200 kb, 24 contacts per bin), su_3mb.npz / su_1mb.npz
(IMR90 onto 3 Mb / 1 Mb bins centred on the Su loci), takei_1mb.npz (mESC onto 1 Mb bins
centred on the Takei loci), su_crosscheck.json.
Checked in: benchmarks/bulk/results/kr_check.json (KR check, range reads) and
benchmarks/bulk/results/plumbing_imr90_chr21_22_igm200kb.json (provenance of the
laptop plumbing run: 4DN IMR90 chr21 + chr22, range-read 10 kb level → 20 kb → 200 kb).
Generated outputs¶
with_loops.chromdata.zarr(older runs:with_loops.h5cd) — created at the end ofloop_calling.ipynbas a round-trip demonstration. Safe to delete; the tutorial will regenerate it on next run.fofct_core.csv— a small synthetic FOF-CT that tutorials 4–7 generate if the real Takei 2021 CSV can’t be downloaded. Has a hand-crafted 3-TAD + 1-loop structure purely to keep the notebook runnable offline.
Pre-staging all data (optional)¶
Tutorials download what they need on first run, so you normally don’t have to do anything. To pre-stage all large files (useful for offline / CI environments), run:
python example-data/download_data.py
python example-data/download_data.py --fig3 # Fig. 3 inputs (opt-in)
python example-data/download_data.py --su2020_genome --su2020_chr21 --takei2021_1mb # bulk benchmark truths
python example-data/build_rao2014_slices.py # Fig. 3 Hi-C slices (built)
--dest DIR stages into another directory (e.g. a shared data folder).
This is idempotent — files already present are skipped.
Adding a new dataset¶
If under ~10 MB and redistributable → commit to
example-data/and add a row to the “Small, in-repo datasets” table above.Otherwise → add a
download_*helper todownload_data.py, a row to “Large datasets”, and afind_*()function in whichever tutorial needs it (follow the existingfind_fofct()pattern).Always cite the paper + accession, and note the licence / terms of use where they matter.