Concepts¶
The structure table¶
The central abstraction in U-Chrom is the structure table: a flat mapping from genomic bins to 3D coordinates.
chrom |
start |
end |
x |
y |
z |
|---|---|---|---|---|---|
chr1 |
0 |
100_000 |
0.34 |
1.20 |
-0.55 |
chr1 |
100_000 |
200_000 |
0.38 |
1.15 |
-0.52 |
… |
… |
… |
… |
… |
… |
This is what both reconstruction output (Hi-C → 3D) and imaging output (chromatin tracing → 3D) produce. Every analysis module in U-Chrom consumes this format.
Hierarchy: Cell → Trace → Spot¶
A single row in the structure table is a Spot — one 3D observation of one genomic bin. Spots are organised hierarchically:
Spot — one observation of one bin in one trace.
Trace — an ordered polymer of spots along a chromatin fibre. One allele’s copy of the locus is one trace. In chromatin tracing, a diploid cell can give two traces per chromosome (maternal + paternal).
Cell — a physical cell containing one or more traces.
For reconstructed data this hierarchy is thin: each chromosome is typically a single trace, and the whole output is one cell.
For imaging data it matters: a single FOF-CT file may have hundreds of traces on the same chromosome, each coming from a different cell (and without explicit cell segmentation, individual traces do not necessarily belong to different cells — they just represent separate allele observations).
ChromData — the container¶
uchrom.ChromData holds the structure table plus hierarchical
metadata and analysis results:
ChromData
├── coords (n_spots, 3) x, y, z coordinates
├── spots DataFrame (n_spots) chrom, start, end, trace_id, [cell_id]
├── cells DataFrame (n_cells) cell-level metadata
├── cellm dict[str, ndarray] per-cell embeddings etc.
├── tracks DataFrame (n_spots) per-bin epigenomic signals (ATAC, ChIP)
├── traces DataFrame (n_traces) trace-level metadata
├── layers dict[str, (n_spots, 3)] alternative coordinate sets
├── results dict analysis outputs (loops, tads, ...)
└── uns dict genome_assembly, xyz_unit, ...
Key design choices are documented in User guide — ChromData.
Data flow¶
┌──────────────────────────────────┐
│ Hi-C / Dip-C (.pairs, .cool) │
└──────────────────────────────────┘
│
▼ uchrom.recon.{sc,bulk}
┌──────────────────────────────────┐
│ 3D coordinates (ChromData) │──┐
└──────────────────────────────────┘ │
│
┌──────────────────────────────────┐ │
│ Imaging (FOF-CT .csv) │──┼──▶ ChromData
└──────────────────────────────────┘ │
│ │
▼ uchrom.im (WIP) │
┌──────────────────────────────────┐ │
│ 3D coordinates (ChromData) │──┘
└──────────────────────────────────┘
│
▼
┌──────────────────────────────────┐
│ Downstream analysis │
│ - uchrom.strc.loop │
│ - uchrom.strc.tad │
│ - uchrom.strc.comp (WIP) │
│ - uchrom.fea │
│ - uchrom.emb (WIP) │
│ - uchrom.pl │
│ - uchrom.browser │
└──────────────────────────────────┘
On-disk format: .chromdata.zarr¶
Format 2.0 is a Zarr v3 directory whose tables are Parquet files (the
SpatialData approach); <name>.cdz is the same store as one zip file:
data.chromdata.zarr/
├── zarr.json @uchrom_format_version = "2.0", @uchrom_version
├── tables/ spots.parquet (sorted by cell, trace, bin; x, y, z),
│ spot_tracks.parquet, layers/, bins, bin_tracks,
│ cells, traces, intervals/, points/, categories/
├── index/ row offsets per trace / cell, chromosome per trace
├── cellm/, binm/ zarr arrays
├── results/ values + provenance attrs
├── uns/ JSON attrs (+ arrays)
└── contacts/, links/ linked .cool / .scool / .h5ad / .h5mu records
The index makes a cell, trace or chromosome a few contiguous row ranges,
so ChromData.read(path, backed=True) can open large data sets without
loading them. .h5cd (HDF5) files — format 1.x and 2.0 — are still read;
python -m uchrom.io.upgrade converts them.
Format version follows MAJOR.MINOR:
same MAJOR, higher MINOR → reader warns and proceeds (unknown fields are ignored);
different MAJOR → reader raises, with a clear upgrade path;
missing attribute → reader warns and assumes legacy 1.0 layout.
See uchrom/core/spec.md
for the complete on-disk spec.