Changelog¶
Unreleased¶
Cell spatial position is first-class (
.chromdata.zarrformat 2.3). Centroids arecellscolumnscentroid_x,centroid_y[,centroid_z] with a recorduns['cell_spatial'][key](unit, frame fov / tissue / global,region_col,in_coords_frame, status measured / inferred, source, optional voxel size):cd.set_cell_positions(),cd.cell_positions(key, physical=). Optional outlinescd.cell_shapes[key](cell_id+ WKB polygons) are stored as GeoParquet (tables/cell_shapes/<key>.parquet, readable bygeopandas.read_parquet);cd.set_cell_shapes()needs no geometry library,cd.cell_shapes_geodataframe()needs shapely + geopandas (new extrau-chrom[spatial]).cd.link_spatialdata()records an external SpatialData store and its cell map (region/instance_key↔cell_id), checked byvalidate_links();cd.load_linked_spatialdata(cells=),cd.to_spatialdata()(extrau-chrom[spatialdata]). Loaders:read_seqfish_multiomicswritescentroid_x/y/z(voxels, per FOV; wasx_centroid…) and afovcells column;from_fofct(cell_table=)registersCent_ROI_x/y[/z]orcell_center_x_global / _y_global;from_fofct(mapping_table=)/to_fofct(mapping_table=)read / write the Cell/ROI Mapping table (ROI_Boundaries, OME polygons).set_cell_spatial_coordinates()is deprecated and now writes the schema (centroid_x, wasspatial_x); older columns stay readable throughcell_positions(). Format 2.3 = the 2.2 layout + these additive parts: 2.2 readers read 2.3 stores with a warning; 2.0–2.2 stores read unchanged. Real data (benchmarks/cell_spatial_validation.py): Takei 2025 cerebellum 1,799 / 1,799 cells, Takei 2021 201 / 201, Liu 2025 MOp 14,733 / 14,733, all equal to the source tables..chromdata.zarrformat 2.2 — every spot value stored once. Coordinates + keys in a chromosome-partitioned table (tables/coords/chrom=<name>/), every other spot-aligned column in cell-sorted primary tables without key columns (tables/primary/), and spot columns that are a function of the bin / cell / trace (probe names, sample ids) once per key (tables/derived/; exact round trip). Float columns use Parquet dictionary pages; the lossless decimal codec (uchrom.core.floatcodec) is opt-in (float_encoding=). Parallel full reads. Takei 2025 (1e7 spots, 62 tracks): 1.81 GB (2.1: 2.58 GB; flat Parquet: 2.11 GB), full read 0.57 s (2.1: 0.81 s; Parquet 0.76 s),get_cell8.5 ms (2.1: 25 ms),get_chrom(columns="coords")15 ms; Liu 2025 MOp 255 MB (2.1: 272 MB). Regression:get_chromwith spot tracks 0.53 s at 1e7 (2.1: 39 ms). In-memory row order is now cell › trace › chromosome › bin (iter_traces/iter_cellsstill stream chromosome › cell › trace). 2.0 / 2.1 stores stay readable;python -m uchrom.io.upgraderewrites them (keeping the original row order); the streaming writer andbenchmarks/fig2/replicate_store.pywrite 2.2..chromdata.zarrformat 2.1 — spot tables partitioned by chromosome.tables/spots/chrom=<name>/part-0.parquet(and the same forspot_tracksandlayers/<key>), sorted cell › trace › bin inside a partition, 16,384-row row groups (was 65,536);index/gainspartition_offsets/partition_groups/partition_chrom/cell_partitionand onecell_offsetsentry per (chromosome, cell). Backedget_chromreads one partition (Takei 2025, 1e7 spots: 0.77 s → 0.04 s; 1e8: 8.9 s → 0.49 s);get_cellreads one row group per chromosome and is slower (8 → 23 ms at 1e7; 22 → 42 ms at 1e8; withcolumns="coords"4.4 / 22 ms). Stores with many per-spot tracks grow (Takei: +35 %, the per-spot z-scores compress worse once a cell’s chromosomes are apart; no Parquet option recovers it —benchmarks/fig2/spot_tracks_encoding.py); Liu 2025 MOp shrinks 3 %. Full reads: 0.98 → 1.11 s at 1e7. Stored row order is now chromosome › cell › trace › bin, also for in-memoryiter_traces/iter_cells. 2.0 stores stay readable (in memory and backed); 2.0 readers cannot read 2.1 stores —python -m uchrom.io.upgrade old.chromdata.zarr new.chromdata.zarrrewrites them.Column selection:
get_cell/get_trace/get_chrom,iter_traces/iter_cells,to_memoryandChromData.readtakecolumns=/tracks=(columns="coords": coordinates and key columns only); the web browser reads geometry withcolumns="coords".iter_traces(chrom=...);batch="auto".Streaming (roadmap step 7):
uchrom.settings.memory_budget;ChromData.writer(path)(uchrom.core.stream.ChromDataWriter: append chunks, spill per chromosome, merge each partition in pieces within the budget — the store equals the in-memory writer’s);ChromData.from_fofct(..., out=...)(chunked) andread_seqfish_multiomics(..., out=...)(FOV by FOV) return the store backed.uchrom.fea.distance_map: exact per-chromosome median / mean distance maps streamed over the traces (median per row band; bitwise equal to the densemean_distance_matrix).call_loops_axiswise_fandcall_tads_by_pvaltakestreaming=(default on backed data;uchrom.fea.arc_stream.axis_cube_streaming): memoryO(n_bins²)instead ofO(n_traces × n_bins²), same calls..chromdata.zarr— the format 2.0 container (Zarr v3 + Parquet, the SpatialData approach) is the new default:cd.write("x.chromdata.zarr")(orx.cdz, the same store in one zip file). Spots are stored sorted by (cell, trace, bin) with anindex/of row offsets, so a round trip keeps every spot but may change the row order (ChromData.read(path, original_order=True)restores it). Backed mode:ChromData.read(path, backed=True)loads the small tables only;get_cell/get_trace/get_chromread the rows they need;iter_traces(batch)/iter_cells(batch)stream chunks (also on in-memory objects);to_memory()loads all..h5cdfiles (1.x and 2.0) are still read; writing.h5cdis deprecated (it warns).python -m uchrom.io.upgrade old.h5cdconverts toold.chromdata.zarr. The web browser opens stores backed; the example-data builders, tutorials and CLIs write the new format. Requires Python ≥ 3.11 and the new core dependencieszarr>=3andpyarrow.FOF-CT writer:
cd.to_fofct(path, cell_table=None, rna_table=None)anduchrom.io.write_fofctwrite the 4DN FOF-CT core table (with its##/#header lines, fromuns['fofct_header']oruns) and optionally the cell and RNA-spot companion tables — the inverse offrom_fofct.from_fofctnow parses floats exactly (float_precision="round_trip"; the default parser could be 1 ulp off), so FOF-CT round trips are value-identical.cd.resultsis now aResultsStoreofResultRecords: every result carries its kind, producing function, parameters, inputs, uchrom version and timestamp. Unwritable values raiseTypeErrorat assignment..h5cdformat 1.4 stores the provenance as attrs; older files still read.One calling convention for the structure callers (
call_tads_by_pval,call_domains_fishnet,call_loops_axiswise_f,call_compartments_axes_pc, newcall_tads_di): keyword-onlychrom=None(all chromosomes, one merged table),trace_ids,cells,params,key_added="<what>.<method>",copy. Positionalchrom,store=,result_key=andstrc.call_*_multiare deprecated (they warn and keep the 1.x keys).uchrom.tl/uchrom.ppalias namespaces..h5cdformat 2.0 — thebinslocus axis (breaking on disk; 1.x files still read).cd.bins(indexbin_id), requiredspots.bin_id(derived fromchrom/start/endby every constructor and loader), bin-levelcd.bin_tracksvs per-spotcd.spot_tracks,cd.binm, typedcd.intervals(domain | pair | peak | segment, withto_bins). Structure callers also fillcd.intervals; compartments add segments and a per-binpc2track.fea/strcfeature writers project onto bins (fea.project_interval_features_to_bins; the spot version is deprecated).cd.tracksis a deprecated spot-aligned view.uchrom.io.upgrade_h5cdandpython -m uchrom.io.upgraderewrite 1.x files; the web browser lists bin- and spot-level tracks and typed intervals. Large integer / string datasets are chunked and gzip-4 compressed (write(compression=…, compress_floats=…)); strings are encoded / decoded vectorised.Fix:
from_fofctkept the last##columns=name intact when it ends in)(n_per_dist(um)was read asn_per_dist(um).seqFISH+ loader (Takei 2025):
tracesnow reportsallele,n_spots,n_loci,n_duplicate_lociand per-column allele majority / agreement instead of the first spot’sdbscan_*values; the docstring warns that the defaultallele_col="dbscan_ldp_nbr_allele"often merges homologs (default unchanged).
0.2.0 — 2026-04¶
Renamed package from
nucboxtouchrom/u-chrom.New core container
ChromDatawith HDF5 persistence (.h5cd) and format versioning (see Concepts — on-disk format).uchrom.io.save_particles/ extendedread_particles— all reconstruction modules now write.h5cdby default.Bulk MDS reconstruction (
uchrom.recon.bulk.mds) — PyTorch SMACOF with inter-chromosomal whole-genome support.Browser updates:
Trace-aware rendering (per
(chrom, trace_id)polymer).“Per Trace” colour mode.
Unified Chromosome / Region / Trace Management tab.
Trace Statistics panel (Distance Matrix / Contact Map / Rg Histogram / Split by Trace).
Click-to-identify trace overlay.
Draggable Layers / Properties divider.
macOS Dock icon + startup logo.
ArcFISH-style axis-wise F-test loop caller (
uchrom.strc.loop.call_loops_axiswise_f) — GPU-accelerated via PyTorch, CLI underpython -m uchrom.strc.loop.New population-aggregate modules:
uchrom.fea.distance(median distance matrix, contact frequency, radius of gyration)uchrom.fea.arc(axis variance cube, LOWESS filter+normalise, axis weights)uchrom.pl.trace_stats(matplotlib helpers)uchrom.utils.stats(Cauchy combination test, log-log LOWESS)
Version stamp on
.h5cdfiles with MAJOR/MINOR semantics and legacy-file handling.