uchrom.io¶
I/O helpers with lazy optional-dependency imports.
- uchrom.io.create_contacts_from_pairs(df, particles, min_skip=2)¶
Count contacts between all particles.
- uchrom.io.create_particles_from_cool(df, particle_size, prev_particles=None, random_seed=0, max_radius=10, *, min_label='start1', max_label='end2')¶
- uchrom.io.create_particles_from_pairs(df, particle_size, prev_particles=None, random_seed=0, max_radius=10, *, min_label='pos1', max_label='pos2')¶
- uchrom.io.read_particles(path)[source]¶
Read reconstruction output. Auto-detects format by extension.
.chromdata.zarr/.cdz/.h5cd→ ChromData, returned as a flat DataFrame in the row order it was saved in.csvor other → CSV
- uchrom.io.save_particles(df, path, cell_id=None)[source]¶
Save reconstruction output as a ChromData file or a CSV.
Format inferred from file extension:
.csv→ CSV viadf.to_csv.chromdata.zarr/.cdz→ ChromData format 2.0 (recommended).h5cdor other → ChromData HDF5 (deprecated)
cell_idis only propagated into the ChromData branch (CSV has no dedicated cell column); it is passed through toChromData.from_dataframe().
Format upgrade¶
Convert ChromData files to the format 2.0 container (.chromdata.zarr).
ChromData.read reads every older file (.h5cd 1.x and 2.0) in
memory; this rewrites them once, so later reads are fast and can be
backed (ChromData.read(path, backed=True)):
python -m uchrom.io.upgrade old.h5cd new.chromdata.zarr
python -m uchrom.io.upgrade data/*.h5cd --out-dir converted/ # <stem>.chromdata.zarr
python -m uchrom.io.upgrade data/*.h5cd --out-dir converted/ --format cdz
What changes (see docs/source/guide/chromdata_2_0_design.md):
1.x → 2.0 data model:
bins= the unique spot loci;spotsstorebin_idinstead ofchrom/start/end(still available viacd.spots/cd.to_dataframe()); the spot-aligned 1.xtrackssplit into bin-leveltracks(columns constant within every bin) andspot_tracks;the container: Zarr v3 + Parquet, spots sorted by (cell, trace, bin) with an
index/of row offsets (index/source_rowkeeps the original order).
The target format follows the destination suffix (.chromdata.zarr,
.cdz, or — deprecated — .h5cd). The source is never modified.
- uchrom.io.upgrade.convert(src: str | Path, dst: str | Path | None = None, *, overwrite: bool = False, coord_dtype: str = 'float64', compression: str | None = 'gzip', compression_opts: int | None = 4, compress_floats: bool = False) dict¶
the conversion is not specific to .h5cd sources
- uchrom.io.upgrade.default_target(src: str | Path, fmt: str = 'zarr', out_dir: str | Path | None = None) Path[source]¶
<dir>/<stem>.chromdata.zarr(or.cdz/.h5cd) forsrc.
- uchrom.io.upgrade.file_format_version(path: str | Path) str[source]¶
The
uchrom_format_versionof a ChromData file ("1.0"for an unversioned.h5cd).
- uchrom.io.upgrade.upgrade_h5cd(src: str | Path, dst: str | Path | None = None, *, overwrite: bool = False, coord_dtype: str = 'float64', compression: str | None = 'gzip', compression_opts: int | None = 4, compress_floats: bool = False) dict[source]¶
Read
src(.h5cd1.x / 2.0, or any ChromData store) and write it todstin the current format (.chromdata.zarr2.2,.h5cd2.0); a zarr / cdz source is read in its original row order.dstdefaults to<src stem>.chromdata.zarrnext tosrc; its suffix picks the container (.chromdata.zarr,.cdz, or the deprecated.h5cd).coord_dtypeapplies to the zarr container;compression/compression_opts/compress_floatsto.h5cdtargets (seeChromData.write()).Returns a summary: source / target versions and containers,
n_spots,n_binsand which tracks are bin- vs spot-level.srcis never modified;dstmust differ from it.