uchrom.io

I/O helpers with lazy optional-dependency imports.

uchrom.io.create_contacts_from_cool(pixel_df, particles, min_count=0, min_skip=2)[source]
uchrom.io.create_contacts_from_pairs(df, particles, min_skip=2)

Count contacts between all particles.

uchrom.io.create_particles_from_cool(df, particle_size, prev_particles=None, random_seed=0, max_radius=10, *, min_label='start1', max_label='end2')
uchrom.io.create_particles_from_pairs(df, particle_size, prev_particles=None, random_seed=0, max_radius=10, *, min_label='pos1', max_label='pos2')
uchrom.io.filter_contacts(contacts, genome_ranges)[source]
uchrom.io.load_cool(path, pos='mid')[source]

Load pixels from cooler.

uchrom.io.read_pairs(path)[source]
uchrom.io.read_particles(path)[source]

Read reconstruction output. Auto-detects format by extension.

  • .chromdata.zarr / .cdz / .h5cd → ChromData, returned as a flat DataFrame in the row order it was saved in

  • .csv or other → CSV

uchrom.io.save_particles(df, path, cell_id=None)[source]

Save reconstruction output as a ChromData file or a CSV.

Format inferred from file extension:

  • .csv → CSV via df.to_csv

  • .chromdata.zarr / .cdz → ChromData format 2.0 (recommended)

  • .h5cd or other → ChromData HDF5 (deprecated)

cell_id is only propagated into the ChromData branch (CSV has no dedicated cell column); it is passed through to ChromData.from_dataframe().

Format upgrade

Convert ChromData files to the format 2.0 container (.chromdata.zarr).

ChromData.read reads every older file (.h5cd 1.x and 2.0) in memory; this rewrites them once, so later reads are fast and can be backed (ChromData.read(path, backed=True)):

python -m uchrom.io.upgrade old.h5cd new.chromdata.zarr
python -m uchrom.io.upgrade data/*.h5cd --out-dir converted/        # <stem>.chromdata.zarr
python -m uchrom.io.upgrade data/*.h5cd --out-dir converted/ --format cdz

What changes (see docs/source/guide/chromdata_2_0_design.md):

  • 1.x → 2.0 data model: bins = the unique spot loci; spots store bin_id instead of chrom/start/end (still available via cd.spots / cd.to_dataframe()); the spot-aligned 1.x tracks split into bin-level tracks (columns constant within every bin) and spot_tracks;

  • the container: Zarr v3 + Parquet, spots sorted by (cell, trace, bin) with an index/ of row offsets (index/source_row keeps the original order).

The target format follows the destination suffix (.chromdata.zarr, .cdz, or — deprecated — .h5cd). The source is never modified.

uchrom.io.upgrade.convert(src: str | Path, dst: str | Path | None = None, *, overwrite: bool = False, coord_dtype: str = 'float64', compression: str | None = 'gzip', compression_opts: int | None = 4, compress_floats: bool = False) → dict

the conversion is not specific to .h5cd sources

uchrom.io.upgrade.default_target(src: str | Path, fmt: str = 'zarr', out_dir: str | Path | None = None) → Path[source]

<dir>/<stem>.chromdata.zarr (or .cdz / .h5cd) for src.

uchrom.io.upgrade.file_format_version(path: str | Path) → str[source]

The uchrom_format_version of a ChromData file ("1.0" for an unversioned .h5cd).

uchrom.io.upgrade.main(argv: list | None = None) → int[source]
uchrom.io.upgrade.upgrade_h5cd(src: str | Path, dst: str | Path | None = None, *, overwrite: bool = False, coord_dtype: str = 'float64', compression: str | None = 'gzip', compression_opts: int | None = 4, compress_floats: bool = False) → dict[source]

Read src (.h5cd 1.x / 2.0, or any ChromData store) and write it to dst in the current format (.chromdata.zarr 2.2, .h5cd 2.0); a zarr / cdz source is read in its original row order.

dst defaults to <src stem>.chromdata.zarr next to src; its suffix picks the container (.chromdata.zarr, .cdz, or the deprecated .h5cd). coord_dtype applies to the zarr container; compression / compression_opts / compress_floats to .h5cd targets (see ChromData.write()).

Returns a summary: source / target versions and containers, n_spots, n_bins and which tracks are bin- vs spot-level. src is never modified; dst must differ from it.