I/O and file formats¶
uchrom.io handles reading sequencing and imaging data, and saving
reconstruction output.
Supported inputs¶
Format |
Reader |
Returns |
|---|---|---|
|
|
DataFrame (read-level pairs) |
|
|
contact pixel DataFrame / bin+matrix |
|
|
bin+matrix |
|
|
DataFrame (chrom, start, x, y, z) |
|
|
DataFrame |
|
|
DataFrame / |
FOF-CT core CSV |
|
|
Reconstruction output: save_particles¶
All reconstruction modules write through
uchrom.io.save_particles(df, path):
pathending in.csv→ CSV (legacy).chromdata.zarr/.cdz→ChromDataformat 2.0 (recommended).h5cdor anything else →ChromDataHDF5 (deprecated)
from uchrom.io import save_particles
save_particles(df, "out.chromdata.zarr") # ChromData
save_particles(df, "out.csv") # plain CSV
read_particles auto-detects the format by extension so downstream code
(browser, analysis) does not care which format is on disk; ChromData
files come back in the row order they were saved in.
FOF-CT import details¶
ChromData.from_fofct(path) handles several real-world FOF-CT quirks:
##Columns=(…)declared case-insensitively (some writers use lower-case)CSV column header row repeated after the
##Columns=metadataHeader lines wrapped in double quotes and padded with trailing commas
Optional extra columns (e.g.
Readout) preserved inspots##XYZ_Unit,##Genome_Assemblypromoted to top-levelunskeys
Unit conversion: FOF-CT files are typically in µm; you can scale in place
by multiplying cd.coords yourself, or pass a custom unit label into
uns['xyz_unit'] so downstream plots label correctly.
FOF-CT export: to_fofct¶
cd.to_fofct(path, cell_table=None, rna_table=None) (or
uchrom.io.write_fofct(cd, path, …)) writes the inverse of from_fofct:
the core table:
##FOF-CT_version,##Table_namespace,##genome_assembly,##XYZ_unitand the#key: valuelines — taken fromuns['fofct_header']when the data came from FOF-CT, else fromuns— then##columns=(Spot_ID, Trace_ID, X, Y, Z, Chrom, Chrom_Start, Chrom_End, [Cell_ID, Sub_Cell_ROI_ID, Extra_Cell_ROI_ID], …extra spot columns, spot tracks), one row per spot;cell_table=:cd.cellsas a4dn_FOF-CT_celltable (Cell_IDfirst;centroid_*→Cent_ROI_*,nucleus_area_um2→area(um2), therna.prefix removed);rna_table=:points['rna']as a4dn_FOF-CT_rnatable.
Coordinates are written with full precision and from_fofct parses them
exactly (float_precision="round_trip"), so from_fofct(to_fofct(cd))
gives back the same coordinates, ids, loci, extra columns, cell table and
RNA spots. Spot tracks come back as spot columns; bin_tracks=True also
writes the per-locus tracks, broadcast to spots. FOF-CT has no place for
cellm, binm, layers, intervals, results or other uns keys.
cd.to_fofct("out_core.csv", cell_table="out_cell.csv", rna_table="out_rna.csv",
header={"lab_name": "My lab"})
ChromData files and format versioning¶
cd.write(path) / ChromData.read(path) pick the container from the
path: .chromdata.zarr (Zarr v3 + Parquet, the default), .cdz (the
same store zipped into one file), .h5cd (HDF5; read indefinitely,
writing deprecated). Convert older files with
python -m uchrom.io.upgrade old.h5cd (→ old.chromdata.zarr) or
uchrom.io.upgrade_h5cd(src, dst).
Every container records two attributes:
Attribute |
Example |
|---|---|
|
|
|
|
Read semantics:
Same MAJOR → read (higher MINOR warns, unknown fields ignored)
Different MAJOR →
ValueErrorwith upgrade guidanceMissing attribute → assume legacy 1.0 with a warning
New MAJOR versions add a _read_vN(cls, f) function in
uchrom/core/cdata.py (HDF5) or a reader in uchrom/core/zarrcd.py
(zarr) plus any migration helper needed; the dispatch in
ChromData.read is the single place that maps path and version to
reader.