uchrom.core¶
- class uchrom.core.UChromProject(chromdata: ChromData | None = None, chromdata_path: Path | None = None, links: dict[str, ~typing.Any]=<factory>)[source]¶
Bases:
objectCoordinate a
ChromDataobject with linked MuData and scool files.This wrapper is intentionally lightweight. The data remain in their native files; the project object centralizes links, validation, and descriptive metadata.
ChromData¶
- class uchrom.core.ChromData(coords: ndarray, spots: DataFrame, *, bins: DataFrame | None = None, cells: DataFrame | None = None, cellm: Dict[str, ndarray] | None = None, tracks: DataFrame | None = None, spot_tracks: DataFrame | None = None, binm: Dict[str, ndarray] | None = None, intervals: Mapping[str, DataFrame] | None = None, traces: DataFrame | None = None, layers: Dict[str, ndarray] | None = None, results: dict | None = None, uns: dict | None = None, points: Dict[str, DataFrame] | None = None, cell_shapes: Dict[str, DataFrame] | None = None, linked_adata=None, validate: bool = True, _own_spots: bool = False)[source]¶
Bases:
objectChromatin Data — the core container for U-Chrom.
See the module docstring of
uchrom.core.cdatafor the full purpose, hierarchy, FOF-CT mapping, and on-disk format contract. Summary:The central abstraction is a “structure table” — genomic bins mapped to 3D coordinates. Each row of
spots(with the corresponding row ofcoords) is one Spot.Spots are grouped hierarchically as Cell → Trace → Spot. A Trace is an ordered chromatin-fibre polymer; a Cell contains one or more traces.
All analysis in U-Chrom consumes or produces
ChromData. Reconstruction modules (uchrom.recon) emit it, structure callers (uchrom.strc) decoratecd.results[...]with TADs / loops / compartments, and the browser / plotters render it.
- Parameters:
coords (ndarray, shape (n_spots, 3)) – 3-D coordinates (x, y, z) per spot.
spots (DataFrame, shape (n_spots, ≥2)) – Per-spot metadata. Required:
trace_id(int or str, will be categorified) and eitherbin_id(withbins=) or the locus columnschrom/start(0-based) /end(non-inclusive), from whichbinsandbin_idare derived. Optional:cell_id(int or str),spot_id, FOF-CTsub_cell_roi_id/extra_cell_roi_id, and any experiment-specific annotation column (carried through verbatim).bins (DataFrame, optional) – Locus axis (
chrom, start, end+ optionalresolution,name, …); rowiisbin_idi. Derived from the spot loci when omitted (ordered by chromosome, start, end).cells (DataFrame, optional) – Per-cell metadata indexed by cell_id.
cellm (dict[str, ndarray], optional) – Per-cell multi-dimensional annotations (embeddings, UMAP, …). Each array’s first axis length =
n_cells.tracks (DataFrame, optional) – Per-locus signals (ATAC, ChIP-seq, …):
n_binsrows (or an index namedbin_id) →bin_tracks. A spot-aligned 1.x table (n_spotsrows) is still accepted with aDeprecationWarningand split (columns constant within every bin →bin_tracks, the rest →spot_tracks).spot_tracks (DataFrame, optional) – Per-spot signals (IF intensity, seqFISH z-scores),
n_spotsrows.binm (dict[str, ndarray], optional) – Per-locus multi-dimensional arrays, first axis
n_bins.intervals (dict[str, DataFrame], optional) – Typed interval tables (see
uchrom.core.intervals).traces (DataFrame, optional) – Per-trace metadata indexed by trace_id.
layers (dict[str, ndarray], optional) – Alternative coordinate sets, each with shape
(n_spots, 3)— e.g. raw / drift-corrected / aligned.points (dict[str, DataFrame], optional) – Named sets of non-genomic 3-D points in the same frame as
coords— e.g. nascent RNA spots or IF puncta. Each DataFrame needsx, y, z;cell_id(if present) links points to cells and is used when subsetting.cell_shapes (dict[str, DataFrame], optional) – Cell outlines:
cell_id+geometry(WKB bytes), one row per cell (seeset_cell_shapes()); subset with the cells.results (dict or ResultsStore, optional) – Analysis outputs, stored as a
ResultsStore(cd.results[key]returns the value,cd.results.record(key)the value plus provenance). Keys follow"<what>.<method>":'tads.arcfish','loops.axiswise_f','compartments.axes_pc', ….uns (dict, optional) – Unstructured metadata preserved on disk. Conventional keys:
'genome_assembly','xyz_unit','fofct_header'. Auto-discovery context may use'dataset_references'for source papers/repositories and'user_annotations'for user-provided priors, constraints, or hypothesis seeds.validate (bool) – If True (default), validate internal consistency on construction (coords shape, spots required columns, tracks / layers alignment).
- n_spots, n_traces, n_cells, chroms
- Type:
derived accessors
- Key methods
- -----------
- from_dataframe, from_fofct, read, write, to_dataframe,
- get_cell, get_trace, get_chrom, compute_distances
- On-disk format — ``.h5cd``
- --------------------------
- Versioned HDF5. See :mod:`uchrom.core.cdata` module docstring for
- the full layout and the :meth:`read` / :meth:`write` round-trip
- contract.
Notes
Subsetting (
cd[mask],get_chrometc.) always returns a newChromData; the source is not mutated.Global pairwise distance matrices are intentionally not stored — they are biologically meaningful per-trace, not across cells, and would be O(n²) memory. Compute on demand via
compute_distances(trace_id=...)().String columns (
chrom,trace_id,cell_id) are auto-converted topd.Categoricalfor ~10× memory savings.
- add_reference(*, reference_id: str | None = None, role: str = 'user_supplied_reference', title: str | None = None, doi: str | None = None, pmid: str | None = None, url: str | None = None, year: int | None = None, notes: str | None = None, **extra: Any) dict[source]¶
Add a source reference to
uns['dataset_references'].Parameters are intentionally metadata-oriented rather than tied to one publication database.
roleshould describe how the reference relates to the dataset, for example'primary_dataset_paper','data_repository','supplementary_table', or'related_biology_prior'.
- add_user_annotation(*, annotation_id: str | None = None, scope: str, text: str, target: str | None = None, tags: List[str] | None = None, confidence: str = 'user_asserted', **extra: Any) dict[source]¶
Add a user annotation to
uns['user_annotations'].Use annotations for cell-type notes, marker priors, analysis constraints, hypothesis seeds, negative constraints, field semantics, or quality warnings. They are surfaced in the discovery schema and agent context but still require notebook validation before becoming evidence.
- property auto_discovery_schema: dict¶
use
uchrom_discovery.get_discovery_schema(cd).- Type:
Deprecated
- property backed: bool¶
Truefor aBackedChromData.
- property bin_tracks: DataFrame¶
Per-locus signals (bulk ATAC, ChIP, GC, annotation features, compartment scores …),
n_binsrows indexed bybin_id.This is the table the 2.0 design calls
tracks; during the 2.x seriescd.tracksremains the deprecated spot-aligned view.
- property bins: DataFrame¶
DataFrameindexed bybin_idwithchrom(category),start,endand optional per-locus columns (resolution,name, …). Shared by all cells; subsetting spots keeps every bin.- Type:
Locus axis
- build_discovery_schema(*, store: bool = True, **kwargs) dict[source]¶
Deprecated: use
uchrom_discovery.build_discovery_schema(cd, ...).
- cell_positions(key: str = 'default', *, physical: bool = False) DataFrame[source]¶
Cell centroids as a DataFrame indexed by
cell_id:x,y[,z] (float, NaN where unknown) [+region]; the record (unit, frame, status, …) is in.attrs['cell_spatial'].Reads the record
uns['cell_spatial'][key](1.x records and theirspatial_x/spatial_ycolumns included). Without a record,key="default"falls back to known column names —centroid_x/y/z, then the olderx_centroid/y_centroid/z_centroidandspatial_x/spatial_y— withstatusunknown (attrs['cell_spatial']['unrecorded'] is True).KeyErrorwhen there is no position set.physical=Trueconverts positions stored in voxels (recordvoxel_size/voxel_unit) to physical units.
- cell_shapes_geodataframe(key: str = 'cell')[source]¶
cell_shapes[key]as ageopandas.GeoDataFrameindexed bycell_id(needs shapely + geopandas; the record is in.attrs['cell_shapes']).
- compute_distances(trace_id=None) ndarray[source]¶
Compute pairwise Euclidean distance matrix.
- Parameters:
trace_id (optional) – If given, compute only for spots in that trace. If None, compute for all spots (use with caution on large data).
- Return type:
np.ndarray, shape (n, n)
- property dataset_references: List[dict]¶
Dataset-level source references used as auto-discovery priors.
References are stored in
uns['dataset_references']and round-trip with.h5cdfiles. They are intended for primary dataset papers, data repositories, supplementary tables, method papers, and related biological priors.
- describe_for_agent(*, max_items: int = 40) str[source]¶
Deprecated: use
uchrom_discovery.describe_for_agent(cd).
- classmethod from_dataframe(df: DataFrame, *, cell_id=None, **kwargs) ChromData[source]¶
Create from a reconstruction output DataFrame.
Expects columns: chrom, start, end, x, y, z. Each chromosome becomes one trace.
- Parameters:
df (DataFrame with columns chrom, start, end, x, y, z.)
cell_id (hashable, optional) – If given, tag every spot with this cell identifier (e.g. derived from the output filename for single-cell reconstruction). The DataFrame’s own
cell_idcolumn, if any, takes precedence.**kwargs – Forwarded to the
ChromDataconstructor.
- classmethod from_fofct(core_path: str | Path, *, cell_table: str | Path | None = None, rna_table: str | Path | None = None, mapping_table: str | Path | None = None, cell_value_prefix: str | None = 'rna.', out: str | Path | None = None, chunksize: int | str = 'auto', memory_budget=None, **kwargs) ChromData[source]¶
Read a FOF-CT core table, optionally with its companion tables.
- Parameters:
core_path (path) – Path to the FOF-CT core table (CSV/TSV/TXT).
cell_table (path, optional) – FOF-CT cell data table (
4dn_FOF-CT_cell). Loaded intocellsindexed bycell_id:cent_ROI_x/ybecomecentroid_x/y,area(um2)/area_cyto(um2)becomenucleus_area_um2/cytoplasm_area_um2, and every other numeric column (per-cell RNA copy numbers in 4DN tables) is prefixed withcell_value_prefix(Nanog→rna.Nanog). Centroid columns become the cell positions (uns['cell_spatial'],cell_positions()):Cent_ROI_x/y[/z](ascentroid_x/y[/z]) orcell_center_x_global/cell_center_y_global(Liu et al. 2025, global frame), in the table’sXYZ_unit.mapping_table (path, optional) – FOF-CT Cell/ROI mapping table (
4dn_FOF-CT_mapping):ROI_Boundariesof eachCell_IDin global coordinates, as OME polygon points strings ("x1,y1 x2,y2 …") →cell_shapes['cell'](uns['cell_shapes']['cell']: frame"global"). Rows keyed by sub- / extra-cell ROI ids and OBJ meshes are skipped with a warning.rna_table (path, optional) – FOF-CT RNA spot table (
4dn_FOF-CT_rna). Loaded intopoints['rna']with columnsx, y, z, gene, gene_id, cell_id, spot_id, …in the same frame ascoords.cell_value_prefix (str or None) – Prefix for unrecognised numeric cell-table columns;
Nonekeeps their names.out (path, optional) – Streaming import: read the core table in chunks of
chunksizerows and write them to this.chromdata.zarrstore withChromDataWriter; returns the store opened backed. Peak memory is one chunk plus the writer’s merge pieces (bounded bymemory_budget), not the table. The store equalsfrom_fofct(core_path, ...).write(out)(same spots, order, dtypes, small tables). Withoutoutthe table is read into memory (aResourceWarningis issued when the file is larger than a fifth of the memory budget).chunksize (int or
"auto") – Rows per chunk of a streaming import ("auto": from the memory budget).memory_budget (int or str, optional) – Defaults to
uchrom.settings.memory_budget.**kwargs – Additional keyword arguments passed to ChromData constructor (e.g. cells, tracks, uns). A streaming import accepts
cells,cellm,traces,pointsanduns.
- classmethod from_pyhim_trace(ecsv_path: str | Path, barcode_dict: dict | DataFrame | None = None, **kwargs) ChromData[source]¶
Read a PyHiM chromatin-trace ECSV table into a ChromData.
PyHiM (Devos et al. 2024) emits one ECSV file per trace-building run. Schema (from
chromatin_trace_table.pyupstream):Spot_ID, Trace_ID, x, y, z, Chrom, Chrom_Start, Chrom_End, ROI #, Mask_id, Barcode #, label
meta['comments']carriesxyz_unit=...andgenome_assembly=....- Parameters:
ecsv_path (path) – Path to the ECSV file written by PyHiM.
barcode_dict (dict[int, (chrom, start, end)] or DataFrame, optional) – Required when
Chrom/Chrom_Start/Chrom_Endare empty in the ECSV (PyHiM does not always populate them). As a DataFrame, expects columnsbarcode, chrom, start, end. IfChromis populated,barcode_dictis ignored.**kwargs – Additional keyword arguments passed to the ChromData constructor (
cells,tracks,uns, …).
Notes
Mask_idbecomescell_id(PyHiM convention).ECSV header comments are captured in
cd.uns['pyhim']['ecsv_comments']and anyxyz_unit/genome_assemblyentries are also promoted tocd.unsdirectly (matchingfrom_fofct()).
- classmethod from_seqfish_multiomics(spot_glob, **kwargs) ChromData[source]¶
Load Takei 2025 DNA seqFISH+ cerebellum data.
Thin shim around
uchrom.io.seqfish_multiomics.read_seqfish_multiomics(). See that function for the full parameter list.
- classmethod from_seqfish_multiomics_linked(spot_glob, **kwargs)[source]¶
Load linked Takei 2025 DNA tracing + RNA AnnData artifacts.
Thin shim around
uchrom.io.seqfish_multiomics.load_seqfish_multiomics_linked(). Returns aChromDatawith RNA expression available atcd.linked_adataand can write paired.h5cd/.h5adfiles.
- classmethod from_takei2025_cerebellum(**kwargs)[source]¶
Load linked Takei 2025 cerebellum data.
Thin shim around
uchrom.io.seqfish_multiomics.load_takei2025_cerebellum(). Returns aChromDatawith RNA expression available atcd.linked_adata.
- get_cell(cell_id, *, columns=None, tracks=None) ChromData[source]¶
The spots of one cell (a new object).
columns/tracksselect the spot-aligned columns to keep —columns="coords"keeps the coordinates and the key columns only,columns=[...]adds the named spot columns / spot tracks / layers,tracks=[...]picks spot tracks (seeresolve_selection()). The default keeps everything. On a backed object the selection decides what is read from disk.
- get_cells(cell_ids, *, columns=None, tracks=None) ChromData[source]¶
The spots of several cells (ids compared as strings; unknown ids are ignored). See
get_cell()forcolumns/tracks.
- get_chrom(chrom: str, *, columns=None, tracks=None) ChromData[source]¶
The spots of one chromosome (see
get_cell()forcolumns/tracks).
- get_trace(trace_id, *, columns=None, tracks=None) ChromData[source]¶
The spots of one trace (see
get_cell()forcolumns/tracks).
- property intervals: IntervalStore¶
Typed interval tables (TADs, loops, peaks, segments) — see
uchrom.core.intervals.
- iter_cells(batch=64, *, columns=None, tracks=None)[source]¶
Yield in-memory
ChromDatachunks ofbatchcells (cell code order; the rows of a chunk in stored order, seeiter_traces()).
- iter_traces(batch=1024, *, chrom=None, columns=None, tracks=None)[source]¶
Yield in-memory
ChromDatachunks ofbatchtraces.Traces are visited in the order a
.chromdata.zarrstore holds them — chromosome › cell › trace, with the spots of a trace in bin order; a trace id shared by several cells (or a trace spanning several chromosomes) counts once per (chromosome, cell) run. A backed object yields identical chunks while reading one row range per chunk.- Parameters:
batch (int or
"auto") – Traces per chunk;"auto"sizes chunks fromuchrom.settings.memory_budget.chrom (str, optional) – Only the traces of this chromosome.
columns – Column selection of the chunks (see
get_cell()), e.g.columns="coords"for coordinates and keys only.tracks – Column selection of the chunks (see
get_cell()), e.g.columns="coords"for coordinates and keys only.
- link_anndata(adata, *, cell_id_col: str | None = None, copy_obs: bool = True, copy_obsm: bool = True) int[source]¶
Import cell-level metadata from an AnnData into this ChromData.
Matches cells by
cell_id: each unique value inspots['cell_id']is looked up inadata.obs(by index, or by the column cell_id_col if given). Matched cells get theiradata.obscolumns merged intoself.cellsand theiradata.obsmarrays copied intoself.cellm.If
self.cellsalready exists, its row order is preserved and AnnData rows are aligned onto that cell axis. This is important for multi-omics loaders such as Takei 2025, where chromatin tracing coordinates live incoords/spots, RNA/IF signals live in spot-leveltracks, and mRNA clustering/UMAP already live incells/cellm.- Parameters:
adata (anndata.AnnData) – The single-cell dataset to link (e.g. scRNA-seq).
cell_id_col (str, optional) – Column in
adata.obsthat holds cell identifiers matchingspots['cell_id']. IfNone,adata.obs.indexis used as the key.copy_obs (bool) – If True (default), copy
adata.obscolumns intoself.cells.copy_obsm (bool) – If True (default), copy
adata.obsmarrays intoself.cellm.
- Returns:
Number of cells matched.
- Return type:
- Raises:
KeyError – If
spotshas nocell_idcolumn.
- link_cool(path: str | Path, *, key: str = 'default', label: str | None = None, genome_assembly: str | None = None, **metadata: Any) dict[source]¶
Record a linked bulk / pseudo-bulk
.coolor.mcoolinuns['linked_cool'](the counterpart oflink_scool()for one matrix per file). The file stays external.
- link_mudata(path: str | Path, *, key: str = 'default', modalities: List[str] | None = None, cell_axis: str | None = None, spatial_source: str | None = None, not_measured_spatial: bool | None = None, **metadata: Any) dict[source]¶
Record a linked MuData file in
uns['linked_mudata'].The MuData object is kept external.
ChromDatastores only lightweight provenance so multi-modal cell matrices are not forced intocoords/spots.
- link_scool(path: str | Path, *, key: str = 'default', cell_name: str | None = None, barcode_source: str | None = None, genome_assembly: str | None = None, bin_size: int | None = None, coordinate_status: str = 'contacts, not xyz coordinates', **metadata: Any) dict[source]¶
Record a linked scool contact store in
uns['linked_scool'].
- link_spatialdata(path: str | Path, *, key: str = 'default', table: str | None = None, region: Any | None = None, region_key: str | None = None, instance_key: str | None = None, cell_col: str | None = None, coordinate_system: str | None = None, read_attrs: bool = True, **metadata: Any) dict[source]¶
Record a linked SpatialData zarr store in
uns['linked_spatialdata'].The store stays external (images, labels, shapes, other tables). The cell map joins its annotating
tableto ChromData cells: rows whoseregion_keycolumn equalsregionare cells, and theirinstance_keyvalue equals the ChromDatacell_id— or the value of thecellscolumncell_colwhen the ids differ. Withread_attrs(and spatialdata installed)table/region/region_key/instance_keydefault to the table’sspatialdata_attrs. Relative paths resolve against the directory of the ChromData store (as forlink_scool()).
- property linked_adata¶
Linked AnnData object, loaded lazily from
unsmetadata if possible.
- load_linked_anndata(path: str | Path | None = None)[source]¶
Load, cache, and return the linked AnnData object.
- load_linked_mudata(*, key: str = 'default')[source]¶
Load a linked MuData file by key.
mudatais an optional dependency. A clear ImportError is raised when it is unavailable.
- load_linked_scool(*, key: str = 'default', cell: str | None = None)[source]¶
Load metadata or one cell cooler from a linked scool file.
If
cellisNone, returns a summary dict with available scool cells. Ifcellis supplied, returnscooler.Coolerfor that cell group.
- load_linked_spatialdata(*, key: str = 'default', cells: Sequence[Any] | None = None, elements: Sequence[str] | None = None)[source]¶
The linked SpatialData object (
spatialdataneeded).cells(ChromData cell ids) subsets the annotating table and the elements it annotates to those cells (spatialdata.match_sdata_to_table);elementsreads only the named elements (plus the tables).
- classmethod read(path: str | Path, *, backed: bool = False, original_order: bool = False, columns=None, tracks=None) ChromData[source]¶
Read a ChromData file; the format follows the path.
.chromdata.zarr/.cdz(format 2.x container) — withbacked=Truereturns aBackedChromData: small tables are loaded,coords/spots/spot_tracks/layersstay on disk, andget_cell/get_trace/get_chrom/iter_traces/iter_cellsread only the rows they need. Spots come back in the stored order — chromosome › cell › trace › bin (format 2.1; cell › trace › bin for 2.0 stores);original_order=Truerestores the order of the object that was written (in-memory reads only).columns/tracksselect the spot-aligned columns (seeget_cell()): an in-memory read loads only those, a backed object uses them as the default ofget_*/iter_*. The default reads everything..h5cd— HDF5, format 1.x or 2.0 (below); always in memory.
HDF5
.h5cd: dispatches to a version-specific reader based onf.attrs['uchrom_format_version']: MAJOR 2 →_read_v2(), MAJOR 1 →_read_v1(), which upgrades the file in memory (bins derived from the spot loci; the spot-aligned 1.xtrackssplit intobin_tracks/spot_tracks). Files written before versioning was introduced are read with a warning as 1.0.Forward compatibility:
Same MAJOR, higher MINOR → read with a warning; unknown fields are ignored silently by the lower-level helpers.
Different MAJOR → raise
ValueErrorwith guidance.
- rebuild_bins() ChromData[source]¶
Re-derive
bins/bin_idfrom the spots’chrom/start/endafter they were edited in place. Bin-level tracks are carried over for loci that still exist;binmmust be empty. Returnsself.
- resolve_link_path(path: str | Path) Path[source]¶
Linked-file path; relative paths resolve against the directory of the
.h5cdthis object was read from (else the working directory), so a dataset folder can be moved or shipped as a whole.
- property results: ResultsStore¶
Analysis outputs — a
ResultsStore.Behaves like a
dictof values; assigning a plaindictconverts it. Values that cannot be written to.h5cdraiseTypeErroron assignment.
- set_cell_positions(positions=None, *, key: str = 'default', cell_ids: Sequence[Any] | None = None, columns: Sequence[str] | None = None, unit: str | None = None, frame: str | None = None, region_col: str | None = None, in_coords_frame: bool | None = None, status: str = 'measured', source: str | None = None, method: str | None = None, shapes: str | None = None, **metadata: Any) dict[source]¶
Store where each cell is (its centroid, 2-D or 3-D) in
cells.Cell positions describe cells, not chromatin spots: they are
cellscolumns —centroid_x,centroid_y[,centroid_z] forkey="default",<key>_centroid_x… otherwise — described by the recorduns['cell_spatial'][key](schema inuchrom/core/spec.md). Read them back withcell_positions().- Parameters:
positions ((n, 2) / (n, 3) array or DataFrame, optional) – Centroids. A DataFrame uses its
x, y[, z]columns (or its first 2 / 3 columns) and, withoutcell_ids, its index as the cell ids.Noneregisters existingcellscolumns given ascolumns(e.g. a loader’scell_center_x_global).cell_ids (sequence, optional) – Cell of each row; rows are aligned onto the
cellsaxis (cells not given get NaN). Without it, rows followcells.columns (sequence of str, optional) –
cellscolumn names to write (or to register).unit (str, optional) – Length unit (
"um","nm","px"…).frame (str, optional) – Coordinate system:
"fov"(local to one field of view — name it withregion_col),"tissue"/"global"(one stitched frame per section / sample) or any other name.region_col (str, optional) –
cellscolumn naming the region / FOV a local frame belongs to.in_coords_frame (bool, optional) – True when the positions share the frame of
cd.coords.status ({"measured", "inferred"}) –
"inferred"(e.g. mapped by Tangram) needssource.source (str, optional) – Where the positions come from and how they were obtained.
method (str, optional) – Where the positions come from and how they were obtained.
shapes (str, optional) – Key of
cell_shapesholding the outlines of these cells.**metadata – Kept in the record, e.g.
voxel_size=[0.103, 0.103, 0.25]andvoxel_unit="um"for positions stored in voxels (unit="voxel";cell_positions(physical=True)applies them).record. (Returns the)
- set_cell_shapes(shapes, *, key: str = 'cell', cell_ids: Sequence[Any] | None = None, unit: str | None = None, frame: str | None = None, in_coords_frame: bool | None = None, status: str = 'measured', source: str | None = None, method: str | None = None, **metadata: Any) dict[source]¶
Store cell outlines (polygons) as
cell_shapes[key].shapes: a GeoDataFrame / GeoSeries (index = cell ids), a DataFrame withcell_id+ WKBgeometry, a mappingcell_id → shape, or a sequence (withcell_ids) of shapely geometries, WKB bytes or(n, 2)/(n, 3)vertex arrays of the exterior ring. Vertex arrays and WKB need no geometry library; shapely / geopandas are only needed to read them back as geometries (cell_shapes_geodataframe()). Common keys:"cell"(segmented cell boundary),"nucleus". Frame / unit / status are recorded inuns['cell_shapes'][key](the same fields asset_cell_positions()). Stored as GeoParquet (tables/cell_shapes/<key>.parquet) in.chromdata.zarr/.cdz. Returns the record.
- set_cell_spatial_coordinates(xy: ndarray, *, key: str = 'default', cell_ids: List[str] | None = None, x_col: str | None = None, y_col: str | None = None, source: str | None = None, method: str | None = None, coordinate_system: str | None = None, inferred: bool = True, **metadata: Any) dict[source]¶
Deprecated: use
set_cell_positions().Kept for 1.x callers:
coordinate_systembecomesframe,inferredthestatus,x_col/y_colthe column names. Columns now default to the schema names (centroid_x…, wasspatial_x…); positions written by older versions stay readable throughcell_positions().
- property spot_tracks: DataFrame¶
Per-observation signals (per-spot IF intensity, seqFISH z-scores …), row-aligned to
spots.
- spots_with_loci() DataFrame[source]¶
A copy of
spotswithchrom/start/endderived frombinsviabin_id(the stable way to get spot loci).
- to_anndata()[source]¶
Export cell-level data as an AnnData object.
Creates an AnnData where each observation is a cell,
obsisself.cells, andobsmisself.cellm. TheXmatrix is left empty (zeros) because ChromData has no cell-by-feature expression matrix. Spot-level RNA-FISH / IF / epigenomic signals, such as Takei 2025 tracks, remain inself.tracksand are not flattened into AnnData.X.- Return type:
- Raises:
ImportError – If anndata is not installed.
- to_dataframe(include_bin_id: bool = False) DataFrame[source]¶
Export as a flat DataFrame:
chrom, start, end, x, y, z(loci derived frombins) followed by the other spot columns.bin_idis left out unlessinclude_bin_id=True.
- to_fofct(path: str | Path, *, cell_table: str | Path | None = None, rna_table: str | Path | None = None, mapping_table: str | Path | None = None, spot_tracks: bool = True, bin_tracks: bool = False, header: Mapping[str, Any] | None = None, cell_value_prefix: str | None = 'rna.') Dict[str, Path][source]¶
Write a 4DN FOF-CT core table, optionally with its companion cell and RNA-spot tables — the inverse of
from_fofct().The core table has the
##/#header lines, thenSpot_ID, Trace_ID, X, Y, Z, Chrom, Chrom_Start, Chrom_End(+Cell_ID/Sub_Cell_ROI_ID/Extra_Cell_ROI_IDwhen present) and every other spot column, one row per spot in the object’s row order.- Parameters:
path (path) – Core table (
.csv).cell_table (path, optional) – Also write
cellsas a4dn_FOF-CT_celltable (Cell_IDfirst;centroid_x/y/z→Cent_ROI_x/y/z,nucleus_area_um2→area(um2),cytoplasm_area_um2→area_cyto(um2), andcell_value_prefixremoved, sorna.Nanog→Nanog).rna_table (path, optional) – Also write
points['rna']as a4dn_FOF-CT_rnatable (Spot_ID, X, Y, Z, RNA_name, Gene_ID, Cell_ID, …).mapping_table (path, optional) – Also write
cell_shapes['cell']as a4dn_FOF-CT_mappingtable (Cell_ID, ROI_Boundaries: the exterior ring as an OME polygon points string,##ROI_Boundaries_Format).spot_tracks (bool) – Append
spot_tracksas extra columns (default). They come back as spot columns fromfrom_fofct().bin_tracks (bool) – Also append
bin_tracks, broadcast to spots.header (mapping, optional) – Header entries to add or override. Defaults come from
uns['fofct_header'](whatfrom_fofct()read), elseFOF-CT_version=v0.1,Table_namespace, andgenome_assembly/XYZ_unitfromuns. The required fields (FOF-CT_version,Table_namespace,genome_assembly,XYZ_unit) are written as##key=value, the others as#key: value.
- Returns:
{"core": path, "cells": path, "rna": path}of the files written.- Return type:
Notes
FOF-CT has no place for
binsrows without spots,cellm,binm,layers,intervals,resultsor otherunskeys; those are not written. Spots withoutspot_idgetSpot_ID = 0 .. n_spots - 1. Categorical extra columns come back as plain columns.
- to_memory(*, columns=None, tracks=None) ChromData[source]¶
An in-memory
ChromData:self(this object already is), or a copy restricted to acolumns/tracksselection.
- to_spatialdata(*, positions_key: str = 'default', shapes_key: str | None = None, spots: bool = True, points: bool = True, cell_table: bool = True, coordinate_system: str | None = None)[source]¶
Export to a
spatialdata.SpatialData(spatialdata + geopandas needed): spots as a 3-D points elementspots(with chrom / start / end / bin_id / trace_id / cell_id columns), eachpoints[key]aspoints_<key>, the cell centroids as pointscell_centroids, the outlinescell_shapes[key]as a shapes elementcell_shapes_<key>, andcells(+cellmin obsm) as the tablecellsannotating the outlines. Spots and cell positions get separate coordinate systems ("spots"and the position frame); the unit and the records are insdata.attrs['uchrom'].
- track_names() Dict[str, List[str]][source]¶
{"bin": [...], "spot": [...]}— names of the stored tracks.
- property tracks: DataFrame | None¶
Deprecated spot-aligned view of all tracks (1.x
cd.tracks).Returns
tracks_spot_view()— bin-level tracks broadcast to spots plus spot-level tracks — orNonewhen there are none. Usebin_tracks/spot_tracksinstead. Assigning a spot-aligned table splits it (columns constant within every bin go tobin_tracks, the rest tospot_tracks); assigning a table withn_binsrows (or index namedbin_id) setsbin_tracks.
- tracks_spot_view(columns: List[str] | None = None) DataFrame[source]¶
All tracks aligned to spots (the 1.x
trackslayout).Bin-level tracks are broadcast through
spots.bin_id; spot-level tracks are appended.columnsselects a subset.
- update_discovery_schema(schema: dict | None = None, **kwargs) dict[source]¶
Deprecated: use
uchrom_discovery.store_discovery_schema(cd, ...).
- property user_annotations: List[dict]¶
User-provided discovery context and analysis constraints.
Annotations are stored in
uns['user_annotations']and are treated as user-supplied priors or constraints by discovery agents, not as validated data evidence.
- validate_discovery_schema(schema: dict | None = None, *, raise_on_error: bool = False) List[str][source]¶
Deprecated: use
uchrom_discovery.validate_discovery_schema(...).
- validate_links(*, check_files: bool = True) List[str][source]¶
Return issues found in linked external modality metadata.
- write(path: str | Path, *, format: str | None = None, coord_dtype: str = 'float64', row_group_rows: int | None = None, compression_level: int | None = None, keep_source_order: bool = True, compression: str | None = 'gzip', compression_opts: int | None = 4, compress_floats: bool = False) None[source]¶
Write to disk; the format follows the path.
<name>.chromdata.zarr(any*.zarr) — the format 2.0 container (format 2.1): a Zarr v3 directory whose large tables are Parquet (seeuchrom.core.zarrcdandspec.md). The spot-aligned tables are partitioned by chromosome and sorted by (cell, trace, bin) within a partition, withindex/offsets, soread()can open the store backed.<name>.cdz— the same tree in one uncompressed zip file.<name>.h5cd— HDF5 format 2.0 (deprecated; readable indefinitely, still written in this release, with aDeprecationWarning). Other suffixes also write HDF5, as before.
format="zarr" | "cdz" | "h5cd"overrides the suffix.Row order. The zarr / cdz writer stores spots sorted by (chromosome,
cell_idcode,trace_idcode,bin_id) — a stable sort, so ties keep their order. Reading returns that order; every spot keeps its values (compare round trips after sorting by a stable key).index/source_rowrecords each spot’s row in the written object, andChromData.read(path, original_order=True)restores it. Objects that are already sorted are written unchanged.Parameters (zarr / cdz)¶
- coord_dtype
"float64"(default) or"float32" Storage dtype of
coordsandlayers(always float64 in memory). float32 halves their size; see the format notes.- row_group_rowsint, optional
Target rows per Parquet row group of the spot-aligned tables (default 16,384). Groups end on cell boundaries within a chromosome partition.
- compression_levelint, optional
zstd level for Parquet and Zarr (default 3).
- keep_source_orderbool
Store
index/source_rowwhen the writer had to sort.
Parameters (h5cd)¶
- compression
"gzip"(default),"lzf"orNone Filter for the large integer / string datasets (
bin_id, categorical codes, string columns, integer tracks). gzip is part of every HDF5 build; lzf ships with h5py. Datasets below 4,096 elements are stored contiguously, uncompressed.- compression_optsint, optional
gzip level (default 4); ignored for
"lzf".- compress_floatsbool
Also compress float datasets (
coords,layers, float tracks). Off by default: on real tracing data it saves only ~15-20 % while making full reads ~4x slower (see the format 2.0 notes indocs/source/guide/chromdata_2_0_design.md).
- classmethod writer(path: str | Path, **kwargs)[source]¶
A streaming
ChromDataWriterfor a.chromdata.zarrstore larger than memory:with ChromData.writer("big.chromdata.zarr", memory_budget="4GB") as w: for chunk in chunks: w.append(chunk) # a ChromData, or coords= / spots= / spot_tracks= cd = ChromData.read("big.chromdata.zarr", backed=True)
UChromProject¶
- class uchrom.core.UChromProject(chromdata: ChromData | None = None, chromdata_path: Path | None = None, links: dict[str, ~typing.Any]=<factory>)[source]¶
Bases:
objectCoordinate a
ChromDataobject with linked MuData and scool files.This wrapper is intentionally lightweight. The data remain in their native files; the project object centralizes links, validation, and descriptive metadata.
Results store¶
Typed, provenance-tracked analysis results — cd.results.
cd.results is a ResultsStore: a MutableMapping from a
result key ("tads.arcfish", "loops.axiswise_f", …) to a
ResultRecord. Mapping access returns the value, so code
written against the 1.x plain dict keeps working:
cd.results["tads.arcfish"] # -> DataFrame (the value)
cd.results.record("tads.arcfish") # -> ResultRecord (value + provenance)
cd.results["my_table"] = df # plain assignment: record without provenance
Analysis functions store their output with ResultsStore.set(),
passing the producing function, its parameters and its inputs.
Serialisation contract¶
Every record kind has a writer and a reader in uchrom.core.cdata.
Assigning a value that cannot be written raises TypeError
at assignment time (not later, in write()):
kind |
value |
|---|---|
table |
|
intervals |
|
array |
|
mapping |
|
scalar |
JSON-compatible scalar or list / tuple (str, int, float, bool, None) |
- class uchrom.core.results.ResultRecord(kind: str, value: Any, params: Dict[str, ~typing.Any]=<factory>, function: str | None = None, uchrom_version: str = <factory>, inputs: Dict[str, ~typing.Any]=<factory>, created_utc: str = <factory>)[source]¶
Bases:
objectOne analysis output plus its provenance.
- value¶
The result itself (what
cd.results[key]returns).- Type:
Any
- params¶
The tuning parameters (
asdict(params)of the caller’s*Paramsdataclass). JSON-compatible.- Type:
- class uchrom.core.results.ResultsStore(data: Mapping[str, Any] | None = None)[source]¶
Bases:
MutableMappingMutableMapping[str, value]backed byResultRecords.Plain item assignment (
store[key] = value) creates a record with an inferred kind and no provenance; analysis functions useset(). Assigning aResultRecordstores it as is.- classmethod coerce(data: Any) ResultsStore[source]¶
Return
dataif it is already a store, else wrap it.
- record(key: str) ResultRecord[source]¶
The full
ResultRecordstored underkey.
- records() Dict[str, ResultRecord][source]¶
A shallow copy of
key -> ResultRecord.
Calling convention helpers¶
Shared plumbing for the ChromData analysis calling convention.
Every cd-level analysis function follows (see
docs/source/guide/chromdata_2_0_design.md, section 3):
def call_x(cd, *, chrom=None, trace_ids=None, cells=None,
params=None, device="auto", key_added="<what>.<method>",
copy=False) -> DataFrame | ChromData
chrom=Noneruns every chromosome and merges the per-chromosome outputs into one table under one key.key_addedis thecd.resultskey (None= do not store).copy=Falsereturns the primary table;copy=Truereturns a newChromDataholding the result and leavescduntouched.
The helpers here implement the parts every caller shares, including the
deprecation path for the 1.x keywords (positional chrom, store=,
result_key=).
- uchrom.core.convention.UNSET: Any = UNSET¶
Sentinel for “argument not passed” (lets us detect deprecated keywords).
- uchrom.core.convention.merge_tables(frames: Sequence[DataFrame], empty: DataFrame) DataFrame[source]¶
Merge per-chromosome outputs into one table.
Empty frames are skipped; a single non-empty frame is returned as is (so a one-chromosome call is identical to the 1.x per-chromosome output); no rows at all →
empty.
- uchrom.core.convention.parse_legacy_call(func_name: str, args: Tuple[Any, ...], *, chrom: Any, key_added: Any, store: Any, result_key: Any, default_key: str, legacy_key: str | None, copy: bool = False, stacklevel: int = 3) Tuple[Any, str | None, bool][source]¶
Map 1.x arguments onto the convention.
Returns
(chrom, key_added, legacy).legacyis True when any deprecated form was used — the caller then keeps the 1.x default key (legacy_key, e.g."tads") so old code that readscd.results["tads"]keeps working.
- uchrom.core.convention.resolve_chroms(cd, chrom: Any | None) List[str][source]¶
chrom=None→ chromosomes that actually have spots, in category order; a string →[chrom]; a sequence →list(chrom).
- uchrom.core.convention.select_spots(cd, *, trace_ids: Iterable | None = None, cells: Iterable | None = None)[source]¶
Subset
cdto the given traces / cells (cditself if neither).
- uchrom.core.convention.selection_inputs(cd, chroms: Sequence[str], trace_ids, cells) dict[source]¶
The
inputsprovenance block shared by the tracing callers.
- uchrom.core.convention.store_and_return(cd, table: DataFrame, *, key_added: str | None, copy: bool, kind: str, function: str, params: Any, inputs: Mapping[str, Any], extra: Sequence[Tuple[str, Any, str]] = (), interval_kind: str | None = None, intervals: Sequence[Tuple[str, DataFrame, str]] = (), bin_tracks: Mapping[str, DataFrame] | None = None)[source]¶
Store
table(andextra(key, value, kind)records) with provenance, then return per the convention.kind="intervals"tables are also exposed ascd.intervals[key_added](the same object, typedinterval_kind).intervalsadds more(key, table, interval_kind)interval tables (e.g. compartment segments) andbin_tracksmaps a track name to a frame withchrom, start, end, valuethat is written tocd.bin_tracks.
Locus axis (bins)¶
The locus axis of ChromData — bins and spots.bin_id.
bins is a DataFrame indexed by bin_id (0 .. n_bins-1) with
columns chrom (category), start / end (int64) and optional
resolution / name / any extra per-locus column. Every spot points
at one bin through spots["bin_id"]; tracing designs with irregular
probes simply get an irregular bin table (one row per probe locus).
Helpers here derive bins from spot loci, map loci onto an existing bin
table, and split a 1.x spot-aligned tracks table into bin-level and
spot-level parts.
- uchrom.core.bins.bins_from_loci(spots: DataFrame) Tuple[DataFrame, ndarray][source]¶
Unique
(chrom, start, end)ofspots→(bins, bin_id).Bins are ordered by chromosome (category order), then start, then end.
- uchrom.core.bins.loci_of(bins: DataFrame, bin_id: ndarray) dict[source]¶
Spot-aligned
chrom(categorical) /start/endcolumns.
- uchrom.core.bins.map_loci_to_bins(spots: DataFrame, bins: DataFrame) ndarray[source]¶
bin_idof every spot locus;-1where the locus is not a bin.
- uchrom.core.bins.normalise_bins(bins: DataFrame) DataFrame[source]¶
Validate a bin table and return it in canonical form (a copy).
- uchrom.core.bins.split_spot_tracks(tracks: DataFrame, bin_id: ndarray, n_bins: int) Tuple[DataFrame, DataFrame, List[str]][source]¶
Split a spot-aligned 1.x
trackstable.Columns whose value is the same for every spot of a bin move to a bin-level table (
n_binsrows; bins without spots are NaN); the rest stay spot-level. Returns(bin_tracks, spot_tracks, order)whereorderis the original column order.
Interval tables¶
Typed genomic interval tables — cd.intervals.
cd.intervals maps a key ("tads.arcfish", "loops.axiswise_f",
…) to an IntervalTable: a pandas.DataFrame whose
attrs["kind"] declares its schema and whose attrs["source_result"]
names the cd.results record that produced it (if any).
kind |
required columns |
examples |
|---|---|---|
domain |
|
TADs, FISHnet domains |
pair |
|
loops |
peak |
|
MACS peaks |
segment |
|
A/B compartment runs |
Tables are validated when assigned. IntervalTable.to_bins() maps
intervals onto a bin table (cd.bins).
- class uchrom.core.intervals.IntervalStore(data=None)[source]¶
Bases:
MutableMappingcd.intervals— validatedkey -> IntervalTable.- add(key: str, value: DataFrame, *, kind: str | None = None, source_result: str | None = None) IntervalTable[source]¶
Store
valueas akindinterval table and return it.
- classmethod coerce(data) IntervalStore[source]¶
- class uchrom.core.intervals.IntervalTable(data=None, index: Axes | None = None, columns: Axes | None = None, dtype: Dtype | None = None, copy: bool | None = None)[source]¶
Bases:
DataFrameA DataFrame of genomic intervals with a declared
kind.- classmethod from_frame(df: DataFrame, kind: str | None = None, source_result: str | None = None) IntervalTable[source]¶
Validate
dfand wrap it (the data is not copied whendfalready is anIntervalTable).
- to_bins(bins: DataFrame, how: str = 'id', column: str | None = None) Series[source]¶
Map intervals onto
bins(cd.bins).- Parameters:
how (
"id" | "bool" | "count") –"id"— row position of the interval with the largest overlap (-1for none);"bool"— any overlap;"count"— number of overlapping intervals. Forpairtables both anchors count.column (str, optional) – With
how="id", return this column of the best interval instead of its row position (NaN / None where no overlap), e.g.column="label"for compartment segments.
Storage: .chromdata.zarr¶
The .chromdata.zarr container (format 2.3): Zarr + Parquet.
A ChromData store is a Zarr v3 group whose large tables are Parquet files
(the SpatialData approach): Parquet for anything with rows (spots, tracks,
cells, traces, bins, intervals, points, result tables), Zarr for
n-dimensional arrays (cellm, binm, the index/ offsets, result
arrays) and JSON attributes for metadata (format version, uns,
provenance, linked files). Layout (2.2):
x.chromdata.zarr/
├── zarr.json root group; attrs["uchrom"] = format metadata
├── tables/ zarr group (attrs: table metadata)
│ ├── coords/chrom=<name>/part-0.parquet
│ │ coordinates + keys, one partition per
│ │ chromosome, sorted by (cell, trace, bin):
│ │ bin_id, trace_id, cell_id (integer codes), x, y, z
│ ├── primary/spots.parquet spot_tracks.parquet layers/<key>.parquet
│ │ every other spot-aligned column, sorted by
│ │ (cell, trace, chromosome, bin), no keys
│ ├── derived/<key>.parquet spot columns that are a function of the
│ │ bin / cell / trace, one row per key code
│ ├── categories/<table>.<column>.parquet values of the coded columns
│ ├── bins.parquet bin_tracks.parquet traces.parquet cells.parquet
│ ├── intervals/<key>.parquet
│ ├── points/<key>.parquet
│ └── cell_shapes/<key>.parquet cell outlines, GeoParquet (cell_id, WKB geometry; 2.3)
├── index/ zarr arrays over the primary order: trace_offsets,
│ │ trace_codes, trace_cells, chrom_trace, cell_offsets,
│ │ cell_codes, row_groups, [source_row]
│ └── coords/ the same over the coordinate table + partition
│ arrays and primary_run (run → primary run)
├── cellm/<key> binm/<key> zarr arrays (zstd)
├── results/<key> zarr groups / arrays; attrs = provenance;
│ tables as <key>/table.parquet
├── uns/ attrs (JSON) + arrays
├── contacts/ attrs: linked .cool / .scool records
└── links/ attrs: linked .h5ad / .h5mu / SpatialData records
Every spot-aligned value is stored once. Each (cell, trace, chromosome)
triple is one run of the primary tables and one run of its chromosome
partition, with the rows in the same order, so the two sides map run by
run. get_chrom(columns="coords") reads one partition, get_cell /
get_trace one primary slice plus their coordinate runs; a full read
gathers the partitions into primary order.
Format 2.1 stores (every spot table partitioned by chromosome,
tables/spots/chrom=<name>/) and 2.0 stores (one tables/spots.parquet
etc., sorted by (cell, trace, bin)) stay readable.
<name>.cdz is the same tree in one uncompressed (ZIP_STORED) zip
archive, so every member can be read in place (Parquet through a byte
range of the archive, Zarr through zarr.storage.ZipStore).
The full specification is in uchrom/core/spec.md; the design
rationale in docs/source/guide/chromdata_2_0_design.md (section 4).
- class uchrom.core.zarrcd.Container(path: str | Path)[source]¶
Bases:
objectRead access to a
.chromdata.zarrdirectory or a.cdzzip.- parquet(rel: str, read_dictionary: Sequence[str] | None = None)[source]¶
pyarrow.parquet.ParquetFileof a member (read_dictionary: columns read as Arrow dictionary arrays).Plain file reads, not memory maps: mapped pages would count as resident memory in backed mode. A
.cdzmember (stored uncompressed) is read in place through a byte-range view of the archive.
- uchrom.core.zarrcd.DEFAULT_CACHE_BYTES = 268435456¶
decoded row groups kept by backed readers (LRU, bytes)
- uchrom.core.zarrcd.DEFAULT_ROW_GROUP_ROWS = 65536¶
target rows per Parquet row group of the cell-sorted spot tables (row groups end on cell boundaries)
- uchrom.core.zarrcd.DEFAULT_ZSTD_LEVEL = 3¶
zstd level for Parquet and Zarr
- class uchrom.core.zarrcd.SpotPartitionWriter(tpath: Path, tables: Sequence[str], *, level: int, row_group_rows: int, has_cells: bool, layer_names: Mapping[str, str], flat: bool = False, codec: str = 'zstd', split_chrom: bool = False, prefix: str = '', dict_floats: bool = False)[source]¶
Bases:
objectWrites the spot-aligned tables of a format 2.1 store, one chromosome partition at a time, and builds
index/as it goes.Rows are handed over already sorted by (cell, trace, bin) within the partition, in any number of
add()calls. Row groups end on unit boundaries (cell runs; (cell, trace) runs withoutcell_id): a group is cut at the first unit boundary where it holds at leastrow_group_rowsrows — the same groups however the rows are split into calls, so the in-memory writer and the streaming writer produce identical files.- add(tables: Mapping[str, Any], cell: ndarray, trace: ndarray, chrom: ndarray | None = None) None[source]¶
Append sorted rows:
tables[name](Arrow tables of equal length) with the cell / trace codes of the rows (cell = -1 without cell_id). A flat writer (one file per table, format 2.0 / 2.2 primary) also takes the chromosome code of every row:index/chrom_traceis the run’s chromosome, or -1 for a run over several.
- plain_columns: Dict[str, set]¶
table → integer columns written without a dictionary (decimal- encoded floats whose plan says plain pages are smaller)
- prefix¶
“primary/”)
- Type:
directory of the files under
tables/(format 2.2
- split_chrom¶
addthen needs the chromosome code of every row- Type:
runs are (cell, trace, chromosome) runs (format 2.2 primary)
- class uchrom.core.zarrcd.SpotRows(c: Container, parts: Dict[str, Any], cache_bytes: int = 268435456, *, coords: bool = False)[source]¶
Bases:
objectRow access to the spot-aligned Parquet tables of a store.
Rows are numbered globally: partition after partition (format 2.1: one per chromosome; format 2.0: a single one). Every spot-aligned table has the same row groups (
index/row_groups, global offsets;index/partition_groupsgives each partition’s first group), so a row range maps onto the same groups in each table.Row groups a request covers completely are read in one call and not cached (
get_chromreads a whole partition); partially covered ones go through a small LRU cache of decoded groups bounded in bytes (cache_bytes), from which small selections are copied so a group can be freed.- columns(name: str) List[str][source]¶
Logical column names (the helper columns of decimal-encoded floats left out).
- decode(name: str, table)[source]¶
Stored (physical) table → logical table: decimal-encoded float columns decoded to float64 (bitwise the written values); columns of several at a time in threads.
- is_coords¶
keys + x, y, z, one partition per chromosome); otherwise the primary / 2.0 / 2.1 tables
- Type:
format 2.2 coordinate table (
"coords"
- physical(name: str, columns: Sequence[str] | None) List[str] | None[source]¶
Stored columns holding the logical
columns(None: all).
- read_all(name: str, columns: Sequence[str] | None = None, dest: Mapping[str, ndarray] | None = None, deferred: bool = False, scatter: ndarray | None = None)[source]¶
The whole table (all partitions), decoded, with single-chunk columns.
Row groups are read in parallel — tasks of about
_FULL_READ_TASK_ROWSrows, oneParquetFileper thread — straight into one preallocated buffer per numeric column (from Arrow’s pool, soto_pandasis zero-copy). Decimal-encoded floats (format 2.2): a dictionary-encoded column is read as an Arrow dictionary, whose values alone are decoded before one gather into the buffer; a plain one is copied as packed integers and decoded in place afterwards, in cache-sized blocks in threads. Exception columns whose Parquet statistics say they are all null are not read.destmaps logical columns to float64 / integer arrays (or strided views) ofnrows to fill instead. Non-numeric columns (strings, booleans) and integer columns with nulls are gathered as chunks and combined.Returns an Arrow table of the logical columns (
destcolumns left out: they are in the given arrays) — or, withdeferred, a function returning it once the in-place decoding started in the background is done (the caller can convert other data meanwhile).scatter(int64,nrows, or a function(a, b)→ the output rows of table rowsa:b): rowiof the table goes to rowscatter[i]of the output (format 2.2: coordinate rows → primary order); numeric columns only.
- read_ranges(name: str, ranges: Sequence[Tuple[int, int]], columns: Sequence[str] | None = None)[source]¶
Rows of the given
[start, stop)ranges (ascending, disjoint) of one table, as one Arrow table.Consecutive row groups a range covers completely are read in one call (not cached); partially covered groups go through the cache. Different partitions are read in parallel threads (a cell touches one row group per chromosome).
- uchrom.core.zarrcd.ZARR_LAYOUTS = {'2.0': 'flat', '2.1': 'partitioned', '2.2': 'primary+coords', '2.3': 'primary+coords'}¶
one file per spot-aligned table, sorted by cell › trace › bin; 2.1: partitioned by chromosome; 2.2 / 2.3: coordinates + keys partitioned by chromosome, everything else in cell-sorted primary tables, stored once)
- Type:
MINOR versions of the container this reader understands (2.0
- uchrom.core.zarrcd.ZARR_LAYOUT_VERSION = '2.2'¶
2.3 is the 2.2 layout plus the optional cell outlines (
tables/cell_shapes/<key>.parquet, GeoParquet) and the cell-position records (uns['cell_spatial'],uns['cell_shapes'],linksattrslinked_spatialdata) — additive, so 2.2 readers read 2.3 stores (with a warning) and ignore the new parts- Type:
the spot-table layout the current version writes
- uchrom.core.zarrcd.container_kind(path: str | Path) str | None[source]¶
"zarr"/"cdz"/"h5cd"for a ChromData path, elseNone.Suffix first (
.chromdata.zarror any.zarrdirectory,.cdz,.h5cd); an existing path without a known suffix is sniffed (a directory withzarr.json, a zip archive, an HDF5 file).
- uchrom.core.zarrcd.partition_dirs(names: Sequence[str | None]) List[str][source]¶
chrom=<name>directory of every partition (portable names; others, and case-insensitive clashes, becomechrom=_<i>).
- uchrom.core.zarrcd.read_small_parts(c: Container) Dict[str, Any][source]¶
Everything that is loaded eagerly — also in backed mode.
- uchrom.core.zarrcd.sort_order(cell: ndarray, trace: ndarray, bin_id: ndarray, chrom: ndarray | None = None, *, chrom_minor: bool = False) ndarray | None[source]¶
Stable permutation sorting spots by ([chrom,] cell, trace, bin), or
Nonewhen they already are sorted.chrom(the chromosome code of every spot) is the format 2.1 partition key; without it the order is the 2.0 one.chrom_minorsorts by (cell, trace, chrom, bin) instead — the format 2.2 primary order, in which every (cell, trace, chromosome) run is contiguous.
- uchrom.core.zarrcd.write_zarr(cd, path: str | Path, *, coord_dtype: str | dtype = 'float64', row_group_rows: int | None = None, compression_level: int = 3, keep_source_order: bool = True, float_encoding=False, coords_row_group_rows: int = 16384, normalize_spot_columns: bool = True, float_dictionary: bool = True, compression='zstd', _layout: str = '2.2') None[source]¶
Write
cdas a.chromdata.zarrdirectory, or a.cdzzip whenpathends with.cdz. SeeChromData.write()._layoutwrites an older layout (“2.0”: one cell-sorted file per spot-aligned table; “2.1”: every spot table partitioned by chromosome; kept for compatibility tests and benchmarks).float_encoding,coords_row_group_rowsandnormalize_spot_columnsapply to format 2.2.
Backed mode¶
Backed (out-of-core) ChromData over a .chromdata.zarr / .cdz store.
ChromData.read(path, backed=True) returns a BackedChromData:
eager (read on open):
bins,bin_tracks,binm,cells,cellm,traces,intervals,points,results,unsand theindex/offsets — all small;lazy:
coords,layers,spots,spot_tracksare proxies that read Parquet row groups on access.
Subsetting — get_cell / get_trace / get_chrom / cd[rows] /
iter_traces / iter_cells — reads only the row ranges it needs (the
spots are sorted by cell, trace, bin; index/ holds the offsets) and
returns an ordinary in-memory ChromData.
BackedChromData.to_memory() loads everything.
The backed object is read-only for the spot-aligned data (coords,
spots, spot_tracks, layers); the small tables are ordinary
in-memory objects that can be changed, and write() (which loads the
data first) saves them.
- class uchrom.core.backed.BackedArray(rows: SpotRows, name: str)[source]¶
Bases:
object(n_spots, 3)coordinates read from Parquet on access.Supports
len,shape/ndim/dtype, row indexing (slices, integer or boolean arrays, optionally followed by a column index) andnp.asarray(a full read). Values are float64, like in-memory coords.- dtype = dtype('float64')¶
- ndim = 2¶
- class uchrom.core.backed.BackedChromData(path, *, columns=None, tracks=None, cache_bytes: int | None = None, _container: Container | None = None)[source]¶
Bases:
ChromDataA
ChromDatawhose spot-aligned data stays on disk.Created by
ChromData.read(path, backed=True)on a.chromdata.zarrdirectory or.cdzfile. See the module docstring for what is eager and what is lazy.columns/tracks(seeChromData.get_cell()) set the default column selection ofget_*,iter_*,cd[rows]andto_memory(); every one of those also takes them per call.- property backed: bool¶
Truefor aBackedChromData.
- compute_distances(trace_id=None) ndarray[source]¶
Compute pairwise Euclidean distance matrix.
- Parameters:
trace_id (optional) – If given, compute only for spots in that trace. If None, compute for all spots (use with caution on large data).
- Return type:
np.ndarray, shape (n, n)
- property coords¶
- get_cell(cell_id, *, columns=None, tracks=None) ChromData[source]¶
The spots of one cell (a new object).
columns/tracksselect the spot-aligned columns to keep —columns="coords"keeps the coordinates and the key columns only,columns=[...]adds the named spot columns / spot tracks / layers,tracks=[...]picks spot tracks (seeresolve_selection()). The default keeps everything. On a backed object the selection decides what is read from disk.
- get_chrom(chrom: str, *, columns=None, tracks=None) ChromData[source]¶
The spots of one chromosome (see
get_cell()forcolumns/tracks).
- get_trace(trace_id, *, columns=None, tracks=None) ChromData[source]¶
The spots of one trace (see
get_cell()forcolumns/tracks).
- property has_coords_table: bool¶
coordinates + keys partitioned by chromosome, the other spot-aligned columns in cell-sorted primary tables.
- Type:
Truefor a format 2.2 store
- iter_cells(batch=64, *, columns=None, tracks=None) Iterator[ChromData][source]¶
Yield in-memory
ChromDatachunks ofbatchcells (cell code order; the rows of a chunk in stored order, seeiter_traces()).
- iter_traces(batch=1024, *, chrom=None, columns=None, tracks=None) Iterator[ChromData][source]¶
Yield in-memory
ChromDatachunks ofbatchtraces.Traces are visited in the order a
.chromdata.zarrstore holds them — chromosome › cell › trace, with the spots of a trace in bin order; a trace id shared by several cells (or a trace spanning several chromosomes) counts once per (chromosome, cell) run. A backed object yields identical chunks while reading one row range per chunk.- Parameters:
batch (int or
"auto") – Traces per chunk;"auto"sizes chunks fromuchrom.settings.memory_budget.chrom (str, optional) – Only the traces of this chromosome.
columns – Column selection of the chunks (see
get_cell()), e.g.columns="coords"for coordinates and keys only.tracks – Column selection of the chunks (see
get_cell()), e.g.columns="coords"for coordinates and keys only.
- rebuild_bins()[source]¶
Re-derive
bins/bin_idfrom the spots’chrom/start/endafter they were edited in place. Bin-level tracks are carried over for loci that still exist;binmmust be empty. Returnsself.
- property spot_tracks: DataFrame¶
Per-observation signals (per-spot IF intensity, seqFISH z-scores …), row-aligned to
spots.
- property spots¶
- spots_with_loci() DataFrame[source]¶
A copy of
spotswithchrom/start/endderived frombinsviabin_id(the stable way to get spot loci).
- to_anndata()[source]¶
Export cell-level data as an AnnData object.
Creates an AnnData where each observation is a cell,
obsisself.cells, andobsmisself.cellm. TheXmatrix is left empty (zeros) because ChromData has no cell-by-feature expression matrix. Spot-level RNA-FISH / IF / epigenomic signals, such as Takei 2025 tracks, remain inself.tracksand are not flattened into AnnData.X.- Return type:
- Raises:
ImportError – If anndata is not installed.
- to_dataframe(include_bin_id: bool = False) DataFrame[source]¶
Export as a flat DataFrame:
chrom, start, end, x, y, z(loci derived frombins) followed by the other spot columns.bin_idis left out unlessinclude_bin_id=True.
- to_memory(*, columns=None, tracks=None) ChromData[source]¶
Load everything (or the selected columns) into an in-memory
ChromData(the small tables are shared with this object, not copied).
- class uchrom.core.backed.BackedFrame(owner: BackedChromData, name: str, columns: List[str])[source]¶
Bases:
objectA spot-aligned table (
spotsorspot_tracks) read on access.frame[col]reads one column (apandas.Series);frame[[cols]]several;frame.iloc[rows]rows (aDataFrame);to_pandas()everything.spotsalso serves the derivedchrom/start/endcolumns (frombinsviabin_id).- property iloc: _ILoc¶
- property index: RangeIndex¶