uchrom.core

class uchrom.core.UChromProject(chromdata: ChromData | None = None, chromdata_path: Path | None = None, links: dict[str, ~typing.Any]=<factory>)[source]

Bases: object

Coordinate a ChromData object with linked MuData and scool files.

This wrapper is intentionally lightweight. The data remain in their native files; the project object centralizes links, validation, and descriptive metadata.

chromdata: ChromData | None = None
chromdata_path: Path | None = None
describe() → dict[source]
classmethod from_chromdata(cdata_or_path: ChromData | str | Path) → UChromProject[source]
load_mudata(*, key: str = 'default')[source]
load_scool(*, key: str = 'default', cell: str | None = None)[source]
load_spatialdata(*, key: str = 'default', cells=None)[source]
set_cell_positions(positions=None, *, key: str = 'default', **metadata: Any) → dict[source]
set_cell_spatial_coordinates(xy, *, key: str = 'default', **metadata: Any) → dict[source]

Deprecated: use set_cell_positions().

validate(*, check_files: bool = True) → list[str][source]

ChromData

class uchrom.core.ChromData(coords: ndarray, spots: DataFrame, *, bins: DataFrame | None = None, cells: DataFrame | None = None, cellm: Dict[str, ndarray] | None = None, tracks: DataFrame | None = None, spot_tracks: DataFrame | None = None, binm: Dict[str, ndarray] | None = None, intervals: Mapping[str, DataFrame] | None = None, traces: DataFrame | None = None, layers: Dict[str, ndarray] | None = None, results: dict | None = None, uns: dict | None = None, points: Dict[str, DataFrame] | None = None, cell_shapes: Dict[str, DataFrame] | None = None, linked_adata=None, validate: bool = True, _own_spots: bool = False)[source]

Bases: object

Chromatin Data — the core container for U-Chrom.

See the module docstring of uchrom.core.cdata for the full purpose, hierarchy, FOF-CT mapping, and on-disk format contract. Summary:

  • The central abstraction is a “structure table” — genomic bins mapped to 3D coordinates. Each row of spots (with the corresponding row of coords) is one Spot.

  • Spots are grouped hierarchically as Cell → Trace → Spot. A Trace is an ordered chromatin-fibre polymer; a Cell contains one or more traces.

  • All analysis in U-Chrom consumes or produces ChromData. Reconstruction modules (uchrom.recon) emit it, structure callers (uchrom.strc) decorate cd.results[...] with TADs / loops / compartments, and the browser / plotters render it.

Parameters:
  • coords (ndarray, shape (n_spots, 3)) – 3-D coordinates (x, y, z) per spot.

  • spots (DataFrame, shape (n_spots, ≥2)) – Per-spot metadata. Required: trace_id (int or str, will be categorified) and either bin_id (with bins=) or the locus columns chrom / start (0-based) / end (non-inclusive), from which bins and bin_id are derived. Optional: cell_id (int or str), spot_id, FOF-CT sub_cell_roi_id / extra_cell_roi_id, and any experiment-specific annotation column (carried through verbatim).

  • bins (DataFrame, optional) – Locus axis (chrom, start, end + optional resolution, name, …); row i is bin_id i. Derived from the spot loci when omitted (ordered by chromosome, start, end).

  • cells (DataFrame, optional) – Per-cell metadata indexed by cell_id.

  • cellm (dict[str, ndarray], optional) – Per-cell multi-dimensional annotations (embeddings, UMAP, …). Each array’s first axis length = n_cells.

  • tracks (DataFrame, optional) – Per-locus signals (ATAC, ChIP-seq, …): n_bins rows (or an index named bin_id) → bin_tracks. A spot-aligned 1.x table (n_spots rows) is still accepted with a DeprecationWarning and split (columns constant within every bin → bin_tracks, the rest → spot_tracks).

  • spot_tracks (DataFrame, optional) – Per-spot signals (IF intensity, seqFISH z-scores), n_spots rows.

  • binm (dict[str, ndarray], optional) – Per-locus multi-dimensional arrays, first axis n_bins.

  • intervals (dict[str, DataFrame], optional) – Typed interval tables (see uchrom.core.intervals).

  • traces (DataFrame, optional) – Per-trace metadata indexed by trace_id.

  • layers (dict[str, ndarray], optional) – Alternative coordinate sets, each with shape (n_spots, 3) — e.g. raw / drift-corrected / aligned.

  • points (dict[str, DataFrame], optional) – Named sets of non-genomic 3-D points in the same frame as coords — e.g. nascent RNA spots or IF puncta. Each DataFrame needs x, y, z; cell_id (if present) links points to cells and is used when subsetting.

  • cell_shapes (dict[str, DataFrame], optional) – Cell outlines: cell_id + geometry (WKB bytes), one row per cell (see set_cell_shapes()); subset with the cells.

  • results (dict or ResultsStore, optional) – Analysis outputs, stored as a ResultsStore (cd.results[key] returns the value, cd.results.record(key) the value plus provenance). Keys follow "<what>.<method>": 'tads.arcfish', 'loops.axiswise_f', 'compartments.axes_pc', ….

  • uns (dict, optional) – Unstructured metadata preserved on disk. Conventional keys: 'genome_assembly', 'xyz_unit', 'fofct_header'. Auto-discovery context may use 'dataset_references' for source papers/repositories and 'user_annotations' for user-provided priors, constraints, or hypothesis seeds.

  • validate (bool) – If True (default), validate internal consistency on construction (coords shape, spots required columns, tracks / layers alignment).

n_spots, n_traces, n_cells, chroms
Type:

derived accessors

Key methods
-----------
from_dataframe, from_fofct, read, write, to_dataframe,
get_cell, get_trace, get_chrom, compute_distances
On-disk format — ``.h5cd``
--------------------------
Versioned HDF5.  See :mod:`uchrom.core.cdata` module docstring for
the full layout and the :meth:`read` / :meth:`write` round-trip
contract.

Notes

  • Subsetting (cd[mask], get_chrom etc.) always returns a new ChromData; the source is not mutated.

  • Global pairwise distance matrices are intentionally not stored — they are biologically meaningful per-trace, not across cells, and would be O(n²) memory. Compute on demand via compute_distances(trace_id=...)().

  • String columns (chrom, trace_id, cell_id) are auto-converted to pd.Categorical for ~10× memory savings.

add_reference(*, reference_id: str | None = None, role: str = 'user_supplied_reference', title: str | None = None, doi: str | None = None, pmid: str | None = None, url: str | None = None, year: int | None = None, notes: str | None = None, **extra: Any) → dict[source]

Add a source reference to uns['dataset_references'].

Parameters are intentionally metadata-oriented rather than tied to one publication database. role should describe how the reference relates to the dataset, for example 'primary_dataset_paper', 'data_repository', 'supplementary_table', or 'related_biology_prior'.

add_user_annotation(*, annotation_id: str | None = None, scope: str, text: str, target: str | None = None, tags: List[str] | None = None, confidence: str = 'user_asserted', **extra: Any) → dict[source]

Add a user annotation to uns['user_annotations'].

Use annotations for cell-type notes, marker priors, analysis constraints, hypothesis seeds, negative constraints, field semantics, or quality warnings. They are surfaced in the discovery schema and agent context but still require notebook validation before becoming evidence.

property auto_discovery_schema: dict

use uchrom_discovery.get_discovery_schema(cd).

Type:

Deprecated

property backed: bool

True for a BackedChromData.

property bin_tracks: DataFrame

Per-locus signals (bulk ATAC, ChIP, GC, annotation features, compartment scores …), n_bins rows indexed by bin_id.

This is the table the 2.0 design calls tracks; during the 2.x series cd.tracks remains the deprecated spot-aligned view.

property bins: DataFrame

DataFrame indexed by bin_id with chrom (category), start, end and optional per-locus columns (resolution, name, …). Shared by all cells; subsetting spots keeps every bin.

Type:

Locus axis

build_discovery_schema(*, store: bool = True, **kwargs) → dict[source]

Deprecated: use uchrom_discovery.build_discovery_schema(cd, ...).

cell_positions(key: str = 'default', *, physical: bool = False) → DataFrame[source]

Cell centroids as a DataFrame indexed by cell_id: x, y [, z] (float, NaN where unknown) [+ region]; the record (unit, frame, status, …) is in .attrs['cell_spatial'].

Reads the record uns['cell_spatial'][key] (1.x records and their spatial_x / spatial_y columns included). Without a record, key="default" falls back to known column names — centroid_x/y/z, then the older x_centroid/y_centroid/z_centroid and spatial_x/spatial_y — with status unknown (attrs['cell_spatial']['unrecorded'] is True). KeyError when there is no position set. physical=True converts positions stored in voxels (record voxel_size / voxel_unit) to physical units.

cell_shapes_geodataframe(key: str = 'cell')[source]

cell_shapes[key] as a geopandas.GeoDataFrame indexed by cell_id (needs shapely + geopandas; the record is in .attrs['cell_shapes']).

property chroms: List[str]
compute_distances(trace_id=None) → ndarray[source]

Compute pairwise Euclidean distance matrix.

Parameters:

trace_id (optional) – If given, compute only for spots in that trace. If None, compute for all spots (use with caution on large data).

Return type:

np.ndarray, shape (n, n)

coord_status() → dict[source]

Return a compact status summary for coords.

copy() → ChromData[source]
property dataset_references: List[dict]

Dataset-level source references used as auto-discovery priors.

References are stored in uns['dataset_references'] and round-trip with .h5cd files. They are intended for primary dataset papers, data repositories, supplementary tables, method papers, and related biological priors.

describe_for_agent(*, max_items: int = 40) → str[source]

Deprecated: use uchrom_discovery.describe_for_agent(cd).

Summarize linked external modality records.

property discovery_schema: dict

use uchrom_discovery.get_discovery_schema(cd).

Type:

Deprecated

classmethod from_dataframe(df: DataFrame, *, cell_id=None, **kwargs) → ChromData[source]

Create from a reconstruction output DataFrame.

Expects columns: chrom, start, end, x, y, z. Each chromosome becomes one trace.

Parameters:
  • df (DataFrame with columns chrom, start, end, x, y, z.)

  • cell_id (hashable, optional) – If given, tag every spot with this cell identifier (e.g. derived from the output filename for single-cell reconstruction). The DataFrame’s own cell_id column, if any, takes precedence.

  • **kwargs – Forwarded to the ChromData constructor.

classmethod from_fofct(core_path: str | Path, *, cell_table: str | Path | None = None, rna_table: str | Path | None = None, mapping_table: str | Path | None = None, cell_value_prefix: str | None = 'rna.', out: str | Path | None = None, chunksize: int | str = 'auto', memory_budget=None, **kwargs) → ChromData[source]

Read a FOF-CT core table, optionally with its companion tables.

Parameters:
  • core_path (path) – Path to the FOF-CT core table (CSV/TSV/TXT).

  • cell_table (path, optional) – FOF-CT cell data table (4dn_FOF-CT_cell). Loaded into cells indexed by cell_id: cent_ROI_x/y become centroid_x/y, area(um2) / area_cyto(um2) become nucleus_area_um2 / cytoplasm_area_um2, and every other numeric column (per-cell RNA copy numbers in 4DN tables) is prefixed with cell_value_prefix (Nanog → rna.Nanog). Centroid columns become the cell positions (uns['cell_spatial'], cell_positions()): Cent_ROI_x/y[/z] (as centroid_x/y[/z]) or cell_center_x_global / cell_center_y_global (Liu et al. 2025, global frame), in the table’s XYZ_unit.

  • mapping_table (path, optional) – FOF-CT Cell/ROI mapping table (4dn_FOF-CT_mapping): ROI_Boundaries of each Cell_ID in global coordinates, as OME polygon points strings ("x1,y1 x2,y2 …") → cell_shapes['cell'] (uns['cell_shapes']['cell']: frame "global"). Rows keyed by sub- / extra-cell ROI ids and OBJ meshes are skipped with a warning.

  • rna_table (path, optional) – FOF-CT RNA spot table (4dn_FOF-CT_rna). Loaded into points['rna'] with columns x, y, z, gene, gene_id, cell_id, spot_id, … in the same frame as coords.

  • cell_value_prefix (str or None) – Prefix for unrecognised numeric cell-table columns; None keeps their names.

  • out (path, optional) – Streaming import: read the core table in chunks of chunksize rows and write them to this .chromdata.zarr store with ChromDataWriter; returns the store opened backed. Peak memory is one chunk plus the writer’s merge pieces (bounded by memory_budget), not the table. The store equals from_fofct(core_path, ...).write(out) (same spots, order, dtypes, small tables). Without out the table is read into memory (a ResourceWarning is issued when the file is larger than a fifth of the memory budget).

  • chunksize (int or "auto") – Rows per chunk of a streaming import ("auto": from the memory budget).

  • memory_budget (int or str, optional) – Defaults to uchrom.settings.memory_budget.

  • **kwargs – Additional keyword arguments passed to ChromData constructor (e.g. cells, tracks, uns). A streaming import accepts cells, cellm, traces, points and uns.

classmethod from_pyhim_trace(ecsv_path: str | Path, barcode_dict: dict | DataFrame | None = None, **kwargs) → ChromData[source]

Read a PyHiM chromatin-trace ECSV table into a ChromData.

PyHiM (Devos et al. 2024) emits one ECSV file per trace-building run. Schema (from chromatin_trace_table.py upstream):

Spot_ID, Trace_ID, x, y, z, Chrom, Chrom_Start, Chrom_End, ROI #, Mask_id, Barcode #, label

meta['comments'] carries xyz_unit=... and genome_assembly=....

Parameters:
  • ecsv_path (path) – Path to the ECSV file written by PyHiM.

  • barcode_dict (dict[int, (chrom, start, end)] or DataFrame, optional) – Required when Chrom/Chrom_Start/Chrom_End are empty in the ECSV (PyHiM does not always populate them). As a DataFrame, expects columns barcode, chrom, start, end. If Chrom is populated, barcode_dict is ignored.

  • **kwargs – Additional keyword arguments passed to the ChromData constructor (cells, tracks, uns, …).

Notes

  • Mask_id becomes cell_id (PyHiM convention).

  • ECSV header comments are captured in cd.uns['pyhim']['ecsv_comments'] and any xyz_unit / genome_assembly entries are also promoted to cd.uns directly (matching from_fofct()).

classmethod from_seqfish_multiomics(spot_glob, **kwargs) → ChromData[source]

Load Takei 2025 DNA seqFISH+ cerebellum data.

Thin shim around uchrom.io.seqfish_multiomics.read_seqfish_multiomics(). See that function for the full parameter list.

classmethod from_seqfish_multiomics_linked(spot_glob, **kwargs)[source]

Load linked Takei 2025 DNA tracing + RNA AnnData artifacts.

Thin shim around uchrom.io.seqfish_multiomics.load_seqfish_multiomics_linked(). Returns a ChromData with RNA expression available at cd.linked_adata and can write paired .h5cd / .h5ad files.

classmethod from_takei2025_cerebellum(**kwargs)[source]

Load linked Takei 2025 cerebellum data.

Thin shim around uchrom.io.seqfish_multiomics.load_takei2025_cerebellum(). Returns a ChromData with RNA expression available at cd.linked_adata.

get_cell(cell_id, *, columns=None, tracks=None) → ChromData[source]

The spots of one cell (a new object).

columns / tracks select the spot-aligned columns to keep — columns="coords" keeps the coordinates and the key columns only, columns=[...] adds the named spot columns / spot tracks / layers, tracks=[...] picks spot tracks (see resolve_selection()). The default keeps everything. On a backed object the selection decides what is read from disk.

get_cells(cell_ids, *, columns=None, tracks=None) → ChromData[source]

The spots of several cells (ids compared as strings; unknown ids are ignored). See get_cell() for columns / tracks.

get_chrom(chrom: str, *, columns=None, tracks=None) → ChromData[source]

The spots of one chromosome (see get_cell() for columns / tracks).

get_trace(trace_id, *, columns=None, tracks=None) → ChromData[source]

The spots of one trace (see get_cell() for columns / tracks).

property intervals: IntervalStore

Typed interval tables (TADs, loops, peaks, segments) — see uchrom.core.intervals.

iter_cells(batch=64, *, columns=None, tracks=None)[source]

Yield in-memory ChromData chunks of batch cells (cell code order; the rows of a chunk in stored order, see iter_traces()).

iter_traces(batch=1024, *, chrom=None, columns=None, tracks=None)[source]

Yield in-memory ChromData chunks of batch traces.

Traces are visited in the order a .chromdata.zarr store holds them — chromosome › cell › trace, with the spots of a trace in bin order; a trace id shared by several cells (or a trace spanning several chromosomes) counts once per (chromosome, cell) run. A backed object yields identical chunks while reading one row range per chunk.

Parameters:
  • batch (int or "auto") – Traces per chunk; "auto" sizes chunks from uchrom.settings.memory_budget.

  • chrom (str, optional) – Only the traces of this chromosome.

  • columns – Column selection of the chunks (see get_cell()), e.g. columns="coords" for coordinates and keys only.

  • tracks – Column selection of the chunks (see get_cell()), e.g. columns="coords" for coordinates and keys only.

Import cell-level metadata from an AnnData into this ChromData.

Matches cells by cell_id: each unique value in spots['cell_id'] is looked up in adata.obs (by index, or by the column cell_id_col if given). Matched cells get their adata.obs columns merged into self.cells and their adata.obsm arrays copied into self.cellm.

If self.cells already exists, its row order is preserved and AnnData rows are aligned onto that cell axis. This is important for multi-omics loaders such as Takei 2025, where chromatin tracing coordinates live in coords/spots, RNA/IF signals live in spot-level tracks, and mRNA clustering/UMAP already live in cells / cellm.

Parameters:
  • adata (anndata.AnnData) – The single-cell dataset to link (e.g. scRNA-seq).

  • cell_id_col (str, optional) – Column in adata.obs that holds cell identifiers matching spots['cell_id']. If None, adata.obs.index is used as the key.

  • copy_obs (bool) – If True (default), copy adata.obs columns into self.cells.

  • copy_obsm (bool) – If True (default), copy adata.obsm arrays into self.cellm.

Returns:

Number of cells matched.

Return type:

int

Raises:

KeyError – If spots has no cell_id column.

Record a linked bulk / pseudo-bulk .cool or .mcool in uns['linked_cool'] (the counterpart of link_scool() for one matrix per file). The file stays external.

Record a linked MuData file in uns['linked_mudata'].

The MuData object is kept external. ChromData stores only lightweight provenance so multi-modal cell matrices are not forced into coords / spots.

Record a linked scool contact store in uns['linked_scool'].

Record a linked SpatialData zarr store in uns['linked_spatialdata'].

The store stays external (images, labels, shapes, other tables). The cell map joins its annotating table to ChromData cells: rows whose region_key column equals region are cells, and their instance_key value equals the ChromData cell_id — or the value of the cells column cell_col when the ids differ. With read_attrs (and spatialdata installed) table / region / region_key / instance_key default to the table’s spatialdata_attrs. Relative paths resolve against the directory of the ChromData store (as for link_scool()).

property linked_adata

Linked AnnData object, loaded lazily from uns metadata if possible.

load_linked_anndata(path: str | Path | None = None)[source]

Load, cache, and return the linked AnnData object.

load_linked_mudata(*, key: str = 'default')[source]

Load a linked MuData file by key.

mudata is an optional dependency. A clear ImportError is raised when it is unavailable.

load_linked_scool(*, key: str = 'default', cell: str | None = None)[source]

Load metadata or one cell cooler from a linked scool file.

If cell is None, returns a summary dict with available scool cells. If cell is supplied, returns cooler.Cooler for that cell group.

load_linked_spatialdata(*, key: str = 'default', cells: Sequence[Any] | None = None, elements: Sequence[str] | None = None)[source]

The linked SpatialData object (spatialdata needed).

cells (ChromData cell ids) subsets the annotating table and the elements it annotates to those cells (spatialdata.match_sdata_to_table); elements reads only the named elements (plus the tables).

property n_bins: int
property n_cells: int
property n_spots: int
property n_traces: int
classmethod read(path: str | Path, *, backed: bool = False, original_order: bool = False, columns=None, tracks=None) → ChromData[source]

Read a ChromData file; the format follows the path.

  • .chromdata.zarr / .cdz (format 2.x container) — with backed=True returns a BackedChromData: small tables are loaded, coords / spots / spot_tracks / layers stay on disk, and get_cell / get_trace / get_chrom / iter_traces / iter_cells read only the rows they need. Spots come back in the stored order — chromosome › cell › trace › bin (format 2.1; cell › trace › bin for 2.0 stores); original_order=True restores the order of the object that was written (in-memory reads only). columns / tracks select the spot-aligned columns (see get_cell()): an in-memory read loads only those, a backed object uses them as the default of get_* / iter_*. The default reads everything.

  • .h5cd — HDF5, format 1.x or 2.0 (below); always in memory.

HDF5 .h5cd: dispatches to a version-specific reader based on f.attrs['uchrom_format_version']: MAJOR 2 → _read_v2(), MAJOR 1 → _read_v1(), which upgrades the file in memory (bins derived from the spot loci; the spot-aligned 1.x tracks split into bin_tracks / spot_tracks). Files written before versioning was introduced are read with a warning as 1.0.

Forward compatibility:

  • Same MAJOR, higher MINOR → read with a warning; unknown fields are ignored silently by the lower-level helpers.

  • Different MAJOR → raise ValueError with guidance.

rebuild_bins() → ChromData[source]

Re-derive bins / bin_id from the spots’ chrom / start / end after they were edited in place. Bin-level tracks are carried over for loci that still exist; binm must be empty. Returns self.

Linked-file path; relative paths resolve against the directory of the .h5cd this object was read from (else the working directory), so a dataset folder can be moved or shipped as a whole.

property results: ResultsStore

Analysis outputs — a ResultsStore.

Behaves like a dict of values; assigning a plain dict converts it. Values that cannot be written to .h5cd raise TypeError on assignment.

set_cell_positions(positions=None, *, key: str = 'default', cell_ids: Sequence[Any] | None = None, columns: Sequence[str] | None = None, unit: str | None = None, frame: str | None = None, region_col: str | None = None, in_coords_frame: bool | None = None, status: str = 'measured', source: str | None = None, method: str | None = None, shapes: str | None = None, **metadata: Any) → dict[source]

Store where each cell is (its centroid, 2-D or 3-D) in cells.

Cell positions describe cells, not chromatin spots: they are cells columns — centroid_x, centroid_y [, centroid_z] for key="default", <key>_centroid_x … otherwise — described by the record uns['cell_spatial'][key] (schema in uchrom/core/spec.md). Read them back with cell_positions().

Parameters:
  • positions ((n, 2) / (n, 3) array or DataFrame, optional) – Centroids. A DataFrame uses its x, y[, z] columns (or its first 2 / 3 columns) and, without cell_ids, its index as the cell ids. None registers existing cells columns given as columns (e.g. a loader’s cell_center_x_global).

  • cell_ids (sequence, optional) – Cell of each row; rows are aligned onto the cells axis (cells not given get NaN). Without it, rows follow cells.

  • columns (sequence of str, optional) – cells column names to write (or to register).

  • unit (str, optional) – Length unit ("um", "nm", "px" …).

  • frame (str, optional) – Coordinate system: "fov" (local to one field of view — name it with region_col), "tissue" / "global" (one stitched frame per section / sample) or any other name.

  • region_col (str, optional) – cells column naming the region / FOV a local frame belongs to.

  • in_coords_frame (bool, optional) – True when the positions share the frame of cd.coords.

  • status ({"measured", "inferred"}) – "inferred" (e.g. mapped by Tangram) needs source.

  • source (str, optional) – Where the positions come from and how they were obtained.

  • method (str, optional) – Where the positions come from and how they were obtained.

  • shapes (str, optional) – Key of cell_shapes holding the outlines of these cells.

  • **metadata – Kept in the record, e.g. voxel_size=[0.103, 0.103, 0.25] and voxel_unit="um" for positions stored in voxels (unit="voxel"; cell_positions(physical=True) applies them).

  • record. (Returns the)

set_cell_shapes(shapes, *, key: str = 'cell', cell_ids: Sequence[Any] | None = None, unit: str | None = None, frame: str | None = None, in_coords_frame: bool | None = None, status: str = 'measured', source: str | None = None, method: str | None = None, **metadata: Any) → dict[source]

Store cell outlines (polygons) as cell_shapes[key].

shapes: a GeoDataFrame / GeoSeries (index = cell ids), a DataFrame with cell_id + WKB geometry, a mapping cell_id → shape, or a sequence (with cell_ids) of shapely geometries, WKB bytes or (n, 2) / (n, 3) vertex arrays of the exterior ring. Vertex arrays and WKB need no geometry library; shapely / geopandas are only needed to read them back as geometries (cell_shapes_geodataframe()). Common keys: "cell" (segmented cell boundary), "nucleus". Frame / unit / status are recorded in uns['cell_shapes'][key] (the same fields as set_cell_positions()). Stored as GeoParquet (tables/cell_shapes/<key>.parquet) in .chromdata.zarr / .cdz. Returns the record.

set_cell_spatial_coordinates(xy: ndarray, *, key: str = 'default', cell_ids: List[str] | None = None, x_col: str | None = None, y_col: str | None = None, source: str | None = None, method: str | None = None, coordinate_system: str | None = None, inferred: bool = True, **metadata: Any) → dict[source]

Deprecated: use set_cell_positions().

Kept for 1.x callers: coordinate_system becomes frame, inferred the status, x_col / y_col the column names. Columns now default to the schema names (centroid_x …, was spatial_x …); positions written by older versions stay readable through cell_positions().

property spot_tracks: DataFrame

Per-observation signals (per-spot IF intensity, seqFISH z-scores …), row-aligned to spots.

spots_with_loci() → DataFrame[source]

A copy of spots with chrom / start / end derived from bins via bin_id (the stable way to get spot loci).

to_anndata()[source]

Export cell-level data as an AnnData object.

Creates an AnnData where each observation is a cell, obs is self.cells, and obsm is self.cellm. The X matrix is left empty (zeros) because ChromData has no cell-by-feature expression matrix. Spot-level RNA-FISH / IF / epigenomic signals, such as Takei 2025 tracks, remain in self.tracks and are not flattened into AnnData.X.

Return type:

anndata.AnnData

Raises:

ImportError – If anndata is not installed.

to_dataframe(include_bin_id: bool = False) → DataFrame[source]

Export as a flat DataFrame: chrom, start, end, x, y, z (loci derived from bins) followed by the other spot columns. bin_id is left out unless include_bin_id=True.

to_fofct(path: str | Path, *, cell_table: str | Path | None = None, rna_table: str | Path | None = None, mapping_table: str | Path | None = None, spot_tracks: bool = True, bin_tracks: bool = False, header: Mapping[str, Any] | None = None, cell_value_prefix: str | None = 'rna.') → Dict[str, Path][source]

Write a 4DN FOF-CT core table, optionally with its companion cell and RNA-spot tables — the inverse of from_fofct().

The core table has the ## / # header lines, then Spot_ID, Trace_ID, X, Y, Z, Chrom, Chrom_Start, Chrom_End (+ Cell_ID / Sub_Cell_ROI_ID / Extra_Cell_ROI_ID when present) and every other spot column, one row per spot in the object’s row order.

Parameters:
  • path (path) – Core table (.csv).

  • cell_table (path, optional) – Also write cells as a 4dn_FOF-CT_cell table (Cell_ID first; centroid_x/y/z → Cent_ROI_x/y/z, nucleus_area_um2 → area(um2), cytoplasm_area_um2 → area_cyto(um2), and cell_value_prefix removed, so rna.Nanog → Nanog).

  • rna_table (path, optional) – Also write points['rna'] as a 4dn_FOF-CT_rna table (Spot_ID, X, Y, Z, RNA_name, Gene_ID, Cell_ID, …).

  • mapping_table (path, optional) – Also write cell_shapes['cell'] as a 4dn_FOF-CT_mapping table (Cell_ID, ROI_Boundaries: the exterior ring as an OME polygon points string, ##ROI_Boundaries_Format).

  • spot_tracks (bool) – Append spot_tracks as extra columns (default). They come back as spot columns from from_fofct().

  • bin_tracks (bool) – Also append bin_tracks, broadcast to spots.

  • header (mapping, optional) – Header entries to add or override. Defaults come from uns['fofct_header'] (what from_fofct() read), else FOF-CT_version=v0.1, Table_namespace, and genome_assembly / XYZ_unit from uns. The required fields (FOF-CT_version, Table_namespace, genome_assembly, XYZ_unit) are written as ##key=value, the others as #key: value.

Returns:

{"core": path, "cells": path, "rna": path} of the files written.

Return type:

dict

Notes

FOF-CT has no place for bins rows without spots, cellm, binm, layers, intervals, results or other uns keys; those are not written. Spots without spot_id get Spot_ID = 0 .. n_spots - 1. Categorical extra columns come back as plain columns.

to_memory(*, columns=None, tracks=None) → ChromData[source]

An in-memory ChromData: self (this object already is), or a copy restricted to a columns / tracks selection.

to_spatialdata(*, positions_key: str = 'default', shapes_key: str | None = None, spots: bool = True, points: bool = True, cell_table: bool = True, coordinate_system: str | None = None)[source]

Export to a spatialdata.SpatialData (spatialdata + geopandas needed): spots as a 3-D points element spots (with chrom / start / end / bin_id / trace_id / cell_id columns), each points[key] as points_<key>, the cell centroids as points cell_centroids, the outlines cell_shapes[key] as a shapes element cell_shapes_<key>, and cells (+ cellm in obsm) as the table cells annotating the outlines. Spots and cell positions get separate coordinate systems ("spots" and the position frame); the unit and the records are in sdata.attrs['uchrom'].

track_names() → Dict[str, List[str]][source]

{"bin": [...], "spot": [...]} — names of the stored tracks.

property tracks: DataFrame | None

Deprecated spot-aligned view of all tracks (1.x cd.tracks).

Returns tracks_spot_view() — bin-level tracks broadcast to spots plus spot-level tracks — or None when there are none. Use bin_tracks / spot_tracks instead. Assigning a spot-aligned table splits it (columns constant within every bin go to bin_tracks, the rest to spot_tracks); assigning a table with n_bins rows (or index named bin_id) sets bin_tracks.

tracks_spot_view(columns: List[str] | None = None) → DataFrame[source]

All tracks aligned to spots (the 1.x tracks layout).

Bin-level tracks are broadcast through spots.bin_id; spot-level tracks are appended. columns selects a subset.

update_discovery_schema(schema: dict | None = None, **kwargs) → dict[source]

Deprecated: use uchrom_discovery.store_discovery_schema(cd, ...).

property user_annotations: List[dict]

User-provided discovery context and analysis constraints.

Annotations are stored in uns['user_annotations'] and are treated as user-supplied priors or constraints by discovery agents, not as validated data evidence.

validate_discovery_schema(schema: dict | None = None, *, raise_on_error: bool = False) → List[str][source]

Deprecated: use uchrom_discovery.validate_discovery_schema(...).

Return issues found in linked external modality metadata.

write(path: str | Path, *, format: str | None = None, coord_dtype: str = 'float64', row_group_rows: int | None = None, compression_level: int | None = None, keep_source_order: bool = True, compression: str | None = 'gzip', compression_opts: int | None = 4, compress_floats: bool = False) → None[source]

Write to disk; the format follows the path.

  • <name>.chromdata.zarr (any *.zarr) — the format 2.0 container (format 2.1): a Zarr v3 directory whose large tables are Parquet (see uchrom.core.zarrcd and spec.md). The spot-aligned tables are partitioned by chromosome and sorted by (cell, trace, bin) within a partition, with index/ offsets, so read() can open the store backed.

  • <name>.cdz — the same tree in one uncompressed zip file.

  • <name>.h5cd — HDF5 format 2.0 (deprecated; readable indefinitely, still written in this release, with a DeprecationWarning). Other suffixes also write HDF5, as before.

format="zarr" | "cdz" | "h5cd" overrides the suffix.

Row order. The zarr / cdz writer stores spots sorted by (chromosome, cell_id code, trace_id code, bin_id) — a stable sort, so ties keep their order. Reading returns that order; every spot keeps its values (compare round trips after sorting by a stable key). index/source_row records each spot’s row in the written object, and ChromData.read(path, original_order=True) restores it. Objects that are already sorted are written unchanged.

Parameters (zarr / cdz)

coord_dtype"float64" (default) or "float32"

Storage dtype of coords and layers (always float64 in memory). float32 halves their size; see the format notes.

row_group_rowsint, optional

Target rows per Parquet row group of the spot-aligned tables (default 16,384). Groups end on cell boundaries within a chromosome partition.

compression_levelint, optional

zstd level for Parquet and Zarr (default 3).

keep_source_orderbool

Store index/source_row when the writer had to sort.

Parameters (h5cd)

compression"gzip" (default), "lzf" or None

Filter for the large integer / string datasets (bin_id, categorical codes, string columns, integer tracks). gzip is part of every HDF5 build; lzf ships with h5py. Datasets below 4,096 elements are stored contiguously, uncompressed.

compression_optsint, optional

gzip level (default 4); ignored for "lzf".

compress_floatsbool

Also compress float datasets (coords, layers, float tracks). Off by default: on real tracing data it saves only ~15-20 % while making full reads ~4x slower (see the format 2.0 notes in docs/source/guide/chromdata_2_0_design.md).

classmethod writer(path: str | Path, **kwargs)[source]

A streaming ChromDataWriter for a .chromdata.zarr store larger than memory:

with ChromData.writer("big.chromdata.zarr", memory_budget="4GB") as w:
    for chunk in chunks:
        w.append(chunk)      # a ChromData, or coords= / spots= / spot_tracks=
cd = ChromData.read("big.chromdata.zarr", backed=True)

UChromProject

class uchrom.core.UChromProject(chromdata: ChromData | None = None, chromdata_path: Path | None = None, links: dict[str, ~typing.Any]=<factory>)[source]

Bases: object

Coordinate a ChromData object with linked MuData and scool files.

This wrapper is intentionally lightweight. The data remain in their native files; the project object centralizes links, validation, and descriptive metadata.

chromdata: ChromData | None = None
chromdata_path: Path | None = None
describe() → dict[source]
classmethod from_chromdata(cdata_or_path: ChromData | str | Path) → UChromProject[source]
link_mudata(path: str | Path, *, key: str = 'default', **metadata: Any) → dict[source]
link_scool(path: str | Path, *, key: str = 'default', **metadata: Any) → dict[source]
link_spatialdata(path: str | Path, *, key: str = 'default', **metadata: Any) → dict[source]
links: dict[str, Any]
load_mudata(*, key: str = 'default')[source]
load_scool(*, key: str = 'default', cell: str | None = None)[source]
load_spatialdata(*, key: str = 'default', cells=None)[source]
set_cell_positions(positions=None, *, key: str = 'default', **metadata: Any) → dict[source]
set_cell_spatial_coordinates(xy, *, key: str = 'default', **metadata: Any) → dict[source]

Deprecated: use set_cell_positions().

validate(*, check_files: bool = True) → list[str][source]

Results store

Typed, provenance-tracked analysis results — cd.results.

cd.results is a ResultsStore: a MutableMapping from a result key ("tads.arcfish", "loops.axiswise_f", …) to a ResultRecord. Mapping access returns the value, so code written against the 1.x plain dict keeps working:

cd.results["tads.arcfish"]            # -> DataFrame (the value)
cd.results.record("tads.arcfish")     # -> ResultRecord (value + provenance)
cd.results["my_table"] = df           # plain assignment: record without provenance

Analysis functions store their output with ResultsStore.set(), passing the producing function, its parameters and its inputs.

Serialisation contract

Every record kind has a writer and a reader in uchrom.core.cdata. Assigning a value that cannot be written raises TypeError at assignment time (not later, in write()):

kind

value

table

pandas.DataFrame or pandas.Series

intervals

pandas.DataFrame of genomic intervals (TADs, loops, …)

array

numpy.ndarray (not object dtype)

mapping

dict whose leaves are any of the value types here

scalar

JSON-compatible scalar or list / tuple (str, int, float, bool, None)

class uchrom.core.results.ResultRecord(kind: str, value: Any, params: Dict[str, ~typing.Any]=<factory>, function: str | None = None, uchrom_version: str = <factory>, inputs: Dict[str, ~typing.Any]=<factory>, created_utc: str = <factory>)[source]

Bases: object

One analysis output plus its provenance.

kind

"table" | "intervals" | "array" | "mapping" | "scalar".

Type:

str

value

The result itself (what cd.results[key] returns).

Type:

Any

params

The tuning parameters (asdict(params) of the caller’s *Params dataclass). JSON-compatible.

Type:

dict

function

Fully qualified name of the producing function, e.g. "uchrom.strc.tad.call_tads_by_pval". None for values assigned directly.

Type:

str or None

uchrom_version

Package version that produced the value.

Type:

str

inputs

What the function read, e.g. {"chrom": ["chr1"], "n_traces": 900}.

Type:

dict

created_utc

ISO-8601 UTC timestamp.

Type:

str

created_utc: str
function: str | None = None
inputs: Dict[str, Any]
kind: str
params: Dict[str, Any]
property provenance: dict

Everything except the value, as a JSON-compatible dict.

uchrom_version: str
value: Any
class uchrom.core.results.ResultsStore(data: Mapping[str, Any] | None = None)[source]

Bases: MutableMapping

MutableMapping[str, value] backed by ResultRecord s.

Plain item assignment (store[key] = value) creates a record with an inferred kind and no provenance; analysis functions use set(). Assigning a ResultRecord stores it as is.

classmethod coerce(data: Any) → ResultsStore[source]

Return data if it is already a store, else wrap it.

record(key: str) → ResultRecord[source]

The full ResultRecord stored under key.

records() → Dict[str, ResultRecord][source]

A shallow copy of key -> ResultRecord.

set(key: str, value: Any, *, kind: str | None = None, function: str | None = None, params: Any = None, inputs: Mapping[str, Any] | None = None) → ResultRecord[source]

Store value with provenance and return the record.

params may be a *Params dataclass or a dict.

to_dict() → Dict[str, Any][source]

Plain key -> value dict (the 1.x representation).

uchrom.core.results.infer_kind(value: Any) → str[source]

Record kind for value (intervals is never inferred).

uchrom.core.results.validate_value(value: Any, where: str = 'value') → None[source]

Raise TypeError if value has no .h5cd writer.

Calling convention helpers

Shared plumbing for the ChromData analysis calling convention.

Every cd-level analysis function follows (see docs/source/guide/chromdata_2_0_design.md, section 3):

def call_x(cd, *, chrom=None, trace_ids=None, cells=None,
           params=None, device="auto", key_added="<what>.<method>",
           copy=False) -> DataFrame | ChromData
  • chrom=None runs every chromosome and merges the per-chromosome outputs into one table under one key.

  • key_added is the cd.results key (None = do not store).

  • copy=False returns the primary table; copy=True returns a new ChromData holding the result and leaves cd untouched.

The helpers here implement the parts every caller shares, including the deprecation path for the 1.x keywords (positional chrom, store=, result_key=).

uchrom.core.convention.UNSET: Any = UNSET

Sentinel for “argument not passed” (lets us detect deprecated keywords).

uchrom.core.convention.merge_tables(frames: Sequence[DataFrame], empty: DataFrame) → DataFrame[source]

Merge per-chromosome outputs into one table.

Empty frames are skipped; a single non-empty frame is returned as is (so a one-chromosome call is identical to the 1.x per-chromosome output); no rows at all → empty.

uchrom.core.convention.parse_legacy_call(func_name: str, args: Tuple[Any, ...], *, chrom: Any, key_added: Any, store: Any, result_key: Any, default_key: str, legacy_key: str | None, copy: bool = False, stacklevel: int = 3) → Tuple[Any, str | None, bool][source]

Map 1.x arguments onto the convention.

Returns (chrom, key_added, legacy). legacy is True when any deprecated form was used — the caller then keeps the 1.x default key (legacy_key, e.g. "tads") so old code that reads cd.results["tads"] keeps working.

uchrom.core.convention.resolve_chroms(cd, chrom: Any | None) → List[str][source]

chrom=None → chromosomes that actually have spots, in category order; a string → [chrom]; a sequence → list(chrom).

uchrom.core.convention.select_spots(cd, *, trace_ids: Iterable | None = None, cells: Iterable | None = None)[source]

Subset cd to the given traces / cells (cd itself if neither).

uchrom.core.convention.selection_inputs(cd, chroms: Sequence[str], trace_ids, cells) → dict[source]

The inputs provenance block shared by the tracing callers.

uchrom.core.convention.store_and_return(cd, table: DataFrame, *, key_added: str | None, copy: bool, kind: str, function: str, params: Any, inputs: Mapping[str, Any], extra: Sequence[Tuple[str, Any, str]] = (), interval_kind: str | None = None, intervals: Sequence[Tuple[str, DataFrame, str]] = (), bin_tracks: Mapping[str, DataFrame] | None = None)[source]

Store table (and extra (key, value, kind) records) with provenance, then return per the convention.

kind="intervals" tables are also exposed as cd.intervals[key_added] (the same object, typed interval_kind). intervals adds more (key, table, interval_kind) interval tables (e.g. compartment segments) and bin_tracks maps a track name to a frame with chrom, start, end, value that is written to cd.bin_tracks.

Locus axis (bins)

The locus axis of ChromData — bins and spots.bin_id.

bins is a DataFrame indexed by bin_id (0 .. n_bins-1) with columns chrom (category), start / end (int64) and optional resolution / name / any extra per-locus column. Every spot points at one bin through spots["bin_id"]; tracing designs with irregular probes simply get an irregular bin table (one row per probe locus).

Helpers here derive bins from spot loci, map loci onto an existing bin table, and split a 1.x spot-aligned tracks table into bin-level and spot-level parts.

uchrom.core.bins.bins_from_loci(spots: DataFrame) → Tuple[DataFrame, ndarray][source]

Unique (chrom, start, end) of spots → (bins, bin_id).

Bins are ordered by chromosome (category order), then start, then end.

uchrom.core.bins.empty_bins() → DataFrame[source]
uchrom.core.bins.loci_of(bins: DataFrame, bin_id: ndarray) → dict[source]

Spot-aligned chrom (categorical) / start / end columns.

uchrom.core.bins.map_loci_to_bins(spots: DataFrame, bins: DataFrame) → ndarray[source]

bin_id of every spot locus; -1 where the locus is not a bin.

uchrom.core.bins.normalise_bins(bins: DataFrame) → DataFrame[source]

Validate a bin table and return it in canonical form (a copy).

uchrom.core.bins.split_spot_tracks(tracks: DataFrame, bin_id: ndarray, n_bins: int) → Tuple[DataFrame, DataFrame, List[str]][source]

Split a spot-aligned 1.x tracks table.

Columns whose value is the same for every spot of a bin move to a bin-level table (n_bins rows; bins without spots are NaN); the rest stay spot-level. Returns (bin_tracks, spot_tracks, order) where order is the original column order.

Interval tables

Typed genomic interval tables — cd.intervals.

cd.intervals maps a key ("tads.arcfish", "loops.axiswise_f", …) to an IntervalTable: a pandas.DataFrame whose attrs["kind"] declares its schema and whose attrs["source_result"] names the cd.results record that produced it (if any).

kind

required columns

examples

domain

chrom, start, end

TADs, FISHnet domains

pair

chrom1, start1, end1, chrom2, start2, end2

loops

peak

chrom, start, end (+ summit, score, …)

MACS peaks

segment

chrom, start, end, label

A/B compartment runs

Tables are validated when assigned. IntervalTable.to_bins() maps intervals onto a bin table (cd.bins).

class uchrom.core.intervals.IntervalStore(data=None)[source]

Bases: MutableMapping

cd.intervals — validated key -> IntervalTable.

add(key: str, value: DataFrame, *, kind: str | None = None, source_result: str | None = None) → IntervalTable[source]

Store value as a kind interval table and return it.

classmethod coerce(data) → IntervalStore[source]
to_bins(key: str, bins: DataFrame, **kwargs) → Series[source]
class uchrom.core.intervals.IntervalTable(data=None, index: Axes | None = None, columns: Axes | None = None, dtype: Dtype | None = None, copy: bool | None = None)[source]

Bases: DataFrame

A DataFrame of genomic intervals with a declared kind.

classmethod from_frame(df: DataFrame, kind: str | None = None, source_result: str | None = None) → IntervalTable[source]

Validate df and wrap it (the data is not copied when df already is an IntervalTable).

property kind: str | None
property source_result: str | None
to_bins(bins: DataFrame, how: str = 'id', column: str | None = None) → Series[source]

Map intervals onto bins (cd.bins).

Parameters:
  • how ("id" | "bool" | "count") – "id" — row position of the interval with the largest overlap (-1 for none); "bool" — any overlap; "count" — number of overlapping intervals. For pair tables both anchors count.

  • column (str, optional) – With how="id", return this column of the best interval instead of its row position (NaN / None where no overlap), e.g. column="label" for compartment segments.

uchrom.core.intervals.infer_interval_kind(df: DataFrame) → str[source]
uchrom.core.intervals.validate_intervals(df: DataFrame, kind: str) → None[source]

Storage: .chromdata.zarr

The .chromdata.zarr container (format 2.3): Zarr + Parquet.

A ChromData store is a Zarr v3 group whose large tables are Parquet files (the SpatialData approach): Parquet for anything with rows (spots, tracks, cells, traces, bins, intervals, points, result tables), Zarr for n-dimensional arrays (cellm, binm, the index/ offsets, result arrays) and JSON attributes for metadata (format version, uns, provenance, linked files). Layout (2.2):

x.chromdata.zarr/
├── zarr.json                 root group; attrs["uchrom"] = format metadata
├── tables/                   zarr group (attrs: table metadata)
│   ├── coords/chrom=<name>/part-0.parquet
│   │                         coordinates + keys, one partition per
│   │                         chromosome, sorted by (cell, trace, bin):
│   │                         bin_id, trace_id, cell_id (integer codes), x, y, z
│   ├── primary/spots.parquet  spot_tracks.parquet  layers/<key>.parquet
│   │                         every other spot-aligned column, sorted by
│   │                         (cell, trace, chromosome, bin), no keys
│   ├── derived/<key>.parquet spot columns that are a function of the
│   │                         bin / cell / trace, one row per key code
│   ├── categories/<table>.<column>.parquet   values of the coded columns
│   ├── bins.parquet  bin_tracks.parquet  traces.parquet  cells.parquet
│   ├── intervals/<key>.parquet
│   ├── points/<key>.parquet
│   └── cell_shapes/<key>.parquet  cell outlines, GeoParquet (cell_id, WKB geometry; 2.3)
├── index/                    zarr arrays over the primary order: trace_offsets,
│   │                         trace_codes, trace_cells, chrom_trace, cell_offsets,
│   │                         cell_codes, row_groups, [source_row]
│   └── coords/               the same over the coordinate table + partition
│                             arrays and primary_run (run → primary run)
├── cellm/<key>  binm/<key>   zarr arrays (zstd)
├── results/<key>             zarr groups / arrays; attrs = provenance;
│                             tables as <key>/table.parquet
├── uns/                      attrs (JSON) + arrays
├── contacts/                 attrs: linked .cool / .scool records
└── links/                    attrs: linked .h5ad / .h5mu / SpatialData records

Every spot-aligned value is stored once. Each (cell, trace, chromosome) triple is one run of the primary tables and one run of its chromosome partition, with the rows in the same order, so the two sides map run by run. get_chrom(columns="coords") reads one partition, get_cell / get_trace one primary slice plus their coordinate runs; a full read gathers the partitions into primary order.

Format 2.1 stores (every spot table partitioned by chromosome, tables/spots/chrom=<name>/) and 2.0 stores (one tables/spots.parquet etc., sorted by (cell, trace, bin)) stay readable.

<name>.cdz is the same tree in one uncompressed (ZIP_STORED) zip archive, so every member can be read in place (Parquet through a byte range of the archive, Zarr through zarr.storage.ZipStore).

The full specification is in uchrom/core/spec.md; the design rationale in docs/source/guide/chromdata_2_0_design.md (section 4).

class uchrom.core.zarrcd.Container(path: str | Path)[source]

Bases: object

Read access to a .chromdata.zarr directory or a .cdz zip.

array(name: str) → ndarray[source]
close() → None[source]
exists(rel: str) → bool[source]
group_attrs(name: str) → dict[source]
parquet(rel: str, read_dictionary: Sequence[str] | None = None)[source]

pyarrow.parquet.ParquetFile of a member (read_dictionary: columns read as Arrow dictionary arrays).

Plain file reads, not memory maps: mapped pages would count as resident memory in backed mode. A .cdz member (stored uncompressed) is read in place through a byte-range view of the archive.

prefetch(names: Iterable[str]) → None[source]

Start reading the arrays names in threads (array then waits for them): zarr reads are small but each has a fixed cost.

read_frame(rel: str) → DataFrame[source]
uchrom.core.zarrcd.DEFAULT_CACHE_BYTES = 268435456

decoded row groups kept by backed readers (LRU, bytes)

uchrom.core.zarrcd.DEFAULT_ROW_GROUP_ROWS = 65536

target rows per Parquet row group of the cell-sorted spot tables (row groups end on cell boundaries)

uchrom.core.zarrcd.DEFAULT_ZSTD_LEVEL = 3

zstd level for Parquet and Zarr

class uchrom.core.zarrcd.SpotPartitionWriter(tpath: Path, tables: Sequence[str], *, level: int, row_group_rows: int, has_cells: bool, layer_names: Mapping[str, str], flat: bool = False, codec: str = 'zstd', split_chrom: bool = False, prefix: str = '', dict_floats: bool = False)[source]

Bases: object

Writes the spot-aligned tables of a format 2.1 store, one chromosome partition at a time, and builds index/ as it goes.

Rows are handed over already sorted by (cell, trace, bin) within the partition, in any number of add() calls. Row groups end on unit boundaries (cell runs; (cell, trace) runs without cell_id): a group is cut at the first unit boundary where it holds at least row_group_rows rows — the same groups however the rows are split into calls, so the in-memory writer and the streaming writer produce identical files.

add(tables: Mapping[str, Any], cell: ndarray, trace: ndarray, chrom: ndarray | None = None) → None[source]

Append sorted rows: tables[name] (Arrow tables of equal length) with the cell / trace codes of the rows (cell = -1 without cell_id). A flat writer (one file per table, format 2.0 / 2.2 primary) also takes the chromosome code of every row: index/chrom_trace is the run’s chromosome, or -1 for a run over several.

begin(chrom_code: int, chrom_name: str | None, pdir: str) → None[source]
end() → Dict[str, Any][source]
index() → Dict[str, ndarray][source]
part_chrom: List[int]
partitions: List[Dict[str, Any]]
plain_columns: Dict[str, set]

table → integer columns written without a dictionary (decimal- encoded floats whose plan says plain pages are smaller)

prefix

“primary/”)

Type:

directory of the files under tables/ (format 2.2

rel_path(name: str, pdir: str | None) → str[source]
split_chrom

add then needs the chromosome code of every row

Type:

runs are (cell, trace, chromosome) runs (format 2.2 primary)

write_empty(tables: Mapping[str, Any]) → None[source]

Create the files of the current partition with the schema of tables and no rows (a store without spots).

class uchrom.core.zarrcd.SpotRows(c: Container, parts: Dict[str, Any], cache_bytes: int = 268435456, *, coords: bool = False)[source]

Bases: object

Row access to the spot-aligned Parquet tables of a store.

Rows are numbered globally: partition after partition (format 2.1: one per chromosome; format 2.0: a single one). Every spot-aligned table has the same row groups (index/row_groups, global offsets; index/partition_groups gives each partition’s first group), so a row range maps onto the same groups in each table.

Row groups a request covers completely are read in one call and not cached (get_chrom reads a whole partition); partially covered ones go through a small LRU cache of decoded groups bounded in bytes (cache_bytes), from which small selections are copied so a group can be freed.

clear_cache() → None[source]
columns(name: str) → List[str][source]

Logical column names (the helper columns of decimal-encoded floats left out).

coords_of(table, columns=('x', 'y', 'z')) → ndarray[source]
decode(name: str, table)[source]

Stored (physical) table → logical table: decimal-encoded float columns decoded to float64 (bitwise the written values); columns of several at a time in threads.

empty(name: str, columns: Sequence[str] | None = None)[source]
encoding: Dict[str, Dict[str, Any]]

table → {logical float column → decimal plan} (format 2.2)

files: Dict[str, bool]

table names (“spots”, “spot_tracks”, “layers/<key>”) → True

is_coords

keys + x, y, z, one partition per chromosome); otherwise the primary / 2.0 / 2.1 tables

Type:

format 2.2 coordinate table ("coords"

property n_partitions: int
partition_range(part: int) → Tuple[int, int][source]
pf(name: str, part: int = 0)[source]
physical(name: str, columns: Sequence[str] | None) → List[str] | None[source]

Stored columns holding the logical columns (None: all).

read_all(name: str, columns: Sequence[str] | None = None, dest: Mapping[str, ndarray] | None = None, deferred: bool = False, scatter: ndarray | None = None)[source]

The whole table (all partitions), decoded, with single-chunk columns.

Row groups are read in parallel — tasks of about _FULL_READ_TASK_ROWS rows, one ParquetFile per thread — straight into one preallocated buffer per numeric column (from Arrow’s pool, so to_pandas is zero-copy). Decimal-encoded floats (format 2.2): a dictionary-encoded column is read as an Arrow dictionary, whose values alone are decoded before one gather into the buffer; a plain one is copied as packed integers and decoded in place afterwards, in cache-sized blocks in threads. Exception columns whose Parquet statistics say they are all null are not read. dest maps logical columns to float64 / integer arrays (or strided views) of n rows to fill instead. Non-numeric columns (strings, booleans) and integer columns with nulls are gathered as chunks and combined.

Returns an Arrow table of the logical columns (dest columns left out: they are in the given arrays) — or, with deferred, a function returning it once the in-place decoding started in the background is done (the caller can convert other data meanwhile).

scatter (int64, n rows, or a function (a, b) → the output rows of table rows a:b): row i of the table goes to row scatter[i] of the output (format 2.2: coordinate rows → primary order); numeric columns only.

read_ranges(name: str, ranges: Sequence[Tuple[int, int]], columns: Sequence[str] | None = None)[source]

Rows of the given [start, stop) ranges (ascending, disjoint) of one table, as one Arrow table.

Consecutive row groups a range covers completely are read in one call (not cached); partially covered groups go through the cache. Different partitions are read in parallel threads (a cell touches one row group per chromosome).

take(name: str, rows: ndarray, columns: Sequence[str] | None = None)[source]

Arbitrary rows (any order) of one table.

to_frame(name: str, table) → DataFrame[source]

Arrow table of a spot-aligned file → DataFrame (codes → categoricals).

uchrom.core.zarrcd.ZARR_LAYOUTS = {'2.0': 'flat', '2.1': 'partitioned', '2.2': 'primary+coords', '2.3': 'primary+coords'}

one file per spot-aligned table, sorted by cell › trace › bin; 2.1: partitioned by chromosome; 2.2 / 2.3: coordinates + keys partitioned by chromosome, everything else in cell-sorted primary tables, stored once)

Type:

MINOR versions of the container this reader understands (2.0

uchrom.core.zarrcd.ZARR_LAYOUT_VERSION = '2.2'

2.3 is the 2.2 layout plus the optional cell outlines (tables/cell_shapes/<key>.parquet, GeoParquet) and the cell-position records (uns['cell_spatial'], uns['cell_shapes'], links attrs linked_spatialdata) — additive, so 2.2 readers read 2.3 stores (with a warning) and ignore the new parts

Type:

the spot-table layout the current version writes

uchrom.core.zarrcd.container_kind(path: str | Path) → str | None[source]

"zarr" / "cdz" / "h5cd" for a ChromData path, else None.

Suffix first (.chromdata.zarr or any .zarr directory, .cdz, .h5cd); an existing path without a known suffix is sniffed (a directory with zarr.json, a zip archive, an HDF5 file).

uchrom.core.zarrcd.partition_dirs(names: Sequence[str | None]) → List[str][source]

chrom=<name> directory of every partition (portable names; others, and case-insensitive clashes, become chrom=_<i>).

uchrom.core.zarrcd.read_small_parts(c: Container) → Dict[str, Any][source]

Everything that is loaded eagerly — also in backed mode.

uchrom.core.zarrcd.sort_order(cell: ndarray, trace: ndarray, bin_id: ndarray, chrom: ndarray | None = None, *, chrom_minor: bool = False) → ndarray | None[source]

Stable permutation sorting spots by ([chrom,] cell, trace, bin), or None when they already are sorted. chrom (the chromosome code of every spot) is the format 2.1 partition key; without it the order is the 2.0 one. chrom_minor sorts by (cell, trace, chrom, bin) instead — the format 2.2 primary order, in which every (cell, trace, chromosome) run is contiguous.

uchrom.core.zarrcd.write_zarr(cd, path: str | Path, *, coord_dtype: str | dtype = 'float64', row_group_rows: int | None = None, compression_level: int = 3, keep_source_order: bool = True, float_encoding=False, coords_row_group_rows: int = 16384, normalize_spot_columns: bool = True, float_dictionary: bool = True, compression='zstd', _layout: str = '2.2') → None[source]

Write cd as a .chromdata.zarr directory, or a .cdz zip when path ends with .cdz. See ChromData.write().

_layout writes an older layout (“2.0”: one cell-sorted file per spot-aligned table; “2.1”: every spot table partitioned by chromosome; kept for compatibility tests and benchmarks). float_encoding, coords_row_group_rows and normalize_spot_columns apply to format 2.2.

Backed mode

Backed (out-of-core) ChromData over a .chromdata.zarr / .cdz store.

ChromData.read(path, backed=True) returns a BackedChromData:

  • eager (read on open): bins, bin_tracks, binm, cells, cellm, traces, intervals, points, results, uns and the index/ offsets — all small;

  • lazy: coords, layers, spots, spot_tracks are proxies that read Parquet row groups on access.

Subsetting — get_cell / get_trace / get_chrom / cd[rows] / iter_traces / iter_cells — reads only the row ranges it needs (the spots are sorted by cell, trace, bin; index/ holds the offsets) and returns an ordinary in-memory ChromData. BackedChromData.to_memory() loads everything.

The backed object is read-only for the spot-aligned data (coords, spots, spot_tracks, layers); the small tables are ordinary in-memory objects that can be changed, and write() (which loads the data first) saves them.

class uchrom.core.backed.BackedArray(rows: SpotRows, name: str)[source]

Bases: object

(n_spots, 3) coordinates read from Parquet on access.

Supports len, shape / ndim / dtype, row indexing (slices, integer or boolean arrays, optionally followed by a column index) and np.asarray (a full read). Values are float64, like in-memory coords.

dtype = dtype('float64')
ndim = 2
property shape: Tuple[int, int]
property size: int
to_numpy() → ndarray[source]
class uchrom.core.backed.BackedChromData(path, *, columns=None, tracks=None, cache_bytes: int | None = None, _container: Container | None = None)[source]

Bases: ChromData

A ChromData whose spot-aligned data stays on disk.

Created by ChromData.read(path, backed=True) on a .chromdata.zarr directory or .cdz file. See the module docstring for what is eager and what is lazy.

columns / tracks (see ChromData.get_cell()) set the default column selection of get_*, iter_*, cd[rows] and to_memory(); every one of those also takes them per call.

property backed: bool

True for a BackedChromData.

close() → None[source]

Release the store (the object is unusable afterwards).

compute_distances(trace_id=None) → ndarray[source]

Compute pairwise Euclidean distance matrix.

Parameters:

trace_id (optional) – If given, compute only for spots in that trace. If None, compute for all spots (use with caution on large data).

Return type:

np.ndarray, shape (n, n)

property coords
copy() → ChromData[source]
property format_version: str
get_cell(cell_id, *, columns=None, tracks=None) → ChromData[source]

The spots of one cell (a new object).

columns / tracks select the spot-aligned columns to keep — columns="coords" keeps the coordinates and the key columns only, columns=[...] adds the named spot columns / spot tracks / layers, tracks=[...] picks spot tracks (see resolve_selection()). The default keeps everything. On a backed object the selection decides what is read from disk.

get_chrom(chrom: str, *, columns=None, tracks=None) → ChromData[source]

The spots of one chromosome (see get_cell() for columns / tracks).

get_trace(trace_id, *, columns=None, tracks=None) → ChromData[source]

The spots of one trace (see get_cell() for columns / tracks).

property has_coords_table: bool

coordinates + keys partitioned by chromosome, the other spot-aligned columns in cell-sorted primary tables.

Type:

True for a format 2.2 store

iter_cells(batch=64, *, columns=None, tracks=None) → Iterator[ChromData][source]

Yield in-memory ChromData chunks of batch cells (cell code order; the rows of a chunk in stored order, see iter_traces()).

iter_traces(batch=1024, *, chrom=None, columns=None, tracks=None) → Iterator[ChromData][source]

Yield in-memory ChromData chunks of batch traces.

Traces are visited in the order a .chromdata.zarr store holds them — chromosome › cell › trace, with the spots of a trace in bin order; a trace id shared by several cells (or a trace spanning several chromosomes) counts once per (chromosome, cell) run. A backed object yields identical chunks while reading one row range per chunk.

Parameters:
  • batch (int or "auto") – Traces per chunk; "auto" sizes chunks from uchrom.settings.memory_budget.

  • chrom (str, optional) – Only the traces of this chromosome.

  • columns – Column selection of the chunks (see get_cell()), e.g. columns="coords" for coordinates and keys only.

  • tracks – Column selection of the chunks (see get_cell()), e.g. columns="coords" for coordinates and keys only.

property n_cells: int
property n_spots: int
property n_traces: int
property partitioned: bool

True for a format 2.1 store (spots partitioned by chromosome).

rebuild_bins()[source]

Re-derive bins / bin_id from the spots’ chrom / start / end after they were edited in place. Bin-level tracks are carried over for loci that still exist; binm must be empty. Returns self.

property spot_tracks: DataFrame

Per-observation signals (per-spot IF intensity, seqFISH z-scores …), row-aligned to spots.

property spots
spots_with_loci() → DataFrame[source]

A copy of spots with chrom / start / end derived from bins via bin_id (the stable way to get spot loci).

to_anndata()[source]

Export cell-level data as an AnnData object.

Creates an AnnData where each observation is a cell, obs is self.cells, and obsm is self.cellm. The X matrix is left empty (zeros) because ChromData has no cell-by-feature expression matrix. Spot-level RNA-FISH / IF / epigenomic signals, such as Takei 2025 tracks, remain in self.tracks and are not flattened into AnnData.X.

Return type:

anndata.AnnData

Raises:

ImportError – If anndata is not installed.

to_dataframe(include_bin_id: bool = False) → DataFrame[source]

Export as a flat DataFrame: chrom, start, end, x, y, z (loci derived from bins) followed by the other spot columns. bin_id is left out unless include_bin_id=True.

to_memory(*, columns=None, tracks=None) → ChromData[source]

Load everything (or the selected columns) into an in-memory ChromData (the small tables are shared with this object, not copied).

tracks_spot_view(columns: List[str] | None = None) → DataFrame[source]

All tracks aligned to spots (the 1.x tracks layout).

Bin-level tracks are broadcast through spots.bin_id; spot-level tracks are appended. columns selects a subset.

write(path, **kwargs) → None[source]

Write a copy (the data is loaded into memory first).

class uchrom.core.backed.BackedFrame(owner: BackedChromData, name: str, columns: List[str])[source]

Bases: object

A spot-aligned table (spots or spot_tracks) read on access.

frame[col] reads one column (a pandas.Series); frame[[cols]] several; frame.iloc[rows] rows (a DataFrame); to_pandas() everything. spots also serves the derived chrom / start / end columns (from bins via bin_id).

property columns: Index
property dtypes: Series

Column dtypes, from the Parquet schema (no data is read).

property empty: bool
head(n: int = 5) → DataFrame[source]
property iloc: _ILoc
property index: RangeIndex
property shape: Tuple[int, int]
to_pandas() → DataFrame[source]