uchrom.emb

Cell embeddings: from per-cell features of any modality (embed_cells with count / z-score / TF-IDF-LSI normalisation, aggregate_tracks for per-spot imaging signals), from single-cell contact maps (scHiCluster, uchrom.emb.hic) and from single-cell Hi-C with Higashi (higashi).

Cell embeddings from per-cell features of any modality.

embed_cells() turns a group of numeric cd.cells columns — the rna.<gene> counts written by from_fofct(), if.<mark> immunofluorescence summaries from aggregate_tracks(), atac.<gene> accessibility, … — or an explicit cells × features matrix into PCA (or LSI) / t-SNE / UMAP coordinates stored in cd.cellm (rows aligned with cd.cells). Parameters, including the normalisation, are recorded in cd.uns["embeddings"] so they round-trip through ChromData files and the web browser can group the embeddings by modality.

Normalisations (normalization=; "auto" picks one from source):

"counts" (rna and unknown prefixes)

log1p then z-score per feature (optionally library-size first).

"zscore" (if, protein, ab)

continuous intensities: z-score per feature, log=True adds log1p.

"tfidf" (atac)

sparse accessibility: TF-IDF → truncated SVD (latent semantic indexing); component 1 is dropped when it tracks sequencing depth.

"none"

features used as given (already normalised, e.g. Hi-C vectors).

uchrom.emb.cells.MODALITY_LABELS = {'ab': 'Protein', 'atac': 'ATAC', 'hic': 'Hi-C', 'higashi': 'Hi-C', 'if': 'Histone/IF', 'protein': 'Protein', 'rna': 'RNA'}

display name of each feature prefix (the browser groups embeddings by it)

uchrom.emb.cells.aggregate_tracks(cd, fields: Sequence[str] | None = None, *, prefix: str = 'if', stat: str = 'mean', by: None | str | int = None, min_spots: int = 1, write: bool = True) → DataFrame[source]

Summarise tracks (cd.spot_tracks + cd.bin_tracks seen per spot) into per-cell features.

Parameters:
  • cd (ChromData) – With tracks and a cell_id column in cd.spots.

  • fields (sequence of str, optional) – Track columns to use (default: every numeric track).

  • prefix (str) – Output columns are named f"{prefix}.{field}".

  • stat ("mean", "median", "std" or "corr") – Per-cell statistic. "corr" gives the within-cell Pearson correlation of every pair of fields (f"{prefix}.{a}~{b}") — how marks co-occur on the same loci, which is insensitive to per-cell intensity scaling.

  • by (None, "chrom" or int) – Also split by chromosome (…@chr1) or by genomic bins of this many bp (…@chr1:10 = 10th bin); not with stat="corr".

  • min_spots (int) – Groups with fewer spots become NaN.

  • write (bool) – Add the columns to cd.cells (replacing ones of the same name).

Return type:

DataFrame indexed like cd.cells (cells without spots are NaN).

uchrom.emb.cells.cluster_cells(cd, *, use: str = 'rna_pca', n_neighbors: int = 15, resolution: float = 1.0, n_components: int | None = None, random_state: int = 0, key_added: str | None = None) → Series[source]

Leiden clusters of the cells from an embedding in cd.cellm.

A k-nearest-neighbour graph (n_neighbors, Euclidean, symmetrised) of cd.cellm[use] (first n_components columns) is partitioned with the Leiden algorithm (leidenalg, RB configuration model at resolution, seeded). Clusters are numbered "0", "1", … by decreasing size and stored as a categorical column cd.cells[key_added] (default f"leiden_{use}"); the parameters go to cd.uns["clusters"][key_added]. Needs leidenalg and python-igraph (pip install leidenalg).

Returns the cluster labels (a Series indexed like cd.cells).

uchrom.emb.cells.embed_cells(cd, *, source: str = 'rna', fields: Sequence[str] | None = None, matrix=None, normalization: str = 'auto', log: bool = False, methods: Sequence[str] = ('pca', 'tsne', 'umap'), n_pcs: int = 10, n_lsi: int = 30, drop_depth_r: float | None = 0.5, random_state: int = 0, tsne_perplexity: float | None = None, umap_neighbors: int = 15, umap_min_dist: float = 0.3, library_size: bool = False, modality: str | None = None, features: str | None = None) → Dict[str, ndarray][source]

Compute cell embeddings and store them in cd.cellm.

Parameters:
  • cd (ChromData) – Needs a non-empty cd.cells; rows of the embeddings follow it.

  • source (str) – Feature group; selects the numeric columns f"{source}.*" and names the outputs f"{source}_{method}" (rna_pca, if_umap, atac_lsi, …).

  • fields (sequence of str, optional) – Explicit cd.cells columns to use instead of the prefix.

  • matrix (array or scipy.sparse matrix, optional) – cells × features input instead of cd.cells columns (rows in cd.cells order) — e.g. a genome-wide ATAC bin matrix too large to keep as cell columns. Describe it with features=.

  • normalization ("auto", "counts", "zscore", "tfidf", "none") – See the module docstring. "auto": atac → tfidf, if / protein → zscore, hic → none, anything else → counts.

  • log (bool) – "zscore" only: log1p first (needs non-negative intensities).

  • methods (subset of ("pca", "tsne", "umap")) – With "tfidf" the linear step is LSI and is stored as f"{source}_lsi". UMAP needs umap-learn (pip install "u-chrom[umap]"); it is skipped with a warning when missing. t-SNE and UMAP run on the PCs / LSI components.

  • n_pcs (int) – Principal components kept (and fed to t-SNE / UMAP).

  • n_lsi (int) – LSI components kept ("tfidf").

  • drop_depth_r (float or None) – "tfidf": drop LSI component 1 if |r(component, log depth)| exceeds this.

  • library_size (bool) – "counts": divide by each cell’s total before log1p (see normalize_counts(); off by default for targeted panels).

  • modality (str, optional) – Display name (default from MODALITY_LABELS, e.g. atac → "ATAC"); the browser groups embeddings by it.

  • features (str, optional) – Free-text description of the features, recorded in the metadata.

Return type:

dict mapping cellm key → array (also written to cd.cellm).

uchrom.emb.cells.marker_features(cd, groups: str | Series, *, prefix: str = 'rna.', n_top: int = 5, log: bool = True) → Dict[str, List[str]][source]

Top features of each group of cells: the prefix columns of cd.cells ranked by the difference of the mean (log1p when log) between the group and the other cells.

uchrom.emb.cells.normalize_counts(counts: ndarray, *, library_size: bool = False) → ndarray[source]

log1p and z-score each feature (optionally library-size first).

library_size=False (default) suits small targeted panels (seqFISH, MERFISH): there one or two housekeeping genes dominate the total, so dividing by it turns PC1 into “that gene vs the rest” (on Takei 2021, Eef2 ≈ 16 % of all counts). Use True for transcriptome-wide data. Constant features become zero.

uchrom.emb.cells.tfidf_lsi(x, n_components: int = 30, *, depth: ndarray | None = None, drop_depth_r: float = 0.5, random_state: int = 0)[source]

TF-IDF + truncated SVD (LSI) of a cells × features count matrix.

Follows Signac’s RunTFIDF (method 1: log1p(tf · idf · 1e4), tf = count / cell total, idf = n_cells / feature total) and RunSVD (each component scaled to mean 0, sd 1). Dense or scipy.sparse input; sparse stays sparse.

Component 1 of scATAC LSI usually tracks sequencing depth. It is dropped when |Pearson r(component, log depth)| > drop_depth_r (only component 1 is tested; pass drop_depth_r=None to keep all).

Returns:

  • (lsi, info) — lsi ((n_cells, n_kept) array; info: dict with)

  • depth_r (per component), dropped (1-based component numbers),

  • explained_variance_ratio.

scHiCluster-style cell features from single-cell contact maps (.scool).

Zhou et al. 2019 (PNAS 116:14011, “Robust single-cell Hi-C clustering by convolution- and random-walk-based imputation”): for each cell and chromosome, the intra-chromosomal contact matrix is

  1. smoothed by a (2·pad+1)² mean filter (linear convolution),

  2. imputed by random walk with restart (Q ← (1-rp)·Q·P + rp·I until convergence, P the row-normalised smoothed matrix),

  3. binarised: the top top_pct % of the upper-triangle entries → 1.

Per chromosome the binarised vectors of all cells go through PCA; the concatenated per-chromosome PCs are the cell features (feed them to uchrom.emb.embed_cells() with source="hic", normalization="none", which runs the final PCA → t-SNE / UMAP).

uchrom.emb.hic.schicluster_features(per_chrom, *, pad: int = 1, rp: float = 0.5, top_pct: float = 20.0, n_per_chrom: int = 20, batch: int = 256, random_state: int = 0) → Tuple[ndarray, List[str]][source]

scHiCluster features, per chromosome.

per_chrom maps chrom → either an (n_cells, n, n) array of raw counts or a callable f(lo, hi) -> (hi - lo, n, n) array for a batch of cells (so all maps never need to be dense at once) together with n_cells — pass (callable, n_cells, n) tuples for that.

Returns (features, column_names) — per-chromosome PCs of the imputed, binarised maps, concatenated (chr1_PC1, …).

uchrom.emb.hic.schicluster_impute(mats: ndarray, *, pad: int = 1, rp: float = 0.5, tol: float = 1e-06, max_iter: int = 100) → ndarray[source]

Convolution + random-walk-with-restart imputation of a batch of symmetric contact matrices, shape (n_cells, n, n) (or (n, n)).

uchrom.emb.hic.scool_cell_features(scool: str, cell_names: Sequence[str], *, chroms: Sequence[str] | None = None, pad: int = 1, rp: float = 0.5, top_pct: float = 20.0, n_per_chrom: int = 20, random_state: int = 0) → Tuple[ndarray, List[str], ndarray][source]

Read one contact map per cell from a .scool and compute schicluster_features().

cell_names are the scool cell names, in the order wanted for the output rows; cells missing from the file get all-zero maps.

Returns (features, column_names, n_contacts_per_cell).

FastHigashi / Higashi wrapper for uchrom.emb.

Design notes

The expensive dependency (fasthigashi / higashi) is imported lazily and can also be injected via wrapper_factory for testing.

FastHigashi’s row order

FastHigashi.fetch_cell_embedding(restore_order=True) returns rows in the same order as label_info.pickle. We always pass restore_order=True and read that order from the pickle next to config.JSON — that’s the canonical row anchor and avoids relying on internal attributes that may change between FastHigashi versions.

Why we don’t store imputed maps in ChromData

Per-cell imputed contact maps are O(n²) per cell. ChromData is deliberately O(n) (no global spotp). When the user wants imputed maps, they live on disk under path2result_dir and we only record the path in cd.uns[embed_key]['result_dir'].

Apple Silicon shim

FastHigashi 0.1.1a0 calls tensor.pin_memory() unconditionally during prep_dataset, which crashes on PyTorch ≥2.x on Apple Silicon when MPS is the only non-CPU device available (PyTorch tries to pin against MPS, which is not a real pinned-memory backend). When CUDA isn’t available we no-op pin_memory for the duration of run() — pinned memory only matters for CUDA host→device transfers anyway.

uchrom.emb.higashi.align_embedding(emb: ndarray, emb_order: list[str], cd_cells: DataFrame, *, cell_id_col: str = 'cell_name') → ndarray[source]

Reorder emb rows to match cd_cells[cell_id_col] exactly.

Misaligning silently poisons every downstream analysis, so we fail loudly on missing / extra names rather than pad with NaNs.

uchrom.emb.higashi.run(cd, contacts_dir: str | Path, *, which: Literal['fast', 'full'] = 'fast', config_path: str | Path | None = None, rank: int = 256, embed_key: str = 'higashi', cell_id_col: str = 'cell_name', embed_variant: str = 'embed_l2_norm', result_dir: str | Path | None = None, cache_dir: str | Path | None = None, wrapper_factory: Callable[[...], Any] | None = None, wrapper_kwargs: dict[str, Any] | None = None, run_model_kwargs: dict[str, Any] | None = None)[source]

Run FastHigashi (or Higashi) and stash the cell embedding on cd.

Parameters:
  • cd (ChromData) – Must have a cells frame containing cell_id_col.

  • contacts_dir (path) – Directory produced by uchrom.io.write_higashi_inputs(). Must contain config.JSON and label_info.pickle.

  • which ({"fast", "full"}) – FastHigashi (default) or full Higashi.

  • rank (int) – Tucker rank passed to FastHigashi’s run_model and used as final_dim in fetch_cell_embedding.

  • embed_variant (str) – Which key from FastHigashi’s embedding dict to use. Default is embed_l2_norm (L2-normalised, what the official tutorial uses).

  • result_dir (path, optional) – Where FastHigashi caches per-cell tensors and writes results. Default to contacts_dir (lets write_higashi_inputs and FastHigashi share one directory tree).

  • cache_dir (path, optional) – Where FastHigashi caches per-cell tensors and writes results. Default to contacts_dir (lets write_higashi_inputs and FastHigashi share one directory tree).

  • wrapper_factory (callable, optional) –

    For tests / alt backends. Called as ``factory(config_path, path2input_cache, path2result_dir,

    off_diag, filter, do_conv, do_rwr, do_col, no_col)``

    — the real FastHigashi constructor takes positional args.

  • wrapper_kwargs (dict, optional) – Override the FastHigashi __init__ defaults (off_diag, filter, do_conv etc.).

  • run_model_kwargs (dict, optional) – Forwarded to run_model. rank from the top-level kwarg wins unless explicitly overridden here.

Returns:

The same cd object, mutated in place.

Return type:

ChromData