uchrom.emb¶
Cell embeddings: from per-cell features of any modality (embed_cells
with count / z-score / TF-IDF-LSI normalisation, aggregate_tracks for
per-spot imaging signals), from single-cell contact maps (scHiCluster,
uchrom.emb.hic) and from single-cell Hi-C with Higashi (higashi).
Cell embeddings from per-cell features of any modality.
embed_cells() turns a group of numeric cd.cells columns — the
rna.<gene> counts written by from_fofct(),
if.<mark> immunofluorescence summaries from aggregate_tracks(),
atac.<gene> accessibility, … — or an explicit cells × features matrix
into PCA (or LSI) / t-SNE / UMAP coordinates stored in cd.cellm (rows
aligned with cd.cells). Parameters, including the normalisation, are
recorded in cd.uns["embeddings"] so they round-trip through ChromData files
and the web browser can group the embeddings by modality.
Normalisations (normalization=; "auto" picks one from source):
"counts"(rnaand unknown prefixes)log1pthen z-score per feature (optionally library-size first)."zscore"(if,protein,ab)continuous intensities: z-score per feature,
log=Trueaddslog1p."tfidf"(atac)sparse accessibility: TF-IDF → truncated SVD (latent semantic indexing); component 1 is dropped when it tracks sequencing depth.
"none"features used as given (already normalised, e.g. Hi-C vectors).
- uchrom.emb.cells.MODALITY_LABELS = {'ab': 'Protein', 'atac': 'ATAC', 'hic': 'Hi-C', 'higashi': 'Hi-C', 'if': 'Histone/IF', 'protein': 'Protein', 'rna': 'RNA'}¶
display name of each feature prefix (the browser groups embeddings by it)
- uchrom.emb.cells.aggregate_tracks(cd, fields: Sequence[str] | None = None, *, prefix: str = 'if', stat: str = 'mean', by: None | str | int = None, min_spots: int = 1, write: bool = True) DataFrame[source]¶
Summarise tracks (
cd.spot_tracks+cd.bin_tracksseen per spot) into per-cell features.- Parameters:
cd (ChromData) – With tracks and a
cell_idcolumn incd.spots.fields (sequence of str, optional) – Track columns to use (default: every numeric track).
prefix (str) – Output columns are named
f"{prefix}.{field}".stat (
"mean","median","std"or"corr") – Per-cell statistic."corr"gives the within-cell Pearson correlation of every pair of fields (f"{prefix}.{a}~{b}") — how marks co-occur on the same loci, which is insensitive to per-cell intensity scaling.by (None,
"chrom"or int) – Also split by chromosome (…@chr1) or by genomic bins of this many bp (…@chr1:10= 10th bin); not withstat="corr".min_spots (int) – Groups with fewer spots become NaN.
write (bool) – Add the columns to
cd.cells(replacing ones of the same name).
- Return type:
DataFrame indexed like
cd.cells(cells without spots are NaN).
- uchrom.emb.cells.cluster_cells(cd, *, use: str = 'rna_pca', n_neighbors: int = 15, resolution: float = 1.0, n_components: int | None = None, random_state: int = 0, key_added: str | None = None) Series[source]¶
Leiden clusters of the cells from an embedding in
cd.cellm.A k-nearest-neighbour graph (
n_neighbors, Euclidean, symmetrised) ofcd.cellm[use](firstn_componentscolumns) is partitioned with the Leiden algorithm (leidenalg, RB configuration model atresolution, seeded). Clusters are numbered"0", "1", …by decreasing size and stored as a categorical columncd.cells[key_added](defaultf"leiden_{use}"); the parameters go tocd.uns["clusters"][key_added]. Needsleidenalgandpython-igraph(pip install leidenalg).Returns the cluster labels (a Series indexed like
cd.cells).
- uchrom.emb.cells.embed_cells(cd, *, source: str = 'rna', fields: Sequence[str] | None = None, matrix=None, normalization: str = 'auto', log: bool = False, methods: Sequence[str] = ('pca', 'tsne', 'umap'), n_pcs: int = 10, n_lsi: int = 30, drop_depth_r: float | None = 0.5, random_state: int = 0, tsne_perplexity: float | None = None, umap_neighbors: int = 15, umap_min_dist: float = 0.3, library_size: bool = False, modality: str | None = None, features: str | None = None) Dict[str, ndarray][source]¶
Compute cell embeddings and store them in
cd.cellm.- Parameters:
cd (ChromData) – Needs a non-empty
cd.cells; rows of the embeddings follow it.source (str) – Feature group; selects the numeric columns
f"{source}.*"and names the outputsf"{source}_{method}"(rna_pca,if_umap,atac_lsi, …).fields (sequence of str, optional) – Explicit
cd.cellscolumns to use instead of the prefix.matrix (array or scipy.sparse matrix, optional) – cells × features input instead of
cd.cellscolumns (rows incd.cellsorder) — e.g. a genome-wide ATAC bin matrix too large to keep as cell columns. Describe it withfeatures=.normalization (
"auto","counts","zscore","tfidf","none") – See the module docstring."auto":atac→ tfidf,if/protein→ zscore,hic→ none, anything else → counts.log (bool) –
"zscore"only:log1pfirst (needs non-negative intensities).methods (subset of
("pca", "tsne", "umap")) – With"tfidf"the linear step is LSI and is stored asf"{source}_lsi". UMAP needsumap-learn(pip install "u-chrom[umap]"); it is skipped with a warning when missing. t-SNE and UMAP run on the PCs / LSI components.n_pcs (int) – Principal components kept (and fed to t-SNE / UMAP).
n_lsi (int) – LSI components kept (
"tfidf").drop_depth_r (float or None) –
"tfidf": drop LSI component 1 if|r(component, log depth)|exceeds this.library_size (bool) –
"counts": divide by each cell’s total beforelog1p(seenormalize_counts(); off by default for targeted panels).modality (str, optional) – Display name (default from
MODALITY_LABELS, e.g.atac→"ATAC"); the browser groups embeddings by it.features (str, optional) – Free-text description of the features, recorded in the metadata.
- Return type:
dict mapping cellm key → array (also written to
cd.cellm).
- uchrom.emb.cells.marker_features(cd, groups: str | Series, *, prefix: str = 'rna.', n_top: int = 5, log: bool = True) Dict[str, List[str]][source]¶
Top features of each group of cells: the
prefixcolumns ofcd.cellsranked by the difference of the mean (log1pwhenlog) between the group and the other cells.
- uchrom.emb.cells.normalize_counts(counts: ndarray, *, library_size: bool = False) ndarray[source]¶
log1pand z-score each feature (optionally library-size first).library_size=False(default) suits small targeted panels (seqFISH, MERFISH): there one or two housekeeping genes dominate the total, so dividing by it turns PC1 into “that gene vs the rest” (on Takei 2021, Eef2 ≈ 16 % of all counts). UseTruefor transcriptome-wide data. Constant features become zero.
- uchrom.emb.cells.tfidf_lsi(x, n_components: int = 30, *, depth: ndarray | None = None, drop_depth_r: float = 0.5, random_state: int = 0)[source]¶
TF-IDF + truncated SVD (LSI) of a cells × features count matrix.
Follows Signac’s
RunTFIDF(method 1:log1p(tf · idf · 1e4), tf = count / cell total, idf = n_cells / feature total) andRunSVD(each component scaled to mean 0, sd 1). Dense orscipy.sparseinput; sparse stays sparse.Component 1 of scATAC LSI usually tracks sequencing depth. It is dropped when
|Pearson r(component, log depth)| > drop_depth_r(only component 1 is tested; passdrop_depth_r=Noneto keep all).- Returns:
(lsi, info) — lsi ((n_cells, n_kept) array; info: dict with)
depth_r(per component),dropped(1-based component numbers),explained_variance_ratio.
scHiCluster-style cell features from single-cell contact maps (.scool).
Zhou et al. 2019 (PNAS 116:14011, “Robust single-cell Hi-C clustering by convolution- and random-walk-based imputation”): for each cell and chromosome, the intra-chromosomal contact matrix is
smoothed by a
(2·pad+1)²mean filter (linear convolution),imputed by random walk with restart (
Q ← (1-rp)·Q·P + rp·Iuntil convergence, P the row-normalised smoothed matrix),binarised: the top
top_pct% of the upper-triangle entries → 1.
Per chromosome the binarised vectors of all cells go through PCA; the
concatenated per-chromosome PCs are the cell features (feed them to
uchrom.emb.embed_cells() with source="hic", normalization="none",
which runs the final PCA → t-SNE / UMAP).
- uchrom.emb.hic.schicluster_features(per_chrom, *, pad: int = 1, rp: float = 0.5, top_pct: float = 20.0, n_per_chrom: int = 20, batch: int = 256, random_state: int = 0) Tuple[ndarray, List[str]][source]¶
scHiCluster features, per chromosome.
per_chrommaps chrom → either an(n_cells, n, n)array of raw counts or a callablef(lo, hi) -> (hi - lo, n, n)array for a batch of cells (so all maps never need to be dense at once) together withn_cells— pass(callable, n_cells, n)tuples for that.Returns
(features, column_names)— per-chromosome PCs of the imputed, binarised maps, concatenated (chr1_PC1, …).
- uchrom.emb.hic.schicluster_impute(mats: ndarray, *, pad: int = 1, rp: float = 0.5, tol: float = 1e-06, max_iter: int = 100) ndarray[source]¶
Convolution + random-walk-with-restart imputation of a batch of symmetric contact matrices, shape
(n_cells, n, n)(or(n, n)).
- uchrom.emb.hic.scool_cell_features(scool: str, cell_names: Sequence[str], *, chroms: Sequence[str] | None = None, pad: int = 1, rp: float = 0.5, top_pct: float = 20.0, n_per_chrom: int = 20, random_state: int = 0) Tuple[ndarray, List[str], ndarray][source]¶
Read one contact map per cell from a
.scooland computeschicluster_features().cell_namesare the scool cell names, in the order wanted for the output rows; cells missing from the file get all-zero maps.Returns
(features, column_names, n_contacts_per_cell).
FastHigashi / Higashi wrapper for uchrom.emb.
Design notes¶
The expensive dependency (fasthigashi / higashi) is imported
lazily and can also be injected via wrapper_factory for testing.
FastHigashi’s row order¶
FastHigashi.fetch_cell_embedding(restore_order=True) returns rows in
the same order as label_info.pickle. We always pass
restore_order=True and read that order from the pickle next to
config.JSON — that’s the canonical row anchor and avoids relying on
internal attributes that may change between FastHigashi versions.
Why we don’t store imputed maps in ChromData¶
Per-cell imputed contact maps are O(n²) per cell. ChromData is
deliberately O(n) (no global spotp). When the user wants imputed
maps, they live on disk under path2result_dir and we only record the
path in cd.uns[embed_key]['result_dir'].
Apple Silicon shim¶
FastHigashi 0.1.1a0 calls tensor.pin_memory() unconditionally during
prep_dataset, which crashes on PyTorch ≥2.x on Apple Silicon when
MPS is the only non-CPU device available (PyTorch tries to pin against
MPS, which is not a real pinned-memory backend). When CUDA isn’t
available we no-op pin_memory for the duration of run() —
pinned memory only matters for CUDA host→device transfers anyway.
- uchrom.emb.higashi.align_embedding(emb: ndarray, emb_order: list[str], cd_cells: DataFrame, *, cell_id_col: str = 'cell_name') ndarray[source]¶
Reorder
embrows to matchcd_cells[cell_id_col]exactly.Misaligning silently poisons every downstream analysis, so we fail loudly on missing / extra names rather than pad with NaNs.
- uchrom.emb.higashi.run(cd, contacts_dir: str | Path, *, which: Literal['fast', 'full'] = 'fast', config_path: str | Path | None = None, rank: int = 256, embed_key: str = 'higashi', cell_id_col: str = 'cell_name', embed_variant: str = 'embed_l2_norm', result_dir: str | Path | None = None, cache_dir: str | Path | None = None, wrapper_factory: Callable[[...], Any] | None = None, wrapper_kwargs: dict[str, Any] | None = None, run_model_kwargs: dict[str, Any] | None = None)[source]¶
Run FastHigashi (or Higashi) and stash the cell embedding on
cd.- Parameters:
cd (ChromData) – Must have a
cellsframe containingcell_id_col.contacts_dir (path) – Directory produced by
uchrom.io.write_higashi_inputs(). Must containconfig.JSONandlabel_info.pickle.which ({"fast", "full"}) – FastHigashi (default) or full Higashi.
rank (int) – Tucker rank passed to FastHigashi’s
run_modeland used asfinal_diminfetch_cell_embedding.embed_variant (str) – Which key from FastHigashi’s embedding dict to use. Default is
embed_l2_norm(L2-normalised, what the official tutorial uses).result_dir (path, optional) – Where FastHigashi caches per-cell tensors and writes results. Default to
contacts_dir(letswrite_higashi_inputsand FastHigashi share one directory tree).cache_dir (path, optional) – Where FastHigashi caches per-cell tensors and writes results. Default to
contacts_dir(letswrite_higashi_inputsand FastHigashi share one directory tree).wrapper_factory (callable, optional) –
For tests / alt backends. Called as ``factory(config_path, path2input_cache, path2result_dir,
off_diag, filter, do_conv, do_rwr, do_col, no_col)``
— the real FastHigashi constructor takes positional args.
wrapper_kwargs (dict, optional) – Override the FastHigashi
__init__defaults (off_diag,filter,do_convetc.).run_model_kwargs (dict, optional) – Forwarded to
run_model.rankfrom the top-level kwarg wins unless explicitly overridden here.
- Returns:
The same
cdobject, mutated in place.- Return type: