Structures from a bulk map: MDS and IGM¶
A bulk map is an average over millions of nuclei. MDS turns it into one consensus structure; IGM into a
population of structures whose contacts together reproduce it. Both return a ChromData (IGM: one cell
per structure, one trace per chromosome copy), measured and compared with imaging like any traced data.
Tutorial¶
MDS: one consensus structure¶
import uchrom as uc
from uchrom.fea import rebin_map
rebin_map(hic, "bulk", resolution=100_000) # a 100 kb copy of the map, linked as "bulk_100kb"
model = uc.tl.reconstruct_bulk(hic, contacts="bulk_100kb") # = uchrom.recon.bulk.mds.reconstruct_mds
The command line does the same from a file:
Input: .hic / .mcool / .cool.
Output: .chromdata.zarr (or per-chromosome for inter mode).
# Single chromosome
python -m uchrom.recon.bulk.mds data.mcool out.chromdata.zarr \
--resolution=100000 --chrom=chr1 --device=auto
# Whole genome with inter-chromosomal contacts
python -m uchrom.recon.bulk.mds data.mcool ./outdir/ \
--resolution=100000 --resolution-inter=1000000 \
--inter=True --device=auto
The solver is SMACOF with CMDS initialisation, implemented in PyTorch so it runs on CUDA / MPS / CPU. Inter-chromosomal mode runs a low-resolution whole-genome scaffold MDS then Procrustes-aligns each chromosome’s high-resolution structure to its scaffold position.
Additional flags:
Flag |
Default |
Meaning |
|---|---|---|
|
4.0 |
Contact → distance exponent |
|
0.05 |
Distance-decay prior weight |
|
1000 |
SMACOF iterations |
|
False |
Split by TAD for better parallelism |
|
None |
TAD regions BED when |
IGM: a population of structures¶
IGM (Integrative Genome Modeling; Boninsegna et al. 2022, Nat Methods 19:938; alberlab/igm) fits a population of structures
so that, together, they reproduce the contact probabilities of the map: each pair of loci in contact
with probability p is placed in contact in a fraction p of the structures (assignment), the
structures are relaxed under those restraints (modelling), and the two steps alternate over decreasing
probability thresholds. U-Chrom runs the modelling on the native engine (protocol "igm"); the
defaults follow the IGM 1.0 protocol, with options from IGM-Plus / IGM 2.0 as fields of IgmParams.
import uchrom.datasets as ds
from uchrom.recon.bulk.deconv import deconvolve
from uchrom.recon.bulk.deconv.preprocess import preprocess_hic
hic = ds.load("rao2014_imr90_chr21") # IMR90 chr21, 5 kb; the map linked as "bulk"
prob = preprocess_hic(hic, contacts="bulk") # Hi-C -> contact probabilities
pop = deconvolve(prob, method="igm", n_structures=50, seed=1)
The result is one cell per structure and one trace per chromosome copy; out="pop.chromdata.zarr"
streams a large population to disk. The tutorial checks it against
the input map and against chromatin tracing of the same cell line.