3D reconstruction¶
U-Chrom ships three reconstruction backends under uchrom.recon, all
writing the same .h5cd output format.
Backend |
Module |
Best for |
Notes |
|---|---|---|---|
EMber |
|
Single-cell Hi-C / Dip-C, noisy contacts |
Native engine (protocol |
NucDynamics |
|
Single-cell Hi-C / Dip-C |
Native engine (Rust; multi-core CPU, wgpu GPU: Metal / Vulkan); ensembles |
GEM |
|
Single-cell, energy minimisation |
Taichi autodiff |
SMACOF MDS |
|
Bulk Hi-C, contact matrices |
PyTorch GPU, inter-chromosomal support |
NucDynamics (single-cell)¶
A re-implementation of NucDynamics (Stevens et al. 2017, Nature 544:59,
doi:10.1038/nature21429; original code tjs23/nuc_dynamics, LGPL-3.0) in the
separate native package u-chrom-nucdyn (native/nucdyn, Rust + PyO3):
pip install u-chrom[nucdyn] # or: pip install ./native/nucdyn
It computes an ensemble of structures at once — models in parallel on the
CPU (f64), all models batched on the GPU (f32, wgpu: Metal on macOS, Vulkan
on Linux/NVIDIA). Without the native package the deprecated Taichi port
runs instead (with a DeprecationWarning); it is not faithful to the
original and will be removed.
Input: .pairs / .pairs.gz, NCC (.ncc), a chr_A pos_A chr_B pos_B
table (the Stevens 2017 GEO files) or a DataFrame. Output: a ChromData
with one spot per particle and one trace per chromosome; coords = model 0
and layers["model_<k>"] = every model (units: particle radii).
from uchrom.recon.sc.nucdyn import reconstruct_nucdyn
cd = reconstruct_nucdyn("cell1.pairs", n_models=10, device="auto", seed=1)
cd.write("cell1.chromdata.zarr")
CLI:
python -m uchrom.recon.sc.nucdyn cell1.pairs out.chromdata.zarr \
--n_models=10 --device=gpu --seed=1
Key parameters:
Param |
Default |
Meaning |
|---|---|---|
|
10 |
Structures in the ensemble |
|
|
|
|
0 (all) |
CPU threads |
|
0 |
Random seed (CPU results are bitwise reproducible per seed) |
|
|
The 2017 code behind the published structures; |
|
8, 4, 2, 0.4, 0.2, 0.1 Mb |
Hierarchical schedule |
|
500 / 100 |
Annealing temperature steps per stage / MD steps per temperature |
|
5000 / 10 |
Annealing temperatures |
|
None |
Restrict the contacts, e.g. |
EMber (single-cell, noise-aware)¶
EMber (EM + annealing / sampling) runs on the same native engine as
NucDynamics (engine protocol "robust", alias "ember"). It keeps the
NucDynamics 2017 base and adds a mixture observation model — each contact is a
true proximity or a false contact; contact restraints are re-weighted by
their posterior probability of being true (EM; false-contact rate and contact
kernel learned from the data, plus local-support evidence) — a short schedule,
a finite-temperature sampling stage (calibrated ensembles) and adaptive
resolution below 25 kb. It was pre-registered, frozen and tested against
NucDynamics on held-out contacts and imaging-truth simulations
(benchmarks/screcon/PREREG.md, TEST_RESULTS.md): better on both primary
endpoints, ~5x faster on one GPU; its real-data ensembles are less tight.
from uchrom.recon.sc import reconstruct_ember
cd = reconstruct_ember("cell1.pairs", n_models=10, device="auto", seed=1)
cd.results["ember.contact_weights"] # posterior weight of every input contact
# also: uchrom.tl.reconstruct_sc(contacts, method="ember")
# reconstruct_nucdyn(contacts, protocol="ember") (same coordinates for the same seed)
GEM (single-cell, energy minimisation)¶
Same CLI pattern:
python -m uchrom.recon.sc.gem cell1.pairs out.h5cd --arch=cuda
Writes four energy terms each step (stretch / bend / exclude / contact)
with Taichi’s ti.Tape for autodiff gradient descent.
Bulk MDS (PyTorch, whole-genome)¶
Input: .hic / .mcool / .cool.
Output: .h5cd (or per-chromosome for inter mode).
# Single chromosome
python -m uchrom.recon.bulk.mds data.mcool out.h5cd \
--resolution=100000 --chrom=chr1 --device=auto
# Whole genome with inter-chromosomal contacts
python -m uchrom.recon.bulk.mds data.mcool ./outdir/ \
--resolution=100000 --resolution-inter=1000000 \
--inter=True --device=auto
The solver is SMACOF with CMDS initialisation, implemented in PyTorch so it runs on CUDA / MPS / CPU. Inter-chromosomal mode runs a low-resolution whole-genome scaffold MDS then Procrustes-aligns each chromosome’s high-resolution structure to its scaffold position.
Additional flags:
Flag |
Default |
Meaning |
|---|---|---|
|
4.0 |
Contact → distance exponent |
|
0.05 |
Distance-decay prior weight |
|
1000 |
SMACOF iterations |
|
False |
Split by TAD for better parallelism |
|
None |
TAD regions BED when |
Programmatic API¶
from uchrom.recon.sc.nucdyn import reconstruct_nucdyn
cd = reconstruct_nucdyn("cell1.pairs", n_models=10) # ChromData ensemble
cd.write("out.chromdata.zarr")
from uchrom import ChromData
cd = ChromData.read("out.h5cd")
All three backends write .h5cd with one trace per chromosome (single
cell) or one chromosome structure (bulk). Downstream tools don’t care
which reconstructor you used.