uchrom.strc.tad

class uchrom.strc.tad.DICallerParams(window: int = 50, smoothing_param: float = 0.1, min_size_frac: float = 0.05)[source]

Bases: object

Knobs of call_tads_di() (defaults = get_domains()).

min_size_frac: float = 0.05
smoothing_param: float = 0.1
window: int = 50
class uchrom.strc.tad.FISHnetParams(thresholds: ndarray | None = None, threshold_step: float | None = None, threshold_min: float | None = None, threshold_max: float | None = None, plateau_size: int = 4, window_size: int = 2, size_exclusion: int = 3, merge_tol: int = 3, n_louvain_runs: int = 20, resolution: float = 1.0, min_coverage: float = 0.6, max_thresholds: int = 200, impute: bool = True)[source]

Bases: object

Runtime parameters for call_domains_fishnet_trace().

Defaults follow Patel et al. 2025.

impute: bool = True
max_thresholds: int = 200
merge_tol: int = 3
min_coverage: float = 0.6
n_louvain_runs: int = 20
plateau_size: int = 4
resolution: float = 1.0
size_exclusion: int = 3
threshold_max: float | None = None
threshold_min: float | None = None
threshold_step: float | None = None
thresholds: ndarray | None = None
window_size: int = 2
class uchrom.strc.tad.TADCallerParams(window_bp: float = 100000.0, fdr_cutoff: float = 0.1, hierarchical: bool = True, max_levels: int = 4, prominence: float = 0.0, distance: int = 1, k_sigma: float = 4.0, frac: float = 0.1)[source]

Bases: object

Runtime parameters for call_tads_by_pval().

Defaults follow ArcFISH’s TADCaller(method='pval').

distance: int = 1
fdr_cutoff: float = 0.1
frac: float = 0.1
hierarchical: bool = True
k_sigma: float = 4.0
max_levels: int = 4
prominence: float = 0.0
window_bp: float = 100000.0
uchrom.strc.tad.calc_directionality_index(contact_mat, window=50, device='auto')[source]

Calculate Directionality Index for each genomic bin.

uchrom.strc.tad.call_domains_fishnet(cd, *args, chrom=None, trace_ids=None, cells=None, params: FISHnetParams | None = None, key_added: str | None = UNSET, ensemble_key: str | None = UNSET, copy: bool = False, verbose: bool = False, max_traces: int | None = None, store=UNSET, result_key=UNSET)[source]

Run FISHnet on every trace of the selected chromosome(s).

Parameters:
  • cd (ChromData)

  • chrom (str, sequence of str, or None) – None (default) runs every chromosome and merges the tables.

  • trace_ids (sequence, optional) – Restrict to these traces / cells.

  • cells (sequence, optional) – Restrict to these traces / cells.

  • params (FISHnetParams, optional)

  • key_added (str or None) – cd.results key of the per-trace domain table (default "tads.fishnet"); None does not store.

  • ensemble_key (str, optional) – cd.results key of the ensemble domain mask (default f"{key_added}.ensemble"). Stored as a mapping {chrom: {"mask": (n_bins, n_bins) int64, "bin_ids": (n_bins, 2) int64 [start, end], "n_traces": int}}.

  • copy (bool) – True → return a new ChromData holding the results.

  • max_traces (int, optional) – Cap on the number of traces processed per chromosome (useful for previews — FISHnet is O(n_bins² · n_thresholds) per trace).

Returns:

One row per (trace, domain): chrom, trace_id, domain_idx, start, end, start_bin, end_bin.

Return type:

DataFrame (copy=False) or ChromData (copy=True)

Notes

Deprecated 1.x forms (positional chrom, store=, result_key=) still work with a DeprecationWarning; they keep the 1.x keys "fishnet_domains" / "fishnet_ensemble" and the 1.x single-chromosome ensemble layout {"chrom", "mask", "bin_ids", "n_traces"}.

uchrom.strc.tad.call_domains_fishnet_trace(dist: ndarray, params: FISHnetParams | None = None, seed_base: int = 0) → dict[source]

Run FISHnet on a single-allele pairwise distance matrix.

Parameters:
  • dist (ndarray (n_bins, n_bins)) – Symmetric pairwise distance matrix, NaN for missing spots, in nm.

  • params (FISHnetParams, optional) – Hyperparameters. Use defaults if None.

  • seed_base (int) – Base seed for the 20 Louvain runs (reproducibility).

Returns:

domains : list of (start, end_exclusive) bin indices per plateau boundaries : sorted list of canonical boundary bin indices domain_mask : (n_bins, n_bins) int — count of plateaus in which

bins i and j fall in the same domain (0 if never).

n_plateaus : number of plateaus detected coverage : fraction of bins with ≥1 non-NaN distance

Return type:

dict with keys

uchrom.strc.tad.call_tads_by_pval(cd, *args, chrom=None, trace_ids=None, cells=None, params: TADCallerParams | None = None, device: str = 'auto', key_added: str | None = UNSET, copy: bool = False, verbose: bool = False, streaming: bool | None = None, batch='auto', memory_budget=None, store=UNSET, result_key=UNSET)[source]

ArcFISH-style TAD caller (chromatin tracing).

Parameters:
  • cd (ChromData)

  • chrom (str, sequence of str, or None) – Chromosome(s) to call. None (default) calls every chromosome with spots and merges the per-chromosome tables.

  • trace_ids (sequence, optional) – Restrict the population to these traces / cells.

  • cells (sequence, optional) – Restrict the population to these traces / cells.

  • params (TADCallerParams, optional)

  • device ('auto' | 'cpu' | 'cuda' | 'mps')

  • key_added (str or None) – cd.results key for the merged table (default "tads.arcfish"); None does not store.

  • copy (bool) – True → return a new ChromData with the result; cd is not modified.

  • streaming (bool, optional) – Compute the axis-variance cube by streaming over the traces (uchrom.fea.arc_stream.axis_cube_streaming(): exact medians per row band, memory O(n_bins²) instead of O(n_traces × n_bins²)). Default: True for a backed ChromData. Results equal the in-memory path up to floating-point rounding of the variance sums.

  • batch – Traces per batch ("auto") and working-memory budget of the streaming path (default uchrom.settings.memory_budget).

  • memory_budget – Traces per batch ("auto") and working-memory budget of the streaming path (default uchrom.settings.memory_budget).

Returns:

Columns chrom, start, end, level, score, pval, fdr. Level 1 is the finest partition (all boundaries); each further level drops the weakest boundary when params.hierarchical=True.

Return type:

DataFrame (copy=False) or ChromData (copy=True)

Notes

Deprecated 1.x forms still work with a DeprecationWarning: positional chrom, store= and result_key=. When any of them is used and key_added is not given, the result is stored under the 1.x key "tads".

uchrom.strc.tad.call_tads_di(cd, *, contacts='default', resolution: int | None = None, balance: bool = True, chrom=None, params: DICallerParams | None = None, device: str = 'auto', key_added: str | None = 'tads.di', copy: bool = False)[source]

Directionality-index TADs on a linked Hi-C contact matrix.

Parameters:
  • cd (ChromData) – Holds the link (cd.link_cool(path, key=...)) and receives the result.

  • contacts (str) – Key of cd.uns['linked_cool'] (default "default") or a path to a .cool / .mcool.

  • resolution (int, optional) – Required for .mcool; checked against a .cool’s bin size.

  • balance (bool) – Use the cooler weight column (falls back to raw counts, with a warning, when the file is not balanced).

  • chrom (str, sequence of str, or None) – None → the chromosomes of cd.spots that the cooler has, or every cooler chromosome when cd has no spots on any of them (e.g. Hi-C-only data).

  • params (DICallerParams, optional)

  • key_added (str or None) – cd.results key (default "tads.di"); None does not store.

  • copy (bool) – True → return a new ChromData holding the result.

Returns:

chrom, start, end, start_bin, end_bin — bin indices are per-chromosome, end_bin exclusive, exactly as returned by get_domains(). (The kernel’s last domain stops one bin short of the chromosome end; this is kept as is.)

Return type:

DataFrame (copy=False) or ChromData (copy=True)

uchrom.strc.tad.detect_tad_boundaries(di, min_size_frac=0.05)[source]

Detect TAD boundaries from sign changes (negative -> positive).

uchrom.strc.tad.get_domains(contact_mat, smoothing_param=0.1, min_size_frac=0.05, window=50, device='auto')[source]

Identify TADs: DI calculation -> smoothing -> boundary detection.

uchrom.strc.tad.load_tad_from_bed(bed_path)[source]

Load TAD regions from BED file.

uchrom.strc.tad.substructures_from_tads(tad_indices, structure_coords)[source]

Extract substructures based on TAD boundaries.

FISHnet (per-allele, graph-theory)

class uchrom.strc.tad.fishnet.FISHnetParams(thresholds: ndarray | None = None, threshold_step: float | None = None, threshold_min: float | None = None, threshold_max: float | None = None, plateau_size: int = 4, window_size: int = 2, size_exclusion: int = 3, merge_tol: int = 3, n_louvain_runs: int = 20, resolution: float = 1.0, min_coverage: float = 0.6, max_thresholds: int = 200, impute: bool = True)[source]

Bases: object

Runtime parameters for call_domains_fishnet_trace().

Defaults follow Patel et al. 2025.

impute: bool = True
max_thresholds: int = 200
merge_tol: int = 3
min_coverage: float = 0.6
n_louvain_runs: int = 20
plateau_size: int = 4
resolution: float = 1.0
size_exclusion: int = 3
threshold_max: float | None = None
threshold_min: float | None = None
threshold_step: float | None = None
thresholds: ndarray | None = None
window_size: int = 2
uchrom.strc.tad.fishnet.call_domains_fishnet(cd, *args, chrom=None, trace_ids=None, cells=None, params: FISHnetParams | None = None, key_added: str | None = UNSET, ensemble_key: str | None = UNSET, copy: bool = False, verbose: bool = False, max_traces: int | None = None, store=UNSET, result_key=UNSET)[source]

Run FISHnet on every trace of the selected chromosome(s).

Parameters:
  • cd (ChromData)

  • chrom (str, sequence of str, or None) – None (default) runs every chromosome and merges the tables.

  • trace_ids (sequence, optional) – Restrict to these traces / cells.

  • cells (sequence, optional) – Restrict to these traces / cells.

  • params (FISHnetParams, optional)

  • key_added (str or None) – cd.results key of the per-trace domain table (default "tads.fishnet"); None does not store.

  • ensemble_key (str, optional) – cd.results key of the ensemble domain mask (default f"{key_added}.ensemble"). Stored as a mapping {chrom: {"mask": (n_bins, n_bins) int64, "bin_ids": (n_bins, 2) int64 [start, end], "n_traces": int}}.

  • copy (bool) – True → return a new ChromData holding the results.

  • max_traces (int, optional) – Cap on the number of traces processed per chromosome (useful for previews — FISHnet is O(n_bins² · n_thresholds) per trace).

Returns:

One row per (trace, domain): chrom, trace_id, domain_idx, start, end, start_bin, end_bin.

Return type:

DataFrame (copy=False) or ChromData (copy=True)

Notes

Deprecated 1.x forms (positional chrom, store=, result_key=) still work with a DeprecationWarning; they keep the 1.x keys "fishnet_domains" / "fishnet_ensemble" and the 1.x single-chromosome ensemble layout {"chrom", "mask", "bin_ids", "n_traces"}.

uchrom.strc.tad.fishnet.call_domains_fishnet_trace(dist: ndarray, params: FISHnetParams | None = None, seed_base: int = 0) → dict[source]

Run FISHnet on a single-allele pairwise distance matrix.

Parameters:
  • dist (ndarray (n_bins, n_bins)) – Symmetric pairwise distance matrix, NaN for missing spots, in nm.

  • params (FISHnetParams, optional) – Hyperparameters. Use defaults if None.

  • seed_base (int) – Base seed for the 20 Louvain runs (reproducibility).

Returns:

domains : list of (start, end_exclusive) bin indices per plateau boundaries : sorted list of canonical boundary bin indices domain_mask : (n_bins, n_bins) int — count of plateaus in which

bins i and j fall in the same domain (0 if never).

n_plateaus : number of plateaus detected coverage : fraction of bins with ≥1 non-NaN distance

Return type:

dict with keys

F-test caller

class uchrom.strc.tad.arc_pval.TADCallerParams(window_bp: float = 100000.0, fdr_cutoff: float = 0.1, hierarchical: bool = True, max_levels: int = 4, prominence: float = 0.0, distance: int = 1, k_sigma: float = 4.0, frac: float = 0.1)[source]

Bases: object

Runtime parameters for call_tads_by_pval().

Defaults follow ArcFISH’s TADCaller(method='pval').

distance: int = 1
fdr_cutoff: float = 0.1
frac: float = 0.1
hierarchical: bool = True
k_sigma: float = 4.0
max_levels: int = 4
prominence: float = 0.0
window_bp: float = 100000.0
uchrom.strc.tad.arc_pval.call_tads_by_pval(cd, *args, chrom=None, trace_ids=None, cells=None, params: TADCallerParams | None = None, device: str = 'auto', key_added: str | None = UNSET, copy: bool = False, verbose: bool = False, streaming: bool | None = None, batch='auto', memory_budget=None, store=UNSET, result_key=UNSET)[source]

ArcFISH-style TAD caller (chromatin tracing).

Parameters:
  • cd (ChromData)

  • chrom (str, sequence of str, or None) – Chromosome(s) to call. None (default) calls every chromosome with spots and merges the per-chromosome tables.

  • trace_ids (sequence, optional) – Restrict the population to these traces / cells.

  • cells (sequence, optional) – Restrict the population to these traces / cells.

  • params (TADCallerParams, optional)

  • device ('auto' | 'cpu' | 'cuda' | 'mps')

  • key_added (str or None) – cd.results key for the merged table (default "tads.arcfish"); None does not store.

  • copy (bool) – True → return a new ChromData with the result; cd is not modified.

  • streaming (bool, optional) – Compute the axis-variance cube by streaming over the traces (uchrom.fea.arc_stream.axis_cube_streaming(): exact medians per row band, memory O(n_bins²) instead of O(n_traces × n_bins²)). Default: True for a backed ChromData. Results equal the in-memory path up to floating-point rounding of the variance sums.

  • batch – Traces per batch ("auto") and working-memory budget of the streaming path (default uchrom.settings.memory_budget).

  • memory_budget – Traces per batch ("auto") and working-memory budget of the streaming path (default uchrom.settings.memory_budget).

Returns:

Columns chrom, start, end, level, score, pval, fdr. Level 1 is the finest partition (all boundaries); each further level drops the weakest boundary when params.hierarchical=True.

Return type:

DataFrame (copy=False) or ChromData (copy=True)

Notes

Deprecated 1.x forms still work with a DeprecationWarning: positional chrom, store= and result_key=. When any of them is used and key_added is not given, the result is stored under the 1.x key "tads".

Directionality index

class uchrom.strc.tad.di_caller.DICallerParams(window: int = 50, smoothing_param: float = 0.1, min_size_frac: float = 0.05)[source]

Bases: object

Knobs of call_tads_di() (defaults = get_domains()).

min_size_frac: float = 0.05
smoothing_param: float = 0.1
window: int = 50
uchrom.strc.tad.di_caller.call_tads_di(cd, *, contacts='default', resolution: int | None = None, balance: bool = True, chrom=None, params: DICallerParams | None = None, device: str = 'auto', key_added: str | None = 'tads.di', copy: bool = False)[source]

Directionality-index TADs on a linked Hi-C contact matrix.

Parameters:
  • cd (ChromData) – Holds the link (cd.link_cool(path, key=...)) and receives the result.

  • contacts (str) – Key of cd.uns['linked_cool'] (default "default") or a path to a .cool / .mcool.

  • resolution (int, optional) – Required for .mcool; checked against a .cool’s bin size.

  • balance (bool) – Use the cooler weight column (falls back to raw counts, with a warning, when the file is not balanced).

  • chrom (str, sequence of str, or None) – None → the chromosomes of cd.spots that the cooler has, or every cooler chromosome when cd has no spots on any of them (e.g. Hi-C-only data).

  • params (DICallerParams, optional)

  • key_added (str or None) – cd.results key (default "tads.di"); None does not store.

  • copy (bool) – True → return a new ChromData holding the result.

Returns:

chrom, start, end, start_bin, end_bin — bin indices are per-chromosome, end_bin exclusive, exactly as returned by get_domains(). (The kernel’s last domain stops one bin short of the chromosome end; this is kept as is.)

Return type:

DataFrame (copy=False) or ChromData (copy=True)

Low-level kernels:

uchrom.strc.tad.di.calc_di_batch(contact_mat, window=50, device='auto')[source]

Batch-optimized DI calculation using matrix operations.

uchrom.strc.tad.di.calc_directionality_index(contact_mat, window=50, device='auto')[source]

Calculate Directionality Index for each genomic bin.

uchrom.strc.tad.di.detect_tad_boundaries(di, min_size_frac=0.05)[source]

Detect TAD boundaries from sign changes (negative -> positive).

uchrom.strc.tad.di.get_device(device='auto')[source]

Get appropriate torch device.

uchrom.strc.tad.di.get_domains(contact_mat, smoothing_param=0.1, min_size_frac=0.05, window=50, device='auto')[source]

Identify TADs: DI calculation -> smoothing -> boundary detection.

uchrom.strc.tad.di.get_dtype(device)[source]

MPS only supports float32.

uchrom.strc.tad.di.smooth_with_moving_average(signal, window_size)[source]

Apply moving average smoothing via 1D convolution.

uchrom.strc.tad.utils.load_tad_from_bed(bed_path)[source]

Load TAD regions from BED file.

uchrom.strc.tad.utils.merge_substructures(substructures, tad_indices, total_bins)[source]

Merge TAD substructures back into full structure.

uchrom.strc.tad.utils.save_tads_to_bed(tad_indices, bins_df, output_path)[source]

Save detected TAD regions to BED file.

uchrom.strc.tad.utils.substructures_from_tads(tad_indices, structure_coords)[source]

Extract substructures based on TAD boundaries.

uchrom.strc.tad.utils.tad_regions_to_bin_indices(regions, bins_df, chrom=None)[source]

Convert TAD genomic regions to bin indices.