An umbrella framework for multi-source & multi-omics clustering in R. One interface, many methods, a complete workflow — from raw modalities to interpretable subgroups.
MosaiClusteR — “MoSaIC” = Multi-Omics Source-Agnostic Integration Clustering in R — treats every data modality (transcriptome, methylome, proteome, fingerprints, targets, imaging, clinical…) as one tile of a larger picture, and assembles those tiles into a single coherent mosaic of sample structure. It unifies a large family of integrative clustering algorithms behind one consistent list‑of‑matrices interface, and wraps them in a complete analysis framework: preprocessing → single‑source baselines → integration across five paradigms → evaluation & interpretation.
No single algorithm is best across all data regimes, so MosaiClusteR gathers many methods on the same footing and gives you the tooling to choose between them: cluster‑number selection, agreement metrics, visual comparison, cluster characterisation, and — for omics — pathway and gene‑set interpretation. It also introduces a data‑nugget feature‑weighting scheme that offers a robust, Big‑Data‑friendly alternative to variance weighting.
Multi‑omics and other multi‑source studies measure the same objects (patients, tumours, compounds, cells) through several heterogeneous lenses. No single modality tells the whole story, and the many integration algorithms in the literature live in scattered packages with incompatible inputs and outputs. MosaiClusteR fixes this:
List of matrices over the same shared objects,
with objects (samples) in rows and features in columns
(n × m) — the same convention across the entire package,
including the M‑ABC family. Distance‑matrix inputs
(type = "dist") are orientation‑free.DistM (a dissimilarity matrix) and Clust (a
hierarchical‑clustering object), so results are directly comparable and
composable.# install.packages("remotes")
remotes::install_github("lucp12891/MosaiClusteR")MosaiClusteR ships C++ kernels for the M‑ABC family, so a C/C++ toolchain is needed to build from source (on Windows install Rtools; on macOS the Xcode command‑line tools). Every method has a pure‑R fallback, so the package still runs if a kernel is unavailable. A few methods rely on suggested packages that are only loaded when you call them:
# optional, enables extra methods / distances
install.packages(c("SNFtool", "datanugget", "WCluster", "lsa", "analogue", "FactoMineR"))| If you want… | Install (Suggests) |
|---|---|
Similarity Network Fusion (SNF) |
SNFtool |
| The reference data‑nugget engine | datanugget, WCluster |
| Cosine / χ² / MCA distances | lsa, analogue,
FactoMineR |
library(MosaiClusteR)
data(mosaic_toy) # two sources, 60 shared objects, 3 planted groups
L <- mosaic_toy$List # a list of matrices, objects in rows
truth <- mosaic_toy$truth
## ---- Integrate: every method shares one interface — swap to change paradigm
fit <- SNF(L, type = "data", distmeasure = c("euclidean", "euclidean"),
normalize = c(FALSE, FALSE), method = list(NULL, NULL))
labels <- mosaic_labels(fit, k = 3) # pull a flat partition from any fitted result
cluster_agreement(labels, truth) # ARI / NMI / Jaccard / purity vs ground truth
## ---- Same score, a different paradigm — e.g. a neighbourhood method
cluster_agreement(NEMO(L, k = 3)$cluster, truth)
## ---- Compare two integrations directly
compare_clusterings(fit$DistM,
ADC(L, distmeasure = "euclidean", normalize = FALSE)$DistM, NC = 3)That is the whole loop: integrate → extract → evaluate → compare — and the same three lines work for any of the methods below.
MosaiClusteR is organised as a four‑layer pipeline. You can enter at
any layer, and every layer speaks the same
List → {DistM, Clust} language.
┌─────────────────────────────────────────────────────────────┐
LAYER 1 │ PREPROCESSING │
prepare │ Normalization() · Distance() │
│ heterogeneous, differently-scaled modalities → comparable │
└─────────────────────────────────────────────────────────────┘
│ List of matrices (features × objects)
▼
┌─────────────────────────────────────────────────────────────┐
LAYER 2 │ SINGLE-SOURCE BASELINES │
explore │ cluster each tile alone → see what each modality recovers │
│ motivates integration: agreements vs. modality-specific │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
LAYER 3 │ INTEGRATION ENGINE (five paradigms) │
integrate │ Direct · Similarity · Graph · Voting-consensus · Hierarchy │
│ M_ABCpp · M_ABCdist · WeightedClust · SNF · CEC · HEC · … │
└─────────────────────────────────────────────────────────────┘
│ {DistM, Clust}
▼
┌─────────────────────────────────────────────────────────────┐
LAYER 4 │ EVALUATION & INTERPRETATION │
conclude │ mosaic_labels() · cluster_agreement() · compare_clusterings│
│ ARI / NMI / Jaccard · cross-method consensus │
└─────────────────────────────────────────────────────────────┘
Choosing a layer‑3 method (decision shortcut):
ADC) or Weighted
(WeightedClust).SNF).M_ABCpp,
CEC, EvidenceAccumulation).HEC).weighting = "nugget" inside the M‑ABC
family.Legend — a named function is implemented in this package · † requires an external graph‑partitioning executable (METIS) or MATLAB, not bundled · “—” marks methods interoperable with, or on the roadmap via, a reference package. All methods share the
List → {DistM, Clust}contract and take the same objects‑in‑rows input.
| Method | Function | What it does | When it’s useful | Ref |
|---|---|---|---|---|
| ADC — Aggregated Data Clustering | ADC |
Concatenate all tiles [D₁│…│Dₗ], one distance, one
hierarchy |
Simplest baseline; when modalities are commensurate and you want a transparent reference | Fodeh 2013 |
| ADEC — Aggregated Data Ensemble Clustering | ADEC |
ADC + resampling of features / cut‑points → co‑clustering matrix | When ADC is unstable and you want a robustness layer | Fodeh 2013 |
| Method | Function | What it does | When it’s useful | Ref |
|---|---|---|---|---|
| Weighted Clustering | WeightedClust |
Convex combination Σ wₖ Dₖ over a weight grid |
When you want to tune and inspect the trade‑off between sources | Perualila‑Tan 2016 |
| SNF — Similarity Network Fusion | SNF |
Cross‑diffuse per‑source similarity networks to one fused network | Noisy, complementary, genomic‑scale sources; the multi‑omics workhorse | Wang 2014 |
| WonM — Weighting on Membership | WonM |
Consensus co‑membership across many cut‑points k | When you don’t want to commit to a single k | — |
| NEMO — Neighborhood‑based Multi‑Omics | NEMO |
Locally‑scaled kNN affinities averaged over the omics measuring each pair → spectral clustering | When samples are missing some modalities (no imputation needed) | Rappoport 2019 |
| Spectral clustering | spectral_clustering |
Normalised‑Laplacian embedding + k‑means on any affinity | Final clusterer for fused/similarity matrices (SNF, NEMO, Spectrum) | Ng 2002 |
| ab‑SNF | — | SNF with association/variance feature weights | When a few features carry the signal and should drive the network | Zhang 2022 |
| CIMLR / rMKL‑LPP | — | Multi‑kernel similarity learning | Many heterogeneous kernels; learn their weights | Zhang 2022 |
| Method | Function | What it does | When it’s useful | Ref |
|---|---|---|---|---|
| CSPA / HGPA / MCLA | EnsembleClustering † |
Cluster‑ensemble via (hyper)graph partitioning | Combining many base partitions into one robust consensus | Strehl 2002 |
| HBGF — Hybrid Bipartite Graph Formulation | HBGF |
Partition an object↔︎cluster bipartite graph (SVD + k‑means) | Ensembles where objects and clusters co‑embed naturally | Fern 2004 |
| Clustering Aggregation (Balls/Agglo/Furthest) | ClusteringAggregation |
Minimise pairwise disagreement cost | When you have several partitions and want the “median” one | Gionis 2007 |
| Method | Function | What it does | When it’s useful | Ref |
|---|---|---|---|---|
| M‑ABC — Multi‑source Aggregating Bundles of Clusters | M_ABCpp |
Per‑source bootstrap of feature bundles → consensus co‑clustering dissimilarity | High‑dimensional sources, where signal hides in feature subsets; robust by design; supports data‑nugget weighting | Amaratunga 2008 |
| M‑ABCdist | M_ABCdist |
M‑ABC that accumulates distances (keeps geometry) then fuses | When you want continuous dissimilarities rather than co‑clustering counts | Amaratunga 2008 |
| M‑ABCdeep — deep‑learning ABC | M_ABCdeep / ABCdeep.SingleInMultiple |
M‑ABC whose base clustering is a neural autoencoder embedding + latent clustering; self‑contained R network, no Python/torch | Nonlinear/complex per‑source structure; when a learned embedding beats a raw distance | Amaratunga 2008 |
| CEC — (Consensus) Ensemble Clustering | CEC |
Incidence accumulation across sources & cut‑points | A solid, weight‑aware consensus baseline | Fodeh 2013 |
| CVAA / W‑CVAA | CVAA |
Cumulative voting aggregation (and weighted) | Aligning and merging many partitions by voting | Saeed 2012/2014 |
| IVC / IPVC / IPC | ConsensusClustering |
Iterative (probabilistic) voting consensus | Iteratively refine a consensus labelling | Nguyen 2007 |
| EA — Evidence Accumulation | EvidenceAccumulation |
Co‑association matrix across partitions (SL / SL‑agnes / MST) | Classic, parameter‑light consensus | Fred 2005 |
| CTS / SRS / ASRS | LinkBasedClustering |
Link‑based cluster ensembles (CTS/SRS pure‑R; ASRS needs MATLAB) | When pairwise link similarity refines the consensus | Iam‑On 2010 |
| Method | Function | What it does | When it’s useful | Ref |
|---|---|---|---|---|
| HEC — Hierarchical Ensemble Clustering | HierarchicalEnsembleClustering |
Aggregate cophenetic distances → closest ultrametric (Floyd–Warshall) | When the dendrogram structure across sources is the object of interest | Zheng 2014 |
| EHC — Ensemble for Hierarchical Clustering | EHC † |
Graph aggregation of dendrogram association strengths | Graph‑flavoured hierarchical consensus | Hossain 2012 |
| Method | Function | What it does | When it’s useful | Ref |
|---|---|---|---|---|
| intNMF — integrative NMF | intNMF |
Shared non‑negative basis W + per‑omic Hₖ;
consensus clustering of W |
Low‑rank, parts‑based integration; non‑negative data | Chalise 2017 |
| LUCID | LUCID ‡ |
Latent clusters jointly from exposures, omics and an outcome (quasi‑mediation, EM) | Supervised subtyping where an outcome should shape the clusters | Zhao 2024 |
| iCluster / iClusterPlus / iClusterBayes | — | Joint latent‑variable integration, mixed data types, feature selection | Zhang 2022 | |
| moCluster / MOFA / JIVE | — | Low‑rank / factor decomposition of shared + individual variation | Zhang 2022 |
‡ LUCID wraps the LUCIDus package (a
Suggests dependency).
| Tool | Function | Role |
|---|---|---|
| Data nuggets | create_data_nuggets,
nugget_feature_weights |
Compress N observations into M weighted representatives; derive robust feature weights for M‑ABC |
| Data‑nugget clustering | nugget_cluster, Wkmeans,
Whclust |
Cluster the compressed representatives directly (weighted k‑means / weighted Ward), then map back to objects — Big‑Data clustering |
M‑ABC generalises the single‑source Aggregating
Bundles of Clusters algorithm (Amaratunga et al. 2008) to many
heterogeneous sources. For each source it repeats, numsim
times: bootstrap the samples, draw a weighted subsample of
features, cluster, and record co‑clustering. Per‑source
evidence is accumulated into one consensus dissimilarity, then
clustered.
The feature subsample is the heart of the method, and how features are weighted matters:
weighting |
Score behind the selection probability | Best when |
|---|---|---|
"var" |
feature variance (classic) | signal lives in high‑variance features |
"cv" |
coefficient of variation | scale‑free relative dispersion |
"nugget" |
data‑nugget between‑group weighted variance | large N, heavy tails, outliers — robust signal |
"equal" |
flat | ablation / no prior |
Why data nuggets? A data nugget summarises a group
of observations by a centre, a weight (how many
observations it represents) and a scale (its internal spread).
Building nuggets over the samples and asking “which features
separate the nugget centres?” yields a feature‑importance score
that rewards real between‑group signal while discounting within‑nugget
noise — and it scales to very large N because it works on
M ≪ N representatives.
# Inspect the nugget compression and the weights it induces
dn <- create_data_nuggets(t(L[[1]]), max_nuggets = 25)
dn # 60 obs -> 25 nuggets (reduction reported)
w <- nugget_feature_weights(dn, type = "between") # robust feature weights
# Drop straight into M-ABC
M_ABCpp(L, weighting = c("nugget", "nugget"),
nugget_args = list(max_nuggets = 25, seed = 1))If the reference datanugget package is installed you can
swap in its engine with
create_data_nuggets(x, engine = "datanugget"); otherwise
the built‑in native engine is used (no extra dependency).
mosaic_labels(fit, k = 3) # extract a partition from any result
cluster_agreement(a, b) # ARI, NMI, pairwise Jaccard
compare_clusterings(D1, D2, NC = 3) # cluster two dissimilarities & comparecluster_agreement() rule of thumb (ARI):
>0.90 excellent · 0.75–0.90 good ·
0.50–0.75 moderate · <0.50 weak.
The package also ships a full downstream analysis suite (Layer 4). These guard their heavier dependencies via Suggests and raise an informative error if one is missing.
| Purpose | Functions | Needs |
|---|---|---|
| Cross‑method comparison | ReorderToReference, SimilarityMeasure,
CompareSilCluster, CompareSvsM,
CompareInteractive, SelectnrClusters |
base / plotrix |
| Source weighting | DetermineWeight_SilClust,
DetermineWeight_SimClust |
base |
| Cluster navigation | FindCluster, FindElement,
SharedComps, TrackCluster,
ChooseCluster |
base |
| Characterisation | CharacteristicFeatures, FeatSelection,
FeaturesOfCluster |
base |
| Dendrograms / colours | ComparePlot, Cyclogram,
ClusterPlot, ColorPalette,
LabelPlot |
circlize, plotrix |
| Heatmaps / boxplots | HeatmapPlot, HeatmapSelection,
SimilarityHeatmap, BoxPlotDistance |
gplots, ggplot2,
gridExtra |
| Feature / profile plots | BinFeaturesPlot_SingleData,
BinFeaturesPlot_MultipleData,
ContFeaturesPlot, ProfilePlot |
plotrix, Biobase |
| Differential expression | DiffGenes, DiffGenesSelection,
FindGenes |
limma |
| Pathway analysis | PathwayAnalysis, Pathways,
PathwaysIter, PathwaysSelection,
PreparePathway, PlotPathways,
Geneset.intersect |
MLP, biomaRt,
org.Hs.eg.db |
| Shared genes/paths/features | SharedGenesPathsFeat, SharedSelection,
SharedSelectionLimma, SharedSelectionMLP |
base / limma / MLP |
mosaic_toy — a bundled two‑source
example (60 shared samples, 3 planted groups) for instant
experimentation and the test suite.mosaic_sim() — generate your own
multi‑source data with controllable group structure, informative‑feature
counts and effect size.The methods are designed for, and have been evaluated on, multi‑omics benchmarks in the spirit of the LUCIDus simulated‑HELIX data, TCGA subtype cohorts, and the MCF7 compound–fingerprint–transcriptome panel used in the companion manuscript.
{DistM, Clust}
return makes any two methods comparable and any result re‑usable.testthat suite covers the
data‑nugget engine, every weighting scheme, the M‑ABC family, distances
and metrics.@Manual{MosaiClusteR2026,
title = {MosaiClusteR: An Umbrella Framework for Multi-Source and
Multi-Omics Clustering},
author = {Osang'ir, Bernard Isekah and Van Moerbeke, Marijke and
Claesen, Jürgen and Gupta, Surya and Cabrera, Javier and
Amaratunga, Dhammika and Shkedy, Ziv},
year = {2026},
note = {R package version 0.1.0},
url = {https://github.com/lucp12891/MosaiClusteR}
}MosaiClusteR is an open, extensible framework — contributions of new integration methods, distances, and evaluation tools are very welcome.