OSCR

PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning.

Code ↔ Paper

3 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 3 matches
  1. [1] § 2 Materials and methods › 2.4 Feature integration ↔ models.py, lines 189–225 · score 0.79 · binary classification, directly concatenates, multi class, refined features, activation, MLP
  2. [2] § 2 Materials and methods › 2.1 Overview of PEARL ↔ models.py, lines 189–225 · score 0.70 · refined features, unifies, max, min, component, concatenation
  3. [3] § 2 Materials and methods › 2.3 Feature refinement ↔ models.py, lines 50–117 · score 0.52 · ReLU, activation function, dropout, layer

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 235 lines · 9 KB · no license · 3 matches

  1. import torch
  2. import torch.nn as nn
  3. import torch.nn.functional as F
  4. from torch_geometric.nn import SSGConv
  5. device = torch.device('cuda' if torch.cuda.is_available() else 'cpu')
  6. def xavier_init(m):
  7. if type(m) == nn.Linear:
  8. nn.init.xavier_normal_(m.weight)
  9. if m.bias is not None:
  10. m.bias.data.fill_(0.0)
  11. class SSGGraphConvolution(nn.Module):
  12. def __init__(self, in_features, out_features, K, alpha):
  13. super(SSGGraphConvolution, self).__init__()
  14. self.conv = SSGConv(in_features, out_features, K=K, alpha=alpha)
  15. def forward(self, x, adj):
  16. if adj.is_sparse:
  17. edge_index, edge_weight = adj._indices(), adj._values()
  18. else:
  19. edge_index = adj.nonzero().t().contiguous()
  20. edge_weight = adj[edge_index[0], edge_index[1]]
  21. return self.conv(x, edge_index, edge_weight)
  22. class GCN_E(nn.Module):
  23. def __init__(self, in_dim, hidden_dim, out_dim, K, alpha):
  24. super(GCN_E, self).__init__()
  25. self.gc1 = SSGGraphConvolution(in_dim, hidden_dim, K=K, alpha=alpha)
  26. self.gc2 = SSGGraphConvolution(hidden_dim, out_dim, K=K, alpha=alpha)
  27. self.dropout = nn.Dropout(0.5)
  28. def forward(self, x, adj):
  29. x = F.relu(self.gc1(x, adj))
  30. x = self.dropout(x)
  31. x = self.gc2(x, adj)
  32. return x
  33. class Classifier_1(nn.Module):
  34. def __init__(self, in_dim, out_dim):
  35. super().__init__()
  36. self.clf = nn.Sequential(nn.Linear(in_dim, out_dim))
  37. self.clf.apply(xavier_init)
  38. def forward(self, x):
  39. x = self.clf(x)
  40. return x
  41. class CombinedPoolingMLP(nn.Module):
  42. def __init__(self, num_view, num_cls, hidden_dims=[64], activation='relu',
  43. normalization='batchnorm', dropout_rates=[0.7], residual=True,
  44. input_normalization=False, output_activation=None):
  45. super().__init__()
  46. self.num_cls = num_cls
  47. self.num_view = num_view
  48. self.input_normalization = input_normalization
  49. self.output_activation = output_activation
  50. if input_normalization:
  51. self.input_norm = nn.BatchNorm1d(num_cls * 4)
  52. input_dim = num_cls * 4
  53. layers = []
  54. in_dim = input_dim
  55. for i, hidden_dim in enumerate(hidden_dims):
  56. layers.append(nn.Linear(in_dim, hidden_dim))
  57. layers.append(self.get_activation(activation))
  58. if normalization == 'batchnorm':
  59. layers.append(nn.BatchNorm1d(hidden_dim))
  60. elif normalization == 'layernorm':
  61. layers.append(nn.LayerNorm(hidden_dim))
  62. dropout_rate = dropout_rates[i] if i < len(dropout_rates) else dropout_rates[-1]
  63. layers.append(nn.Dropout(dropout_rate))
  64. if residual and in_dim == hidden_dim:
  65. layers.append(ResidualConnection())
  66. in_dim = hidden_dim
  67. layers.append(nn.Linear(in_dim, num_cls))
  68. if self.output_activation:
  69. layers.append(self.get_activation(self.output_activation))
  70. self.mlp = nn.Sequential(*layers)
  71. def get_activation(self, activation):
  72. if activation == 'relu':
  73. return nn.ReLU()
  74. elif activation == 'leaky_relu':
  75. return nn.LeakyReLU()
  76. elif activation == 'elu':
  77. return nn.ELU()
  78. elif activation == 'gelu':
  79. return nn.GELU()
  80. elif activation == 'tanh':
  81. return nn.Tanh()
  82. elif activation == 'sigmoid':
  83. return nn.Sigmoid()
  84. else:
  85. raise ValueError(f"Unsupported activation function: {activation}")
  86. def forward(self, in_list):
  87. stacked = torch.stack(in_list)
  88. max_pooled = torch.max(stacked, dim=0)[0]
  89. min_pooled = torch.min(stacked, dim=0)[0]
  90. avg_pooled = torch.mean(stacked, dim=0)
  91. x = torch.cat([max_pooled, min_pooled, avg_pooled, max_pooled - min_pooled], dim=1)
  92. if self.input_normalization:
  93. x = self.input_norm(x)
  94. return self.mlp(x)
  95. class EnhancedConcatMLPIntegration(nn.Module):
  96. def __init__(self, num_view, num_cls, hidden_dims=[64], activation='relu',
  97. normalization='layernorm', dropout_rates=[0.7], residual=True,
  98. input_normalization=False, output_activation=None):
  99. super().__init__()
  100. self.num_cls = num_cls
  101. self.num_view = num_view
  102. self.input_dim = num_view * num_cls
  103. self.input_normalization = input_normalization
  104. self.output_activation = output_activation
  105. if input_normalization:
  106. self.input_norm = nn.BatchNorm1d(self.input_dim)
  107. layers = []
  108. in_dim = self.input_dim
  109. for i, hidden_dim in enumerate(hidden_dims):
  110. layers.append(nn.Linear(in_dim, hidden_dim))
  111. layers.append(self.get_activation(activation))
  112. if normalization == 'batchnorm':
  113. layers.append(nn.BatchNorm1d(hidden_dim))
  114. elif normalization == 'layernorm':
  115. layers.append(nn.LayerNorm(hidden_dim))
  116. dropout_rate = dropout_rates[i] if i < len(dropout_rates) else dropout_rates[-1]
  117. layers.append(nn.Dropout(dropout_rate))
  118. if residual and in_dim == hidden_dim:
  119. layers.append(ResidualConnection())
  120. in_dim = hidden_dim
  121. layers.append(nn.Linear(in_dim, num_cls))
  122. if self.output_activation:
  123. layers.append(self.get_activation(self.output_activation))
  124. self.mlp = nn.Sequential(*layers)
  125. def get_activation(self, activation):
  126. if activation == 'relu':
  127. return nn.ReLU()
  128. elif activation == 'leaky_relu':
  129. return nn.LeakyReLU()
  130. elif activation == 'elu':
  131. return nn.ELU()
  132. elif activation == 'gelu':
  133. return nn.GELU()
  134. elif activation == 'tanh':
  135. return nn.Tanh()
  136. elif activation == 'sigmoid':
  137. return nn.Sigmoid()
  138. else:
  139. raise ValueError(f"Unsupported activation function: {activation}")
  140. def forward(self, in_list):
  141. x = torch.cat(in_list, dim=1)
  142. if self.input_normalization:
  143. x = self.input_norm(x)
  144. return self.mlp(x)
  145. class ResidualConnection(nn.Module):
  146. def forward(self, x):
  147. return x + self.branch(x)
  148. def branch(self, x):
  149. return x
  150. def init_model_dict(num_view, num_class, dim_list, dim_he_list, dim_hc,
  151. aggregation='combined_pooling',
  152. K=1, alpha=0.7, hidden_dims=[64], activation='relu',
  153. normalization='layernorm', dropout_rates=[0.7], residual=True,
  154. input_normalization=False, output_activation=None):
  155. """Initialize PEARL model components.
  156. Args:
  157. aggregation (str): Feature integration strategy for multi-view fusion.
  158. - 'combined_pooling': CombinedPoolingMLP — unifies max, min, mean,
  159. and difference pooling across views (recommended for multi-class).
  160. - 'concatenation': EnhancedConcatMLPIntegration — directly concatenates
  161. refined features from all views (recommended for binary classification).
  162. """
  163. model_dict = {}
  164. for i in range(num_view):
  165. model_dict["E{:}".format(i+1)] = GCN_E(dim_list[i], dim_he_list[0], dim_he_list[-1], K=K, alpha=alpha).to(device)
  166. model_dict["C{:}".format(i+1)] = Classifier_1(dim_he_list[-1], num_class).to(device)
  167. if num_view >= 2:
  168. if aggregation == 'combined_pooling':
  169. model_dict["C"] = CombinedPoolingMLP(
  170. num_view, num_class, hidden_dims=hidden_dims, activation=activation,
  171. normalization=normalization, dropout_rates=dropout_rates, residual=residual,
  172. input_normalization=input_normalization, output_activation=output_activation
  173. ).to(device)
  174. elif aggregation == 'concatenation':
  175. model_dict["C"] = EnhancedConcatMLPIntegration(
  176. num_view, num_class, hidden_dims=hidden_dims, activation=activation,
  177. normalization=normalization, dropout_rates=dropout_rates, residual=residual,
  178. input_normalization=input_normalization, output_activation=output_activation
  179. ).to(device)
  180. else:
  181. raise ValueError(
  182. f"Unsupported aggregation: '{aggregation}'. "
  183. f"Choose 'combined_pooling' or 'concatenation'."
  184. )
  185. return model_dict
  186. def init_optim(num_view, model_dict, lr_e=1e-4, lr_c=1e-4):
  187. optim_dict = {}
  188. for i in range(num_view):
  189. optim_dict["C{:}".format(i+1)] = torch.optim.Adam(
  190. list(model_dict["E{:}".format(i+1)].parameters()) + list(model_dict["C{:}".format(i+1)].parameters()),
  191. lr=lr_e)
  192. if num_view >= 2:
  193. optim_dict["C"] = torch.optim.Adam(model_dict["C"].parameters(), lr=lr_c)
  194. return optim_dict

models.py at commit 0316950, no license · at the source

Overview

Authors: Quan Zhao1, Jiawen Du2, Muqing Zhou3, Xu-Wen Wang4, Quan Sun5,6, Can Chen1,2,7,8
  1. Carolina Health Informatics Program, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, United States
  2. Department of Biostatistics, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, United States
  3. Department of Genetics, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, United States
  4. Channing Division of Network Medicine, Department of Medicine, Brigham and Women’s Hospital, Harvard Medical School, Boston, MA 02115, United States
  5. Center for Computational and Genomic Medicine, Children’s Hospital of Philadelphia, Phildaelphia, PA 19104, United States
  6. Department of Pathology and Laboratory Medicine, University of Pennsylvania Perelman School of Medicine, Philadelphia, PA 19104, United States
  7. School of Data Science and Society, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, United States
  8. Department of Mathematics, University of North Carolina at Chapel Hill, Chapel Hill, NC 27599, United States
Institutions: University of North Carolina at Chapel Hill (United States); Brigham and Women's Hospital (United States); Harvard University (United States); Children's Hospital of Philadelphia (United States); University of Pennsylvania (United States)
Journal: Bioinformatics (Oxford, England), volume 42, issue 6, article btag253
Dates: received 19 October 2025; accepted 28 April 2026; published online 9 May 2026; in print June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1093/bioinformatics/btag253 · PMID 42108553 · PMCID PMC13224962 · OpenAlex W4410716617
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), Alzheimer's / dementia (population), methods / tools (subfield)
Methods: Statistics, Connectivity, Machine learning
MeSH: Computational Biology*, Deep Learning*, Genomics*, Multiomics*, Software*, Algorithms, Alzheimer Disease, Graph Neural Networks, Humans (* major topic)
Journal subjects: Data and Text Mining
Topic: Bioinformatics and Genomic Networks (Molecular Biology, Biochemistry, Genetics and Molecular Biology), according to OpenAlex
Citations: cited by 1 paper (Europe PMC); 71 references in the paper

Abstract

Motivation: Integrating multi-omics data provides valuable insights into biological processes by capturing information across multiple molecular layers, enabling a comprehensive understanding of complex diseases and driving advancements in precision medicine. However, existing computational methods for multi-omics integration face significant challenges, such as low reliability and poor generalizability, due to the high dimensionality and low sample size nature of omics data.

Results: To address these challenges, we present PEARL (Pearson-Enhanced spectrAl gRaph convoLutional networks), a novel deep graph learning method for biomedical classification and functional important omics features identification. PEARL leverages a simple yet effective learning architecture to achieve superior and robust performance in high-dimensional, low-sample-size multi-omics settings. Our results demonstrate that PEARL significantly outperforms existing state-of-the-art methods on both synthetic and real biomedical datasets. Furthermore, applied to Alzheimer’s disease (AD) brain multi-omics data, features prioritized by PEARL lead to functionally important genes that demonstrate significant enrichment in AD-related pathways. These findings highlight PEARL’s practical utility in biomedical research and its potential to enhance biological interpretability in multi-omics studies.

Availability and implementation: The source code of our computational framework is available at https://github.com/zqq121017/PEARL.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 3 matches between paragraphs and lines of code.

zqq121017/PEARL

License: none: the authors keep all their rights
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: 031695035891bfcf212bd78ab0a948404d5360e0, 31 March 2026
Languages: Python (5)
Size: 100 files, 5 scripts
Software Heritage: not checked
Found in: “Data availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: PyTorch (4 files), NumPy (3 files), pandas (2 files), scikit-learn (2 files), PyTorch Geometric (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
6 files

Availability and implementation

The source code of our computational framework is available at https://github.com/zqq121017/PEARL.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 5 scripts, each with its path and the digest of its content;
  • 3 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability

The ROSMAP dataset was obtained from AMP-AD Knowledge Portal (https://adknowledgeportal.synapse.org/). Omics data of BRCA and LUAD was obtained from The Cancer Genome Atlas Program (TCGA) through Broad GDAC Firehose (https://gdac.broadinstitute.org/). PAM50 breast cancer subtypes of TCGA BRCA patients were obtained through the TCGAbiolinks R package v2.12.6. The source data and code of our computational framework is available at https://github.com/zqq121017/PEARL. The archival version of the code is preserved on Zenodo at https://doi.org/10.5281/zenodo.19875866 (zqq121017 2026).

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 9 MeSH terms, 59 references.

Cite

This paper

Zhao, Q., Du, J., Zhou, M., Wang, X.-W., Sun, Q., & Chen, C. (2026). PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning. Bioinformatics (Oxford, England), 42(6), btag253. https://doi.org/10.1093/bioinformatics/btag253

BibTeX

@article{zhao2026pearl,
author = {Zhao, Quan and Du, Jiawen and Zhou, Muqing and Wang, Xu-Wen and Sun, Quan and Chen, Can},
title = {{PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning}},
journal = {Bioinformatics (Oxford, England)},
year = {2026},
month = jun,
volume = {42},
number = {6},
pages = {btag253},
publisher = {Oxford University Press},
issn = {1367-4803},
doi = {10.1093/bioinformatics/btag253},
url = {https://doi.org/10.1093/bioinformatics/btag253},
pmid = {42108553},
pmcid = {PMC13224962}
}

RIS

TY - JOUR
AU - Zhao, Quan
AU - Du, Jiawen
AU - Zhou, Muqing
AU - Wang, Xu-Wen
AU - Sun, Quan
AU - Chen, Can
TI - PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning
T2 - Bioinformatics (Oxford, England)
J2 - Bioinformatics
PY - 2026
DA - 2026/06/01
VL - 42
IS - 6
SP - btag253
SN - 1367-4803
PB - Oxford University Press
DO - 10.1093/bioinformatics/btag253
UR - https://doi.org/10.1093/bioinformatics/btag253
LA - en
ER -

CSL-JSON

{
"id": "10.1093/bioinformatics/btag253",
"type": "article-journal",
"title": "PEARL: integrative multi-omics classification and omics feature discovery via deep graph learning",
"container-title": "Bioinformatics (Oxford, England)",
"author": [
{
"family": "Zhao",
"given": "Quan"
},
{
"family": "Du",
"given": "Jiawen"
},
{
"family": "Zhou",
"given": "Muqing"
},
{
"family": "Wang",
"given": "Xu-Wen"
},
{
"family": "Sun",
"given": "Quan"
},
{
"family": "Chen",
"given": "Can"
}
],
"container-title-short": "Bioinformatics",
"volume": "42",
"issue": "6",
"page": "btag253",
"DOI": "10.1093/bioinformatics/btag253",
"PMID": "42108553",
"PMCID": "PMC13224962",
"ISSN": "1367-4803",
"publisher": "Oxford University Press",
"URL": "https://doi.org/10.1093/bioinformatics/btag253",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
1
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.3389/fsysb.2026.1873899 [code]
A systems microbiology framework for reproducible multi-dataset omics integration with application to long COVID.
Journal: Frontiers in systems biology
In common: PyTorch Geometric, PyTorch, scikit-learn, 3 other tools, methods / tools, genetics / omics, 1 reference
[2] doi:10.1038/s41467-026-74694-6 [code]
Semi-supervised Omics Factor Analysis (SOFA) disentangles known and latent sources of variation in multi-omic data.
Journal: Nature communications
In common: PyTorch, scikit-learn, pandas, 2 other tools, methods / tools, genetics / omics, 2 references
[3] doi:10.1093/bioinformatics/btag540 [code]
Deciphering spatial heterogeneity by multimodal spatial transcriptomics modelling with SpatialModal.
Journal: Bioinformatics (Oxford, England)
In common: PyTorch Geometric, PyTorch, scikit-learn, 3 other tools, methods / tools, Alzheimer's / dementia, genetics / omics
[4] doi:10.1093/bib/bbag259 [code]
PVAED: prior-guided variational autoencoders with diffusion denoising for interpretable single-cell representation learning.
Journal: Briefings in bioinformatics
In common: PyTorch, scikit-learn, pandas, 2 other tools, genetics / omics, 2 references
[5] doi:10.1038/s41467-026-71391-2 [code]
Accelerating Leigh syndrome drug discovery through deep learning screening in brain organoids.
Journal: Nature communications
In common: PyTorch Geometric, PyTorch, scikit-learn, 3 other tools, 1 reference
[6] doi:10.1038/s41592-026-03194-8 [code]
Beyond benchmarking: an expert-guided consensus approach to spatially aware clustering.
Journal: Nature methods
In common: PyTorch Geometric, PyTorch, scikit-learn, 3 other tools, methods / tools, genetics / omics
[7] doi:10.1093/bib/bbag298 [code]
Empowering multifaceted analysis of spatial transcriptomics data with RGAST.
Journal: Briefings in bioinformatics
In common: PyTorch Geometric, PyTorch, scikit-learn, 3 other tools, methods / tools, genetics / omics
[8] doi:10.1002/advs.75969 [code]
Accurately Deciphering Tissue Heterogeneity From Spatial Multi-Modal and Multi-Omics With STransformer.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: PyTorch Geometric, PyTorch, scikit-learn, 3 other tools, Alzheimer's / dementia, genetics / omics
[9] doi:10.1371/journal.pcbi.1014327 [code]
Supervised deep learning with gene functional annotation for cell classification.
Journal: PLoS computational biology
In common: PyTorch Geometric, PyTorch, scikit-learn, 3 other tools, Alzheimer's / dementia, genetics / omics
[10] doi:10.1038/s41746-026-02735-x [code]
Multimodal interpretable deep learning for transcriptome-informed precision oncology and drug mechanism analysis.
Journal: NPJ digital medicine
In common: PyTorch, scikit-learn, pandas, 2 other tools, genetics / omics, 2 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.