OSCR

Transcriptomic analysis in autism spectrum disorder suggests three molecular subtypes with distinct phenotypic profiles and functional pathways.

Code ↔ Paper

6 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 6 matches · 1 of them tie a paragraph to a whole file, not to given lines: a weak match, whose lines are not tinted
  1. [1] § Methods › Identification of genes associated with ASD core symptoms ↔ 1.rlm_single_task.R, lines 1–73 · score 0.96 · considered symptom related, randomly subsampled, rlm function, MASS package, ADOS modules, core symptom
  2. [2] § Methods › RNA-seq dataset for subtyping analysis ↔ 1.rlm_single_task.R, lines 1–73 · score 0.60 · ASD symptoms, core symptom, genes associated, ADOS, ADI, phenotypic
  3. [3] § Methods › Validation analysis ↔ code/02_DEGenesIsoforms/02_01_A_DEGenes.R, lines 1–47 · score 0.56 · SeqBatch, limma, ancestry, covariates, pipeline, sex
  4. [4] § Methods › Validation analysis ↔ code/02_DEGenesIsoforms/02_01_B_DEGenes.R, lines 1–44 · score 0.56 · SeqBatch, limma, ancestry, covariates, pipeline, sex
  5. [5] § Methods › Subtyping using nonnegative matrix factorization (NMF) ↔ 2.NMF_random_n30_1711ASD_5pheno_686_genes_overlap_genes.R, the whole file · a weak match · score 0.55 · optimal factorization rank, NMF, matrix, gene
  6. [6] § Methods › Validation analysis ↔ code/01_RNAseqProcessing/01_02_A_CountsProcessing.R, lines 207–251 · score 0.51 · Brain Bank, age, BA41, BA17, seq, ASD

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

R · 129 lines · 5.6 KB · no license · 2 matches

  1. ## This script is to obtain genes related to trait.
  2. ## And we performed 100 replicates, each time randomly subsampled 90% of the participants (N = 1,540)
  3. ## and testing the associations between gene expression and ASD symptoms using the rlm function from the MASS package.
  4. ## Each replicate may take up to 2 hours, so we recommend running multiple tasks simultaneously.
  5. ## Once all replicates have finished, a data frame containing the coefficients and p-values will be generated.
  6. ## And Genes associated with one core symptom measure (P value < 0.05) across all replicates were considered symptom-related genes,
  7. ## and genes related to at least one core symptom were retained.
  8. rm(list = ls())
  9. gc()
  10. # -------------------------- Load packages & parse command line arguments --------------------------
  11. library(data.table)
  12. library(MASS)
  13. library(sfsmisc)
  14. library(optparse) # Parse command line arguments (install if needed: install.packages("optparse"))
  15. # Parse command line arguments
  16. option_list <- list(
  17. make_option(c("-p", "--pheno"), type = "character", help = "Phenotype name", metavar = "character"),
  18. make_option(c("-b", "--boot_id"), type = "integer", help = "Bootstrap ID", metavar = "integer"),
  19. make_option(c("-d", "--data_path"), type = "character", help = "Path to data file", metavar = "character"),
  20. make_option(c("-o", "--output_dir"), type = "character", help = "Output directory", metavar = "character"),
  21. make_option(c("-r", "--sample_ratio"), type = "numeric", default = 0.9, help = "Sampling ratio [default: 0.9]")
  22. )
  23. opt <- parse_args(OptionParser(option_list = option_list))
  24. # Validate required arguments
  25. if (is.null(opt$pheno) || is.null(opt$boot_id) || is.null(opt$data_path) || is.null(opt$output_dir)) {
  26. stop("Must specify: --pheno phenotype name --boot_id bootstrap ID --data_path data path --output_dir output directory")
  27. }
  28. # Define parameters (simplify subsequent code)
  29. pheno <- opt$pheno
  30. boot_id <- opt$boot_id
  31. data_path <- opt$data_path
  32. output_dir <- opt$output_dir
  33. sample_ratio <- opt$sample_ratio
  34. # -------------------------- Data reading and preprocessing --------------------------
  35. cat(paste0("[", Sys.time(), "] Start processing: phenotype = ", pheno, " | Bootstrap = ", boot_id, "\n"))
  36. # Read data
  37. dat_raw <- fread(data_path, na.strings = c("", "NA", "NaN"))
  38. dat <- as.data.frame(dat_raw)
  39. # Preprocessing
  40. dat$sex <- factor(dat$sex)
  41. dat$age_at_ados <- as.numeric(dat$age_at_ados)
  42. # New: handle ados_module type (adjust according to actual data type; if numeric, change to as.numeric)
  43. dat$ados_module <- factor(dat$ados_module)
  44. colnames(dat) <- gsub("-", "_", colnames(dat))
  45. genes <- colnames(dat)[2:15534] # Gene column range (adjust as needed)
  46. # Filter missing phenotype + missing covariates (critical: avoid model errors)
  47. # Construct dynamic filtering condition: for adi_r phenotypes additionally require non‑missing ados_module
  48. filter_cond <- !is.na(dat[[pheno]]) & !is.na(dat$sex)
  49. if (grepl("^adi_r", pheno)) { # Check if phenotype starts with adi_r
  50. filter_cond <- filter_cond & !is.na(dat$age_at_ados) & !is.na(dat$ados_module)
  51. } else {
  52. filter_cond <- filter_cond & !is.na(dat$age_at_ados)
  53. }
  54. dat_pheno <- dat[filter_cond, ]
  55. cat(paste0("[", Sys.time(), "] Sample size after filtering: ", nrow(dat_pheno), "\n"))
  56. # -------------------------- Core functions --------------------------
  57. # Generate bootstrap sample
  58. #load(paste0("../bin/bootstrap_2023_samples_list_",pheno,".Rdat"))
  59. #load("bootstrap_samples_list.Rdat")
  60. load("bootstrap_1711_samples_list_5_pheno.Rdat")
  61. dat_boot_ids <- sampled_id_list[[boot_id]]
  62. #dat_boot_ids <- random_samples[[boot_id]]
  63. dat_boot <- dat_pheno[match(dat_boot_ids, dat_pheno$V1), ]
  64. cat(paste0("[", Sys.time(), "] Bootstrap sample size: ", nrow(dat_boot), "\n"))
  65. # 2. RLM analysis (with tryCatch error handling)
  66. run_rlm <- function(gene, pheno, data) {
  67. # Filter out invalid gene values
  68. idx_valid <- !is.infinite(data[[gene]]) & !is.na(data[[gene]])
  69. data_valid <- data[idx_valid, ]
  70. # Core modification: dynamically construct formula based on phenotype prefix
  71. if (grepl("^adi_r", pheno)) {
  72. # adi_r phenotype: covariates = interaction of ados_module and age_at_ados + sex
  73. formula_str <- as.formula(paste(pheno, "~", gene, "+ ados_module:age_at_ados + sex"))
  74. } else {
  75. # Non‑adi_r phenotype: keep original covariates (age_at_ados + sex)
  76. formula_str <- as.formula(paste(pheno, "~", gene, "+ age_at_ados + sex"))
  77. }
  78. rlm_model <- rlm(formula_str, data = data_valid)
  79. # Extract coefficient results
  80. coef_summary <- summary(rlm_model)$coefficients
  81. # Wald test (f.robftest)
  82. wald_test <- f.robftest(rlm_model, var = gene)
  83. # Compile results
  84. res <- data.frame(
  85. gene = gene,
  86. phenotype = pheno,
  87. value = coef_summary[2, "Value"],
  88. SE = coef_summary[2, "Std. Error"],
  89. t_value = coef_summary[2, "t value"],
  90. p_value = as.numeric(wald_test$p.value), # Wald test p‑value
  91. bootstrap_id = NA # Will be filled with bootstrap ID later
  92. )
  93. return(res)
  94. }
  95. # -------------------------- Execute analysis --------------------------
  96. # Batch run RLM
  97. rlm_res <- lapply(genes, function(g) run_rlm(g, pheno, dat_boot))
  98. rlm_res_df <- do.call(rbind, rlm_res)
  99. # Fill bootstrap ID
  100. rlm_res_df$bootstrap_id <- boot_id
  101. # -------------------------- Save results --------------------------
  102. # Output file name: phenotype_BootstrapID.Rdat
  103. output_file <- file.path(output_dir, sprintf("RLM_%s_boot_%d.Rdat", pheno, boot_id))
  104. save(rlm_res_df, file = output_file)
  105. cat(paste0("[", Sys.time(), "] Done! Results saved to: ", output_file, "\n"))
  106. cat(paste0("Number of valid results: ", sum(!is.na(rlm_res_df$p_value)), "/", nrow(rlm_res_df), "\n"))
  107. # Clean memory
  108. rm(list = ls())
  109. gc()

1.rlm_single_task.R at commit ad4b502, no license · at the source

Overview

Authors: Tao Pang1, Xiangyu Zheng1, Jia-Jia Liu2, Lin Lu1, Li Yang1,3, Suhua Chang1,3,4
  1. Peking University Sixth Hospital, Peking University Institute of Mental Health, NHC Key Laboratory of Mental Health (Peking University), National Clinical Research Center for Mental Disorders (Peking University Sixth Hospital), Beijing, China
  2. School of Nursing, Peking University, Beijing, China
  3. Beijing Key Laboratory for Big Data Innovative Application of Child and Adolescent Mental Disorders, Beijing, China
  4. Henan Collaborative Innovation Center of Prevention and Treatment of Mental Disorder, the Second Affiliated Hospital of Xinxiang Medical University, Xinxiang, Henan China
Journal: Communications biology, volume 9, issue 1, article 883
Dates: received 17 September 2025; accepted 2 April 2026; published online 27 April 2026
Type: Research article · Language: English
License: CC BY-NC-ND
Identifiers: DOI 10.1038/s42003-026-10059-5 · PMID 42045359 · PMCID PMC13324688 · OpenAlex W7155723103
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), human (organism), autism (population), cellular / molecular (subfield)
Methods: Statistics, Smoothing, state filtering, decompositions, Machine learning, Preprocessing, Connectivity
Keywords: Transcriptomics, Autism spectrum disorders
MeSH: Autism Spectrum Disorder*, Gene Expression Profiling*, Transcriptome*, Clustering Algorithms, Humans, Phenotype (* major topic)
Topic: Autism Spectrum Disorder Research (Cognitive Neuroscience, Neuroscience), according to OpenAlex
Funding: National Natural Science Foundation of China (National Science Foundation of China) (82471565)
Citations: cited by 1 paper (Europe PMC); 62 references in the paper

Abstract

The abstract is not reproduced here: the paper's license (CC BY-NC-ND) does not allow it. Read it in the paper, at the publisher or on Europe PMC.

Repositories

Its files are read in the Code ↔ Paper reader above, with 6 matches between paragraphs and lines of code.

dhglab/Broad-transcriptomic-dysregulation-across-the-cerebral-cortex-in-ASD

License: none: the authors keep all their rights
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: 6821cc55aaf17879a4e0d8eec454d98b2f87dd6a, 10 August 2022
Languages: R (58), C (10), C/C++ (6), Shell (1)
Size: 232 files, 75 scripts
Software Heritage: not archived
Found in: “Data availability”
Holds: README, environment (code/01_RNAseqProcessing/earth-infGenes/DESCRIPTION), tests, documentation
Not found: license file, CITATION.cff, continuous integration
Tools: limma (12 files), ggplot2 (10 files), WGCNA (8 files), reshape2 (6 files), mgcv (3 files), edgeR (2 files), ggpubr (2 files), lmerTest (2 files), FastQC (1 file), igraph (1 file), randomForest (1 file), SAMtools (1 file), Seurat (1 file), STAR (1 file)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
76 files

PerperPKU/SSC_subtype_NMF

License: none: the authors keep all their rights
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Commit: ad4b50205b0466f1433ba194552f45ad1be2ae9f, 24 March 2026
Languages: R (3)
Size: 5 files, 3 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README
Not found: license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: data.table (2 files)
Availability: 1 check, the latest on 30 September 2026: the link answers
  • 30 September 2026: the link answers
4 files

Zenodo 19124178

License: CC-BY-4.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Size: 1 file, 0 scripts
Software Heritage: not checked
Found in: “Code availability”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)
At the source:

Zenodo 19124179

License: CC-BY-4.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Languages: R (3)
Size: 3 files, 3 scripts
Software Heritage: not checked
Found in: the references
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: data.table (2 files)
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)
3 files
At the source:

Code availability statement

The paper has a code availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s42003-026-10059-5.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 4 repositories of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 81 scripts, each with its path and the digest of its content;
  • 6 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

Datasets cited

Code and data availability statement

The paper has a code and data availability statement. Its license (CC BY-NC-ND) does not allow reproducing it here; in short, from what the harvester recognized in it:

Read it in the paper: doi.org/10.1038/s42003-026-10059-5.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 6 authors, 2 keywords, 6 MeSH terms, 1 funder, 60 references.

Cite

This paper

Pang, T., Zheng, X., Liu, J.-J., Lu, L., Yang, L., & Chang, S. (2026). Transcriptomic analysis in autism spectrum disorder suggests three molecular subtypes with distinct phenotypic profiles and functional pathways. Communications biology, 9(1), 883. https://doi.org/10.1038/s42003-026-10059-5

BibTeX

@article{pang2026transcriptomic,
author = {Pang, Tao and Zheng, Xiangyu and Liu, Jia-Jia and Lu, Lin and Yang, Li and Chang, Suhua},
title = {{Transcriptomic analysis in autism spectrum disorder suggests three molecular subtypes with distinct phenotypic profiles and functional pathways}},
journal = {Communications biology},
year = {2026},
month = apr,
volume = {9},
number = {1},
pages = {883},
publisher = {Nature Publishing Group},
issn = {2399-3642},
doi = {10.1038/s42003-026-10059-5},
url = {https://doi.org/10.1038/s42003-026-10059-5},
pmid = {42045359},
pmcid = {PMC13324688}
}

RIS

TY - JOUR
AU - Pang, Tao
AU - Zheng, Xiangyu
AU - Liu, Jia-Jia
AU - Lu, Lin
AU - Yang, Li
AU - Chang, Suhua
TI - Transcriptomic analysis in autism spectrum disorder suggests three molecular subtypes with distinct phenotypic profiles and functional pathways
T2 - Communications biology
J2 - Commun Biol
PY - 2026
DA - 2026/04/27
VL - 9
IS - 1
SP - 883
SN - 2399-3642
PB - Nature Publishing Group
DO - 10.1038/s42003-026-10059-5
UR - https://doi.org/10.1038/s42003-026-10059-5
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s42003-026-10059-5",
"type": "article-journal",
"title": "Transcriptomic analysis in autism spectrum disorder suggests three molecular subtypes with distinct phenotypic profiles and functional pathways",
"container-title": "Communications biology",
"author": [
{
"family": "Pang",
"given": "Tao"
},
{
"family": "Zheng",
"given": "Xiangyu"
},
{
"family": "Liu",
"given": "Jia-Jia"
},
{
"family": "Lu",
"given": "Lin"
},
{
"family": "Yang",
"given": "Li"
},
{
"family": "Chang",
"given": "Suhua"
}
],
"container-title-short": "Commun Biol",
"volume": "9",
"issue": "1",
"page": "883",
"DOI": "10.1038/s42003-026-10059-5",
"PMID": "42045359",
"PMCID": "PMC13324688",
"ISSN": "2399-3642",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s42003-026-10059-5",
"language": "en",
"issued": {
"date-parts": [
[
2026,
4,
27
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: FastQC, randomForest, STAR, 10 other tools, genetics / omics, 1 reference
[2] doi:10.1038/s41467-026-76675-1 [code]
Long-read proteogenomic atlas of human neuronal differentiation reveals isoform diversity informing neurodevelopmental risk mechanisms.
Journal: Nature communications
In common: STAR, WGCNA, SAMtools, 8 other tools, autism, genetics / omics, 1 reference
[3] doi:10.1016/j.xhgg.2026.100652 [code]
CRISPR-engineered deletion of POGZ alters transcription factor binding at promoters of genes involved in synaptic signaling.
Journal: HGG advances
In common: STAR, WGCNA, SAMtools, 5 other tools, autism, cellular / molecular, 4 references
[4] doi:10.1038/s42003-026-10957-8 [code]
Brain defence by the extracellular matrix protein Cochlin.
Journal: Communications biology
In common: randomForest, WGCNA, SAMtools, 8 other tools, cellular / molecular
[5] doi:10.1038/s41467-026-74753-y [code]
A human-specific microRNA controls the timing of excitatory synaptogenesis.
Journal: Nature communications
In common: STAR, SAMtools, edgeR, 7 other tools, cellular / molecular, 1 reference
[6] doi:10.1038/s41380-026-03578-4 [code]
Assessing molecular gene by treatment interactions using a population of neural progenitors exposed to valproic acid and lithium.
Journal: Molecular psychiatry
In common: FastQC, STAR, SAMtools, 6 other tools, genetics / omics, cellular / molecular, 1 reference
[7] doi:10.1016/j.xcrm.2026.102766 [code]
A longitudinal single-cell and spatial multiomic atlas of pediatric high-grade glioma.
Journal: Cell reports. Medicine
In common: WGCNA, edgeR, limma, 7 other tools, genetics / omics, cellular / molecular
[8] doi:10.1038/s41380-026-03585-5 [code]
Multiomics analysis identifies VPA-induced changes in neural progenitor cells, ventricular-like regions, and cellular microenvironment in dorsal forebrain organoids.
Journal: Molecular psychiatry
In common: WGCNA, edgeR, limma, 5 other tools, autism, genetics / omics, 2 references
[9] doi:10.1016/j.cpblue.2026.100007 [code]
An integrated single-cell and spatial proteotranscriptomics atlas of fibroblast-driven immunoregulation within the human adult oral cavity.
Journal: Cell press blue
In common: STAR, SAMtools, edgeR, 7 other tools
[10] doi:10.1016/j.xgen.2026.101278 [code]
Single-cell profiling of DNA methylation in autism spectrum disorder prefrontal cortex reveals distinct regulatory and aging signatures.
Journal: Cell genomics
In common: FastQC, STAR, SAMtools, autism, genetics / omics, 6 references

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.