OSCR

An in silico protocol for predicting genetic biomarkers in rare diseases: a case study in sporadic amyotrophic lateral sclerosis.

Code ↔ Paper

13 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 13 matches
  1. [1] § Methods › Data collection ↔ Project-script.py, lines 20–53 · score 0.84 · STRONGEST SNP RISK, CHR_ID, CHR_POS, MAPPED_GENE, ALLELE, TRAIT
  2. [2] § Methods › Data collection ↔ 1.py, lines 22–43 · score 0.84 · STRONGEST SNP RISK, CHR_ID, CHR_POS, MAPPED_GENE, ALLELE, TRAIT
  3. [3] § Methods › Data preprocessing ↔ Project-script.py, lines 226–303 · score 0.75 · positive negative SNP, dissimilar pairs, model training, intergenic, preprocessing, positional
  4. [4] § Methods › Machine learning model building ↔ 2.py, lines 18–161 · score 0.74 · max_depth, n_estimators, class weight, ROC AUC, Random Forest, split
  5. [5] § Materials and equipment › Visualization and results interpretation ↔ Project-script.py, lines 102–158 · score 0.70 · Precision Recall curve, confusion matrix, ROC curve, metrics, classifier, predictions
  6. [6] § Methods › Machine learning model building ↔ Project-script.py, lines 85–99 · score 0.67 · max_depth, n_estimators, class weight, Random Forest, classification, training
  7. [7] § Methods › Data collection ↔ Project-script.py, lines 20–53 · score 0.62 · riskAllele, mappedGenes, pValue, locations, TSV, chromosome
  8. [8] § Methods › Data collection ↔ 2.py, lines 164–224 · score 0.62 · riskAllele, mappedGenes, pValue, locations, TSV, chromosome
  9. [9] § Methods › Model validation ↔ 1.py, lines 228–294 · score 0.55 · Logistic Regression, Ridge Regression, Random Forest, metrics, model
  10. [10] § Methods › Data preprocessing ↔ Project-script.py, lines 56–82 · score 0.53 · gene related features, mapped genes, intergenic, numeric, Chromosomes, positions
  11. [11] § Methods › Data preprocessing ↔ Project-script.py, lines 226–303 · score 0.53 · chr_diff, pos_diff, dissimilarity, preprocessing, positional, SNPs
  12. [12] § Methods › Data preprocessing ↔ 2.py, lines 18–161 · score 0.52 · chr_diff, pos_diff, dissimilarity, chromosome, positional, SNPs
  13. [13] § Results › Performance of the prediction model ↔ Project-script.py, lines 102–158 · score 0.51 · Precision Recall curve, ROC curve, class, AUC, model, prediction

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 307 lines · 12 KB · CC-BY-4.0 · 8 matches

  1. import pandas as pd
  2. import numpy as np
  3. from sklearn.ensemble import RandomForestClassifier
  4. from sklearn.preprocessing import StandardScaler
  5. from sklearn.pipeline import Pipeline
  6. from sklearn.model_selection import train_test_split
  7. from sklearn.metrics import (classification_report, roc_auc_score,
  8. average_precision_score, confusion_matrix,
  9. precision_recall_curve, roc_curve)
  10. from sklearn.impute import SimpleImputer
  11. import matplotlib.pyplot as plt
  12. import joblib
  13. import warnings
  14. import seaborn as sns
  15. from tqdm import tqdm
  16. warnings.filterwarnings('ignore')
  17. def load_data():
  18. """Load and prepare positive (sALS) and negative (GWAS) SNPs"""
  19. # Load known sALS SNPs
  20. sals_df = pd.read_csv("EFO_0001357_associations_export.tsv", sep='\t')
  21. sals_df['target'] = 1
  22. # Load GWAS data for negative examples
  23. gwas_df = pd.read_csv("GWAS_cache.tsv", sep='\t', low_memory=False)
  24. # Process sALS SNPs
  25. sals_df['rsID'] = sals_df['riskAllele'].str.split('-').str[0]
  26. sals_df[['chromosome', 'position']] = sals_df['locations'].str.extract(r'([XYMT\d]+):(\d+)')
  27. sals_df['position'] = pd.to_numeric(sals_df['position'], errors='coerce')
  28. # Process GWAS SNPs
  29. if 'SNPS' in gwas_df.columns:
  30. gwas_df['rsID'] = gwas_df['SNPS'].str.split('-').str[0]
  31. else:
  32. gwas_df['rsID'] = gwas_df['STRONGEST SNP-RISK ALLELE'].str.split('-').str[0]
  33. gwas_df['chromosome'] = gwas_df['CHR_ID']
  34. gwas_df['position'] = gwas_df['CHR_POS'].astype(str).str.split(' x').str[0].str.split(';').str[0]
  35. gwas_df['position'] = pd.to_numeric(gwas_df['position'], errors='coerce')
  36. gwas_df['mappedGenes'] = gwas_df['MAPPED_GENE']
  37. gwas_df['pValue'] = gwas_df['P-VALUE']
  38. # Filter non-ALS SNPs
  39. als_keywords = ["amyotrophic lateral sclerosis", "ALS", "motor neuron disease"]
  40. gwas_df['is_als'] = gwas_df['DISEASE/TRAIT'].str.contains('|'.join(als_keywords), case=False, na=False)
  41. neg_df = gwas_df[~gwas_df['is_als']].copy()
  42. neg_df = neg_df[~neg_df['rsID'].isin(sals_df['rsID'])]
  43. neg_df['target'] = 0
  44. return sals_df, neg_df
  45. def prepare_features(df):
  46. """Prepare features for similarity analysis"""
  47. # Gene-related features
  48. df['gene_count'] = df['mappedGenes'].apply(lambda x: len(str(x).split(','))) if 'mappedGenes' in df.columns else 0
  49. df['is_intergenic'] = df['mappedGenes'].isna().astype(int) if 'mappedGenes' in df.columns else 1
  50. # p-value transformation
  51. if 'pValue' in df.columns:
  52. df['pValue'] = df['pValue'].replace('NR', np.nan)
  53. df['pValue'] = pd.to_numeric(df['pValue'], errors='coerce')
  54. df['log_pvalue'] = -np.log10(df['pValue'].replace(0, 1e-300))
  55. df['log_pvalue'] = df['log_pvalue'].fillna(df['log_pvalue'].median())
  56. else:
  57. df['log_pvalue'] = 0
  58. # Chromosome encoding
  59. chr_map = {str(i): i for i in range(1, 23)}
  60. chr_map.update({'X': 23, 'Y': 24, 'MT': 25})
  61. df['chr_encoded'] = df['chromosome'].map(chr_map).fillna(26)
  62. # Position handling
  63. if 'position' in df.columns:
  64. df['position'] = df['position'].fillna(df['position'].median())
  65. else:
  66. df['position'] = 0
  67. return df
  68. def train_model(X_train, y_train):
  69. """Train the similarity prediction model"""
  70. model = Pipeline([
  71. ('imputer', SimpleImputer(strategy='median')),
  72. ('scaler', StandardScaler()),
  73. ('clf', RandomForestClassifier(
  74. n_estimators=100,
  75. max_depth=5,
  76. class_weight='balanced',
  77. random_state=42,
  78. n_jobs=-1
  79. ))
  80. ])
  81. model.fit(X_train, y_train)
  82. return model
  83. def evaluate_model(model, X_test, y_test):
  84. """Evaluate model performance with comprehensive metrics"""
  85. print("\n=== Model Performance Evaluation ===")
  86. # Predictions
  87. y_pred = model.predict(X_test)
  88. y_proba = model.predict_proba(X_test)
  89. # Handle single class case
  90. if y_proba.shape[1] == 1:
  91. y_proba = np.column_stack([y_proba, 1 - y_proba])
  92. # Classification metrics
  93. print("\nDetailed Classification Report:")
  94. print(classification_report(y_test, y_pred))
  95. # Confusion matrix
  96. cm = confusion_matrix(y_test, y_pred)
  97. plt.figure(figsize=(6, 6))
  98. sns.heatmap(cm, annot=True, fmt='d', cmap='Blues',
  99. xticklabels=['Negative', 'Positive'],
  100. yticklabels=['Negative', 'Positive'])
  101. plt.title('Confusion Matrix')
  102. plt.xlabel('Predicted')
  103. plt.ylabel('Actual')
  104. plt.show()
  105. # ROC Curve
  106. fpr, tpr, _ = roc_curve(y_test, y_proba[:, 1])
  107. roc_auc = roc_auc_score(y_test, y_proba[:, 1])
  108. plt.figure(figsize=(6, 6))
  109. plt.plot(fpr, tpr, color='darkorange', lw=2,
  110. label=f'ROC curve (AUC = {roc_auc:.2f})')
  111. plt.plot([0, 1], [0, 1], color='navy', lw=2, linestyle='--')
  112. plt.xlabel('False Positive Rate')
  113. plt.ylabel('True Positive Rate')
  114. plt.title('Receiver Operating Characteristic')
  115. plt.legend(loc="lower right")
  116. plt.show()
  117. # Precision-Recall Curve
  118. precision, recall, _ = precision_recall_curve(y_test, y_proba[:, 1])
  119. avg_precision = average_precision_score(y_test, y_proba[:, 1])
  120. plt.figure(figsize=(6, 6))
  121. plt.plot(recall, precision, color='blue', lw=2,
  122. label=f'Precision-Recall (AP = {avg_precision:.2f})')
  123. plt.xlabel('Recall')
  124. plt.ylabel('Precision')
  125. plt.title('Precision-Recall Curve')
  126. plt.legend(loc="lower left")
  127. plt.show()
  128. # Key metrics
  129. print("\nKey Performance Metrics:")
  130. print(f"- ROC-AUC Score: {roc_auc:.3f}")
  131. print(f"- Average Precision: {avg_precision:.3f}")
  132. print(f"- Accuracy: {np.mean(y_pred == y_test):.3f}")
  133. def predict_similar_snps(model, reference_snps, candidate_snps, top_n=10):
  134. """Predict similar SNPs using the trained model with progress tracking"""
  135. results = []
  136. feature_names = ['chr_diff', 'pos_diff', 'pval_diff', 'gene_diff', 'intergenic_diff']
  137. total_snps = len(reference_snps)
  138. print(f"\nStarting prediction for {total_snps} reference SNPs...")
  139. print(f"Comparing against {len(candidate_snps)} candidate SNPs")
  140. print(f"Finding top {top_n} similar SNPs for each reference SNP\n")
  141. # Initialize progress bar
  142. pbar = tqdm(reference_snps.iterrows(), total=total_snps, desc="Processing SNPs")
  143. for idx, (_, ref_snp) in enumerate(pbar, 1):
  144. # Update progress bar description
  145. pbar.set_description(f"Processing {ref_snp['rsID']}")
  146. # Calculate similarity features
  147. features = []
  148. for _, cand_snp in candidate_snps.iterrows():
  149. features.append([
  150. abs(ref_snp['chr_encoded'] - cand_snp['chr_encoded']),
  151. abs(ref_snp['position'] - cand_snp['position']),
  152. abs(ref_snp['log_pvalue'] - cand_snp['log_pvalue']),
  153. abs(ref_snp['gene_count'] - cand_snp['gene_count']),
  154. abs(ref_snp['is_intergenic'] - cand_snp['is_intergenic'])
  155. ])
  156. features_df = pd.DataFrame(features, columns=feature_names)
  157. # Predict similarity scores
  158. proba = model.predict_proba(features_df)
  159. similarity_scores = proba[:, 1] if proba.shape[1] > 1 else np.zeros(len(features_df))
  160. # Get top matches
  161. top_matches = candidate_snps.copy()
  162. top_matches['similarity_score'] = similarity_scores
  163. top_matches = top_matches.nlargest(top_n, 'similarity_score')
  164. for _, match in top_matches.iterrows():
  165. results.append({
  166. 'reference_rsID': ref_snp['rsID'],
  167. 'reference_chr': ref_snp['chromosome'],
  168. 'reference_pos': ref_snp['position'],
  169. 'reference_genes': ref_snp['mappedGenes'],
  170. 'reference_pval': ref_snp['pValue'],
  171. 'predicted_rsID': match['rsID'],
  172. 'predicted_chr': match['chromosome'],
  173. 'predicted_pos': match['position'],
  174. 'predicted_genes': match['mappedGenes'],
  175. 'predicted_pval': match['pValue'],
  176. 'similarity_score': match['similarity_score']
  177. })
  178. # Update progress bar postfix with current SNP info
  179. pbar.set_postfix({
  180. 'Current SNP': ref_snp['rsID'],
  181. 'Top Match': top_matches.iloc[0]['rsID'],
  182. 'Top Score': f"{top_matches.iloc[0]['similarity_score']:.3f}"
  183. })
  184. print(f"\nPrediction completed for all {total_snps} reference SNPs!")
  185. return pd.DataFrame(results)
  186. def main():
  187. print("sALS SNP Similarity Prediction Pipeline")
  188. print("=" * 50)
  189. try:
  190. # 1. Data loading and preparation
  191. print("\n[1/4] Loading and preprocessing data...")
  192. pos_df, neg_df = load_data()
  193. pos_df = prepare_features(pos_df)
  194. neg_df = prepare_features(neg_df)
  195. print(f"- Positive SNPs: {len(pos_df)}")
  196. print(f"- Negative SNPs: {len(neg_df)}")
  197. # 2. Training data preparation
  198. print("\n[2/4] Preparing training data...")
  199. # Create similar pairs (positive-positive)
  200. similar_pairs = []
  201. for i in range(min(500, len(pos_df))): # Limit to 500 positive SNPs for efficiency
  202. for j in range(i + 1, min(i + 5, len(pos_df))): # Compare with next 5 SNPs
  203. similar_pairs.append([
  204. abs(pos_df.iloc[i]['chr_encoded'] - pos_df.iloc[j]['chr_encoded']),
  205. abs(pos_df.iloc[i]['position'] - pos_df.iloc[j]['position']),
  206. abs(pos_df.iloc[i]['log_pvalue'] - pos_df.iloc[j]['log_pvalue']),
  207. abs(pos_df.iloc[i]['gene_count'] - pos_df.iloc[j]['gene_count']),
  208. abs(pos_df.iloc[i]['is_intergenic'] - pos_df.iloc[j]['is_intergenic'])
  209. ])
  210. # Create dissimilar pairs (positive-negative)
  211. dissimilar_pairs = []
  212. for i in range(min(500, len(pos_df))): # Same 500 positive SNPs
  213. for j in range(min(5, len(neg_df))): # Compare with 5 negative SNPs
  214. dissimilar_pairs.append([
  215. abs(pos_df.iloc[i]['chr_encoded'] - neg_df.iloc[j]['chr_encoded']),
  216. abs(pos_df.iloc[i]['position'] - neg_df.iloc[j]['position']),
  217. abs(pos_df.iloc[i]['log_pvalue'] - neg_df.iloc[j]['log_pvalue']),
  218. abs(pos_df.iloc[i]['gene_count'] - neg_df.iloc[j]['gene_count']),
  219. abs(pos_df.iloc[i]['is_intergenic'] - neg_df.iloc[j]['is_intergenic'])
  220. ])
  221. # Combine and split data
  222. X = pd.DataFrame(
  223. similar_pairs + dissimilar_pairs,
  224. columns=['chr_diff', 'pos_diff', 'pval_diff', 'gene_diff', 'intergenic_diff']
  225. )
  226. y = np.array([1] * len(similar_pairs) + [0] * len(dissimilar_pairs))
  227. X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
  228. print(f"\nTraining data summary:")
  229. print(f"- Total samples: {len(X)}")
  230. print(f"- Training samples: {len(X_train)}")
  231. print(f"- Test samples: {len(X_test)}")
  232. print(f"- Positive/Negative ratio: {np.mean(y):.2f}")
  233. # 3. Model training and evaluation
  234. print("\n[3/4] Training and evaluating model...")
  235. model = train_model(X_train, y_train)
  236. evaluate_model(model, X_test, y_test)
  237. # 4. Similar SNP prediction
  238. print("\n[4/4] Predicting similar SNPs...")
  239. print(f"Reference SNPs to process: {len(pos_df)}")
  240. print(f"Candidate SNPs to compare against: {len(neg_df)}")
  241. print(f"Top 10 similar SNPs will be identified for each reference SNP")
  242. predictions = predict_similar_snps(model, pos_df, neg_df, top_n=10)
  243. # Save results
  244. output_file = "sals_similar_snps_predictions.tsv"
  245. predictions.to_csv(output_file, sep='\t', index=False)
  246. print(f"\nSaved {len(predictions)} predictions to {output_file}")
  247. joblib.dump(model, 'sals_similarity_model.pkl')
  248. print("Model saved to 'sals_similarity_model.pkl'")
  249. except Exception as e:
  250. print(f"\nError occurred: {str(e)}")
  251. raise
  252. if __name__ == "__main__":
  253. main()

Project-script.py, under CC-BY-4.0 · at the source

Overview

Authors: Ali Aguerd1, Badreddine Nouadi1, Abdelkarim Ezaouine1, Imad Fenjar1, Faiza Bennis1, Fatima Chegdani1
ORCID iDs: Ali Aguerd
  1. Laboratory of Integrative Biology, Faculty of Science Ain Chock, University Hassan II, Casablanca, Morocco
Journal: Frontiers in genetics, volume 17, article 1742595
Dates: received 9 November 2025; accepted 26 February 2026; published online 12 March 2026
Type: Methods article · Language: English
License: CC BY
Identifiers: DOI 10.3389/fgene.2026.1742595 · PMID 41890230 · PMCID PMC13016588 · OpenAlex W7135010805
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: genetics / omics (modality), none (in silico) (organism), other condition (population), clinical / translational (subfield)
Methods: Machine learning, Statistics
Keywords: genetic biomarkers, genome-wide-associations studies (GWAS), in silico prediction, machine learning, rare diseases, single nucleotide polymorphisms (SNPs), sporadic amyotrophic lateral sclerosis (SALS)
Topic: Amyotrophic Lateral Sclerosis Research (Neurology, Medicine), according to OpenAlex
Citations: not cited yet (Europe PMC); 46 references in the paper

Abstract

Studying the genetics of rare diseases is challenging because small sample sizes limit the statistical power of standard methods like Genome-wide association studies (GWAS). We created a new machine-learning approach to find candidate Single Nucleotide Polymorphisms (SNPs) when data is scarce. Our method trains a Random Forest model to spot similarities between SNPs. We used 189 known Sporadic Amyotrophic Lateral Sclerosis (sALS)-linked SNPs as positive examples and 938,544 unrelated SNPs as negatives. The model learns from genomic location, significance levels, nearby genes, and other features. When we tested it on sALS, it performed exceptionally well, with 93.8% accuracy and near-perfect AUC scores. The method uncovered 1,890 new SNP candidates for sALS. Among these, 209 reached genome-wide significance, and 50 appeared repeatedly in our analyses, making them strong candidates. Key genes like SARM1, OPHN1, and BPTF emerged from the results, all connected to neural health and survival pathways. Our examination revealed a notable excess of SNPs on chromosome 18 compared to expectations. This non-random distribution underscores the region’s particular interest. Here, our approach demonstrates its ability to extract meaningful signals from a restricted sample. The results generated by this approach enable early diagnosis of the disease under study, explanation of its mechanism, and identification of therapeutic targets.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 13 matches between paragraphs and lines of code.

Zenodo 18789012

License: CC-BY-4.0
State: the link answers, verified on 30 September 2026
Evidence: files inventoried
Size: 1 file
Software Heritage: not checked
Found in: “Data availability statement”
Not found: README, license file, CITATION.cff, environment file, tests, continuous integration, documentation
Tools: Matplotlib (3 files), NumPy (3 files), pandas (3 files), scikit-learn (3 files), seaborn (3 files)
Availability: 1 check, the latest on 30 September 2026: the link answers (HTTP 200)
  • 30 September 2026: the link answers (HTTP 200)
3 files

The paper's code and data availability statement is in the Data section.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 3 scripts, each with its path and the digest of its content;
  • 13 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data availability statement

The original contributions presented in the study are publicly available. The SNP dataset and associated Python scripts have been deposited in Zenodo at: https://doi.org/10.5281/zenodo.18789012.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 30 September 2026: the first record

Recorded: type, language, journal, volume, pages, dates, 6 authors, 7 keywords, 43 references.

Cite

This paper

Aguerd, A., Nouadi, B., Ezaouine, A., Fenjar, I., Bennis, F., & Chegdani, F. (2026). An in silico protocol for predicting genetic biomarkers in rare diseases: a case study in sporadic amyotrophic lateral sclerosis. Frontiers in genetics, 17, 1742595. https://doi.org/10.3389/fgene.2026.1742595

BibTeX

@article{aguerd2026silico,
author = {Aguerd, Ali and Nouadi, Badreddine and Ezaouine, Abdelkarim and Fenjar, Imad and Bennis, Faiza and Chegdani, Fatima},
title = {{An in silico protocol for predicting genetic biomarkers in rare diseases: a case study in sporadic amyotrophic lateral sclerosis}},
journal = {Frontiers in genetics},
year = {2026},
month = mar,
volume = {17},
pages = {1742595},
publisher = {Frontiers Media SA},
issn = {1664-8021},
doi = {10.3389/fgene.2026.1742595},
url = {https://doi.org/10.3389/fgene.2026.1742595},
pmid = {41890230},
pmcid = {PMC13016588}
}

RIS

TY - JOUR
AU - Aguerd, Ali
AU - Nouadi, Badreddine
AU - Ezaouine, Abdelkarim
AU - Fenjar, Imad
AU - Bennis, Faiza
AU - Chegdani, Fatima
TI - An in silico protocol for predicting genetic biomarkers in rare diseases: a case study in sporadic amyotrophic lateral sclerosis
T2 - Frontiers in genetics
J2 - Front Genet
PY - 2026
DA - 2026/03/12
VL - 17
SP - 1742595
SN - 1664-8021
PB - Frontiers Media SA
DO - 10.3389/fgene.2026.1742595
UR - https://doi.org/10.3389/fgene.2026.1742595
LA - en
ER -

CSL-JSON

{
"id": "10.3389/fgene.2026.1742595",
"type": "article-journal",
"title": "An in silico protocol for predicting genetic biomarkers in rare diseases: a case study in sporadic amyotrophic lateral sclerosis",
"container-title": "Frontiers in genetics",
"author": [
{
"family": "Aguerd",
"given": "Ali"
},
{
"family": "Nouadi",
"given": "Badreddine"
},
{
"family": "Ezaouine",
"given": "Abdelkarim"
},
{
"family": "Fenjar",
"given": "Imad"
},
{
"family": "Bennis",
"given": "Faiza"
},
{
"family": "Chegdani",
"given": "Fatima"
}
],
"container-title-short": "Front Genet",
"volume": "17",
"page": "1742595",
"DOI": "10.3389/fgene.2026.1742595",
"PMID": "41890230",
"PMCID": "PMC13016588",
"ISSN": "1664-8021",
"publisher": "Frontiers Media SA",
"URL": "https://doi.org/10.3389/fgene.2026.1742595",
"language": "en",
"issued": {
"date-parts": [
[
2026,
3,
12
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1371/journal.pbio.3003856 [code]
Aging and metabolism contribute separately to brain-body health.
Journal: PLoS biology
In common: seaborn, scikit-learn, pandas, 2 other tools, clinical / translational, 3 references
[2] doi:10.1038/s41597-026-07077-7 [code]
Everyday Activity Science and Engineering Table Setting Dataset.
Journal: Scientific data
In common: seaborn, scikit-learn, pandas, 2 other tools, 3 references
[3] doi:10.1016/j.crmeth.2026.101421 [code]
EthoPy provides an accessible platform for reproducible behavioral neuroscience.
Journal: Cell reports methods
In common: seaborn, scikit-learn, pandas, 2 other tools, 3 references
[4] doi:10.1016/j.isci.2026.115766 [code]
RNA-binding protein family diversification correlates with neural complexity across metazoan evolution.
Journal: iScience
In common: scikit-learn, pandas, NumPy, 4 references
[5] doi:10.1002/advs.202521254 [code]
Persistently Increased Expression of PKMzeta and Unbiased Gene Expression Profiles Identify Hippocampal Molecular Traces of a Long-Term Active Place Avoidance Memory and "Shadow" Proteins.
Journal: Advanced science (Weinheim, Baden-Wurttemberg, Germany)
In common: seaborn, scikit-learn, pandas, 2 other tools, genetics / omics, 2 references
[6] doi:10.7554/elife.93664 [code]
Drug-induced changes in connectivity to midbrain dopamine cells revealed by rabies monosynaptic tracing.
Journal: eLife
In common: seaborn, scikit-learn, pandas, 2 other tools, other condition, 2 references
[7] doi:10.7554/elife.110074 [code]
Disentangling cephalopod chromatophores motor units with computer vision.
Journal: eLife
In common: scikit-learn, pandas, Matplotlib, 1 other tool, 3 references
[8] doi:10.1371/journal.pgen.1012242 [code]
Wiz regulates clustered protocadherin genes by restricting CTCF/cohesin loop extrusion in a genomic-distance biased manner.
Journal: PLoS genetics
In common: seaborn, scikit-learn, pandas, 2 other tools, 2 references
[9] doi:10.1093/bioinformatics/btag592 [code]
Network-based stratification of allele-specific expression reveals patient subgroups in Huntington's disease.
Journal: Bioinformatics (Oxford, England)
In common: seaborn, scikit-learn, pandas, 2 other tools, genetics / omics, other condition, 1 reference
[10] doi:10.1016/j.isci.2026.116439 [code]
Decoding the role of transcriptomic clocks in the human prefrontal cortex.
Journal: iScience
In common: seaborn, scikit-learn, pandas, 2 other tools, genetics / omics, 1 reference

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.