OSCR

Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus.

Code ↔ Paper

5 matches between paragraphs of the paper and lines of its authors' code, computed by the harvester (lexical-v1). Click a colored paragraph or line to see its counterpart.

The 5 matches
  1. [1] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Architecture of Deflexformer ↔ deflex/deflexformer/blocks.py, lines 12–109 · score 0.86 · Residual connections, causal attention, layer normalization, multi head, Deflexformer block, transforms
  2. [2] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Pre-training the Deflexformer block ↔ deflex/core/estimator.py, lines 14–52 · score 0.81 · single Deflexformer block, stage training, post training, training Deflexformer, pre train, deep
  3. [3] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Architecture of Deflexformer ↔ deflex/deflexformer/models.py, lines 13–102 · score 0.71 · Residual connections, layer normalization, Deflexformer block, head, modules, dimension
  4. [4] § Results › Evaluation metrics ↔ deflex/core/estimator.py, lines 14–52 · score 0.65 · post training, pre training, Deflexformer block, formula discovery, symbolic regression, fitting
  5. [5] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Post-training the full Deflexformer ↔ deflex/deflexformer/estimator.py, lines 218–251 · score 0.54 · pre trained block, post trained, Deflexformer, model

Paper

Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC

The paper is loaded when this pane is shown.

The authors' code

Python · 231 lines · 8.3 KB · MIT · 2 matches

  1. """
  2. Main Deflex estimator implementing the three-stage training process.
  3. """
  4. import numpy as np
  5. from sklearn.base import BaseEstimator, RegressorMixin
  6. from sklearn.utils.validation import check_X_y, check_array
  7. from typing import Optional, Dict, Any, Union
  8. from ..deflexformer import DeflexformerEstimator
  9. from ..lambda_regression import LambdaRegressionEstimator
  10. class DeflexEstimator(BaseEstimator, RegressorMixin):
  11. """
  12. Deep Formula Discovery estimator combining neural networks and symbolic regression.
  13. This estimator implements a three-stage training process:
  14. 1. Pre-training: Train a single Deflexformer block on synthetic data
  15. 2. Post-training: Stack pre-trained blocks and fine-tune on experimental data
  16. 3. Symbolic regression: Extract symbolic formulas from the trained model
  17. Parameters
  18. ----------
  19. n_blocks : int, default=5
  20. Number of Deflexformer blocks to use in the model.
  21. embedding_dim : int, default=64
  22. Maximum embedding dimension for the Deflexformer.
  23. n_synthetic_samples : int, default=10000
  24. Number of synthetic samples to generate for pre-training.
  25. pretrain_epochs : int, default=100
  26. Number of epochs for pre-training stage.
  27. finetune_epochs : int, default=50
  28. Number of epochs for post-training/fine-tuning stage.
  29. operator_set : list, optional
  30. Set of operators for symbolic regression. If None, uses default set.
  31. lambda_regression_params : dict, optional
  32. Additional parameters for the LambdaRegression estimator.
  33. random_state : int, optional
  34. Random seed for reproducibility.
  35. Attributes
  36. ----------
  37. deflexformer_ : DeflexformerEstimator
  38. The trained Deflexformer model.
  39. lambda_regression_ : LambdaRegressionEstimator
  40. The symbolic regression model.
  41. is_fitted_ : bool
  42. Whether the estimator has been fitted.
  43. formula_ : str
  44. The discovered symbolic formula (available after fitting).
  45. """
  46. def __init__(
  47. self,
  48. n_blocks: int = 5,
  49. embedding_dim: int = 64,
  50. n_synthetic_samples: int = 10000,
  51. pretrain_epochs: int = 100,
  52. finetune_epochs: int = 50,
  53. operator_set: Optional[list] = None,
  54. lambda_regression_params: Optional[Dict[str, Any]] = None,
  55. random_state: Optional[int] = None
  56. ):
  57. self.n_blocks = n_blocks
  58. self.embedding_dim = embedding_dim
  59. self.n_synthetic_samples = n_synthetic_samples
  60. self.pretrain_epochs = pretrain_epochs
  61. self.finetune_epochs = finetune_epochs
  62. self.operator_set = operator_set
  63. self.lambda_regression_params = lambda_regression_params or {}
  64. self.random_state = random_state
  65. def fit(self, X: np.ndarray, y: np.ndarray) -> 'DeflexEstimator':
  66. """
  67. Fit the Deflex model using the three-stage training process.
  68. Parameters
  69. ----------
  70. X : array-like of shape (n_frames, n_elements, embedding_dim)
  71. Training data as 3D tensor.
  72. y : array-like of shape (n_frames, n_elements, output_dim)
  73. Target values.
  74. Returns
  75. -------
  76. self : DeflexEstimator
  77. Fitted estimator.
  78. """
  79. # Validate input
  80. X, y = check_X_y(X, y, multi_output=True, allow_nd=True)
  81. if X.ndim != 3:
  82. raise ValueError(f"X must be 3D tensor, got shape {X.shape}")
  83. if y.ndim < 2:
  84. raise ValueError(f"y must be at least 2D, got shape {y.shape}")
  85. # Stage 1: Pre-training on synthetic data
  86. self._pretrain_stage()
  87. # Stage 2: Post-training on experimental data
  88. self._posttrain_stage(X, y)
  89. # Stage 3: Symbolic regression
  90. self._symbolic_regression_stage(X, y)
  91. self.is_fitted_ = True
  92. return self
  93. def _pretrain_stage(self) -> None:
  94. """Stage 1: Pre-train single block on synthetic data."""
  95. print("Stage 1: Pre-training on synthetic data...")
  96. # Initialize LambdaRegression for synthetic data generation
  97. self.lambda_regression_ = LambdaRegressionEstimator(
  98. operator_set=self.operator_set,
  99. random_state=self.random_state,
  100. **self.lambda_regression_params
  101. )
  102. # Generate synthetic data
  103. X_synthetic, y_synthetic = self.lambda_regression_.generate_synthetic_data(
  104. n_samples=self.n_synthetic_samples
  105. )
  106. # Initialize single-block Deflexformer for pre-training
  107. self.deflexformer_ = DeflexformerEstimator(
  108. n_blocks=1, # Only one block for pre-training
  109. embedding_dim=self.embedding_dim,
  110. enable_clipping=True, # Strict clipping for synthetic data
  111. random_state=self.random_state
  112. )
  113. # Pre-train the single block
  114. self.deflexformer_.fit(
  115. X_synthetic, y_synthetic,
  116. epochs=self.pretrain_epochs,
  117. stage="pretrain"
  118. )
  119. def _posttrain_stage(self, X: np.ndarray, y: np.ndarray) -> None:
  120. """Stage 2: Stack blocks and fine-tune on experimental data."""
  121. print("Stage 2: Post-training on experimental data...")
  122. # Expand to full architecture with pre-trained weights
  123. self.deflexformer_.expand_to_full_model(self.n_blocks)
  124. # Fine-tune on experimental data
  125. self.deflexformer_.fit(
  126. X, y,
  127. epochs=self.finetune_epochs,
  128. stage="posttrain"
  129. )
  130. def _symbolic_regression_stage(self, X: np.ndarray, y: np.ndarray) -> None:
  131. """Stage 3: Extract symbolic formulas."""
  132. print("Stage 3: Symbolic regression...")
  133. # Extract features from trained Deflexformer blocks
  134. block_outputs = self.deflexformer_.extract_block_outputs(X)
  135. # Generate additional synthetic data for symbolic regression
  136. X_additional, y_additional = self.lambda_regression_.generate_synthetic_data(
  137. n_samples=self.n_synthetic_samples // 2
  138. )
  139. # Combine experimental and synthetic data for symbolic regression
  140. X_combined = np.concatenate([block_outputs, X_additional], axis=0)
  141. y_combined = np.concatenate([y, y_additional], axis=0)
  142. # Perform symbolic regression
  143. self.lambda_regression_.fit(X_combined, y_combined)
  144. self.formula_ = self.lambda_regression_.get_formula()
  145. def predict(self, X: np.ndarray) -> np.ndarray:
  146. """
  147. Predict using the fitted model.
  148. Parameters
  149. ----------
  150. X : array-like of shape (n_frames, n_elements, embedding_dim)
  151. Input data.
  152. Returns
  153. -------
  154. y_pred : ndarray of shape (n_frames, n_elements, output_dim)
  155. Predicted values.
  156. """
  157. if not hasattr(self, 'is_fitted_'):
  158. raise ValueError("This DeflexEstimator instance is not fitted yet.")
  159. X = check_array(X, allow_nd=True)
  160. # Use symbolic formula if available, otherwise use Deflexformer
  161. if hasattr(self, 'formula_') and self.formula_:
  162. return self.lambda_regression_.predict(X)
  163. else:
  164. return self.deflexformer_.predict(X)
  165. def get_formula(self) -> str:
  166. """
  167. Get the discovered symbolic formula.
  168. Returns
  169. -------
  170. formula : str
  171. The symbolic formula in human-readable format.
  172. """
  173. if not hasattr(self, 'formula_'):
  174. raise ValueError("Model must be fitted before getting formula.")
  175. return self.formula_
  176. def score(self, X: np.ndarray, y: np.ndarray) -> float:
  177. """
  178. Return the coefficient of determination R^2 of the prediction.
  179. Parameters
  180. ----------
  181. X : array-like of shape (n_frames, n_elements, embedding_dim)
  182. Test samples.
  183. y : array-like of shape (n_frames, n_elements, output_dim)
  184. True values.
  185. Returns
  186. -------
  187. score : float
  188. R^2 coefficient of determination.
  189. """
  190. from sklearn.metrics import r2_score
  191. y_pred = self.predict(X)
  192. return r2_score(y, y_pred, multioutput='uniform_average')

estimator.py at commit df664b1, under MIT · at the source

Overview

Authors: Hanqiao Yu1,2, Shusen Yang1,2, Xuebin Ren1,3, Cong Zhao1,2
  1. National Engineering Laboratory for Big Data Analytics, Xi’an Jiaotong University, Xi’an, Shaanxi China
  2. School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an, Shaanxi China
  3. School of Computer Science and Technology, Faculty of Electronic and Information Engineering, Xi’an Jiaotong University, Xi’an, Shaanxi China
Institutions: Xi'an Jiaotong University (China)
Journal: Nature communications, volume 17, issue 1, article 7639
Dates: received 7 April 2025; accepted 3 June 2026; published online 16 June 2026
Type: Research article · Language: English
License: CC BY
Identifiers: DOI 10.1038/s41467-026-74299-z · PMID 42303622 · PMCID PMC13434754 · OpenAlex W7164903124
Open access: gold, a free copy (OpenAlex)
Status: code verified
Categories: computational (subfield)
Methods: Machine learning
Keywords: Computational science, Information theory and computation
Topic: Machine Learning in Materials Science (Materials Chemistry, Materials Science), according to OpenAlex
Funding: National Natural Science Foundation of China (National Science Foundation of China) (62172329, U21A6005)
Citations: not cited yet (Europe PMC); 54 references in the paper

Abstract

A fundamental problem in science is identifying underlying patterns of complex systems in the form of concise mathematical formulas. Current Artificial Intelligence (AI)-based methods have shown strong performance in single-scale systems, yet remain limited in identifying scale-specific formulas in multiscale complex systems. We present Deflex, an end-to-end AI method to automatically extract multiscale formulas with potentially different forms, including invariants and distributions, from complex systems. Deflex consists of two subsystems named Deflexformer and Deflexpressor. Deflexpressor is a lambda-calculus symbolic regression model for higher-order formulas. Deflexformer is a decomposable deep energy model for learning unified representations across scales. Deflexpressor generates synthetic data to pre-train Deflexformer, which then guides formula discovery by decoupling multiscale latent relationships. Across six representative complex systems with diverse behaviors, Deflex achieves up to 7-fold higher efficiency than the state-of-the-art methods while enabling automated multiscale discovery. Our work could be a useful tool for scientific discovery across disciplines.

Reproduced under the paper's license (CC BY), from the paper cited above.

Repository

Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.

yhqjohn/deflex

License: MIT
State: the link answers, verified on 27 September 2026
Evidence: files inventoried
Commit: df664b1d422f802543087f415eb9a5304b2c79ef, 4 August 2025
Languages: Python (27), Julia (16)
Size: 64 files, 43 scripts
Software Heritage: not archived
Found in: “Code availability”
Holds: README, license file, environment (pyproject.toml, requirements-dev.txt, requirements.txt, setup.py, LambdaRegression/Project.toml, LambdaRegression/docs/Project.toml, LambdaRegression/test/Project.toml), tests, documentation
Not found: CITATION.cff, continuous integration
Tools: NumPy (15 files), scikit-learn (7 files), PyTorch (5 files), pandas (2 files), Matplotlib (1 file), SciPy (1 file)
Availability: 1 check, the latest on 27 September 2026: the link answers
  • 27 September 2026: the link answers
45 files

Code availability

The code used in the study is available in the open-source development repository at https://github.com/yhqjohn/deflex.

Reproduced under the paper's license (CC BY), from the paper cited above.

Tracing map

Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.

What the map holds:

  • 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
  • 43 scripts, each with its path and the digest of its content;
  • 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
  • neither the text of the paper nor the code itself.

Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

The data generated or processed in this study are available in the open-source project repository at https://github.com/yhqjohn/deflex. For the simulation-based projects, the repository provides the generated datasets for rare gas particle motion, pollen-in-water particle motion, cross-scale water particles, and two-dimensional cylinder-wake fluid dynamics. The original third-party datasets used in this study are publicly available from their source repositories or publications: the Zurich Carnival human-mobility dataset at 10.4108/icst.urb-iot.2014.257190; the eBird Status and Trends Data Version 2022 at 10.2173/ebirdst.2022; the JHTDB dataset at https://turbulence.pha.jhu.edu/; and the Feynman Symbolic Regression Database at https://space.mit.edu/home/tegmark/aifeynman.html, with associated publication 10.1126/sciadv.aay2631. No additional access restrictions apply to the data made available through the project repository.

The code used in the study is available in the open-source development repository at https://github.com/yhqjohn/deflex.

Reproduced under the paper's license (CC BY), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 2 keywords, 1 funder, 21 references.

Cite

This paper

Yu, H., Yang, S., Ren, X., & Zhao, C. (2026). Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus. Nature communications, 17(1), 7639. https://doi.org/10.1038/s41467-026-74299-z

BibTeX

@article{yu2026discovering,
author = {Yu, Hanqiao and Yang, Shusen and Ren, Xuebin and Zhao, Cong},
title = {{Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus}},
journal = {Nature communications},
year = {2026},
month = jun,
volume = {17},
number = {1},
pages = {7639},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/s41467-026-74299-z},
url = {https://doi.org/10.1038/s41467-026-74299-z},
pmid = {42303622},
pmcid = {PMC13434754}
}

RIS

TY - JOUR
AU - Yu, Hanqiao
AU - Yang, Shusen
AU - Ren, Xuebin
AU - Zhao, Cong
TI - Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/06/16
VL - 17
IS - 1
SP - 7639
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/s41467-026-74299-z
UR - https://doi.org/10.1038/s41467-026-74299-z
LA - en
ER -

CSL-JSON

{
"id": "10.1038/s41467-026-74299-z",
"type": "article-journal",
"title": "Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus",
"container-title": "Nature communications",
"author": [
{
"family": "Yu",
"given": "Hanqiao"
},
{
"family": "Yang",
"given": "Shusen"
},
{
"family": "Ren",
"given": "Xuebin"
},
{
"family": "Zhao",
"given": "Cong"
}
],
"container-title-short": "Nat Commun",
"volume": "17",
"issue": "1",
"page": "7639",
"DOI": "10.1038/s41467-026-74299-z",
"PMID": "42303622",
"PMCID": "PMC13434754",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://doi.org/10.1038/s41467-026-74299-z",
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
16
]
]
}
}

The tracing map gets a citation of its own once an author has validated it and it has a DOI.

Similar papers

The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.

[1] doi:10.1016/j.molcel.2026.07.006 [code]
DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.
Journal: Molecular cell
In common: PyTorch, scikit-learn, pandas, 3 other tools, 1 reference
[2] doi:10.1038/s41586-026-10658-6 [code]
An AI system to help scientists write expert-level empirical software.
Journal: Nature
In common: PyTorch, scikit-learn, pandas, 3 other tools, 1 reference
[3] doi:10.1038/s41467-026-74002-2 [code]
Learning missing physics from legacy simulators with alternating neural integrators.
Journal: Nature communications
In common: PyTorch, scikit-learn, pandas, 3 other tools, 1 reference
[4] doi:10.3390/biomimetics11070516 [code]
Evolutionary, Neural, or LLM-Driven Heuristic Generation? A Unified Ant Colony Optimization Benchmark for Nature-Inspired Routing Heuristics on the TSP and CVRP.
Journal: Biomimetics (Basel, Switzerland)
In common: PyTorch, scikit-learn, SciPy, 2 other tools, 1 reference
[5] doi:10.1098/rstb.2024.0461 [code]
Shallow recurrent decoders for neural and behavioural dynamics.
Journal: Philosophical transactions of the Royal Society of London. Series B, Biological sciences
In common: PyTorch, scikit-learn, pandas, 3 other tools, computational
[6] doi:10.1002/hbm.70600 [code]
A Data-Driven Closed-Loop Control Approach to Drive Neural State Transitions for Mechanistic Insight.
Journal: Human brain mapping
In common: PyTorch, scikit-learn, pandas, 3 other tools, computational
[7] doi:10.1038/s41467-026-74243-1 [code]
Spiking neural network decoders of finger forces from high-density intramuscular microelectrode arrays.
Journal: Nature communications
In common: PyTorch, scikit-learn, pandas, 3 other tools, computational
[8] doi:10.1371/journal.pcbi.1014364 [code]
A comparative study of simulation-based inference methods for epidemic models with identifiability considerations.
Journal: PLoS computational biology
In common: PyTorch, scikit-learn, pandas, 3 other tools, computational
[9] doi:10.1371/journal.pcbi.1014337 [code]
Fast reconstruction of degenerate populations of conductance-based neuron models from spike times.
Journal: PLoS computational biology
In common: PyTorch, scikit-learn, pandas, 3 other tools, computational
[10] doi:10.1162/imag.a.1147 [code]
The Virtual Brain links transcranial magnetic stimulation evoked potentials and inhibitory neurotransmitter changes in major depressive disorder.
Journal: Imaging neuroscience (Cambridge, Mass.)
In common: PyTorch, scikit-learn, pandas, 3 other tools, computational

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.