Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus.
The 5 matches
- [1] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Architecture of Deflexformer ↔ deflex/deflexformer/blocks.py, lines 12–109 · score 0.86 · Residual connections, causal attention, layer normalization, multi head, Deflexformer block, transforms
- [2] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Pre-training the Deflexformer block ↔ deflex/core/estimator.py, lines 14–52 · score 0.81 · single Deflexformer block, stage training, post training, training Deflexformer, pre train, deep
- [3] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Architecture of Deflexformer ↔ deflex/deflexformer/models.py, lines 13–102 · score 0.71 · Residual connections, layer normalization, Deflexformer block, head, modules, dimension
- [4] § Results › Evaluation metrics ↔ deflex/core/estimator.py, lines 14–52 · score 0.65 · post training, pre training, Deflexformer block, formula discovery, symbolic regression, fitting
- [5] § Methods › Deflexformer: energy-based NNs with decomposable blocks › Post-training the full Deflexformer ↔ deflex/deflexformer/estimator.py, lines 218–251 · score 0.54 · pre trained block, post trained, Deflexformer, model
Paper
Loaded from Europe PMC by your browser, not stored by OSCR: doi.org · Europe PMC
The paper is loaded when this pane is shown.
The authors' code
Python · 231 lines · 8.3 KB · MIT · 2 matches
- """
- Main Deflex estimator implementing the three-stage training process.
- """
- import numpy as np
- from sklearn.base import BaseEstimator, RegressorMixin
- from sklearn.utils.validation import check_X_y, check_array
- from typing import Optional, Dict, Any, Union
- from ..deflexformer import DeflexformerEstimator
- from ..lambda_regression import LambdaRegressionEstimator
- class DeflexEstimator(BaseEstimator, RegressorMixin):
- """
- Deep Formula Discovery estimator combining neural networks and symbolic regression.
- This estimator implements a three-stage training process:
- 1. Pre-training: Train a single Deflexformer block on synthetic data
- 2. Post-training: Stack pre-trained blocks and fine-tune on experimental data
- 3. Symbolic regression: Extract symbolic formulas from the trained model
- Parameters
- ----------
- n_blocks : int, default=5
- Number of Deflexformer blocks to use in the model.
- embedding_dim : int, default=64
- Maximum embedding dimension for the Deflexformer.
- n_synthetic_samples : int, default=10000
- Number of synthetic samples to generate for pre-training.
- pretrain_epochs : int, default=100
- Number of epochs for pre-training stage.
- finetune_epochs : int, default=50
- Number of epochs for post-training/fine-tuning stage.
- operator_set : list, optional
- Set of operators for symbolic regression. If None, uses default set.
- lambda_regression_params : dict, optional
- Additional parameters for the LambdaRegression estimator.
- random_state : int, optional
- Random seed for reproducibility.
- Attributes
- ----------
- deflexformer_ : DeflexformerEstimator
- The trained Deflexformer model.
- lambda_regression_ : LambdaRegressionEstimator
- The symbolic regression model.
- is_fitted_ : bool
- Whether the estimator has been fitted.
- formula_ : str
- The discovered symbolic formula (available after fitting).
- """
- def __init__(
- self,
- n_blocks: int = 5,
- embedding_dim: int = 64,
- n_synthetic_samples: int = 10000,
- pretrain_epochs: int = 100,
- finetune_epochs: int = 50,
- operator_set: Optional[list] = None,
- lambda_regression_params: Optional[Dict[str, Any]] = None,
- random_state: Optional[int] = None
- ):
- self.n_blocks = n_blocks
- self.embedding_dim = embedding_dim
- self.n_synthetic_samples = n_synthetic_samples
- self.pretrain_epochs = pretrain_epochs
- self.finetune_epochs = finetune_epochs
- self.operator_set = operator_set
- self.lambda_regression_params = lambda_regression_params or {}
- self.random_state = random_state
- def fit(self, X: np.ndarray, y: np.ndarray) -> 'DeflexEstimator':
- """
- Fit the Deflex model using the three-stage training process.
- Parameters
- ----------
- X : array-like of shape (n_frames, n_elements, embedding_dim)
- Training data as 3D tensor.
- y : array-like of shape (n_frames, n_elements, output_dim)
- Target values.
- Returns
- -------
- self : DeflexEstimator
- Fitted estimator.
- """
- # Validate input
- X, y = check_X_y(X, y, multi_output=True, allow_nd=True)
- if X.ndim != 3:
- raise ValueError(f"X must be 3D tensor, got shape {X.shape}")
- if y.ndim < 2:
- raise ValueError(f"y must be at least 2D, got shape {y.shape}")
- # Stage 1: Pre-training on synthetic data
- self._pretrain_stage()
- # Stage 2: Post-training on experimental data
- self._posttrain_stage(X, y)
- # Stage 3: Symbolic regression
- self._symbolic_regression_stage(X, y)
- self.is_fitted_ = True
- return self
- def _pretrain_stage(self) -> None:
- """Stage 1: Pre-train single block on synthetic data."""
- print("Stage 1: Pre-training on synthetic data...")
- # Initialize LambdaRegression for synthetic data generation
- self.lambda_regression_ = LambdaRegressionEstimator(
- operator_set=self.operator_set,
- random_state=self.random_state,
- **self.lambda_regression_params
- )
- # Generate synthetic data
- X_synthetic, y_synthetic = self.lambda_regression_.generate_synthetic_data(
- n_samples=self.n_synthetic_samples
- )
- # Initialize single-block Deflexformer for pre-training
- self.deflexformer_ = DeflexformerEstimator(
- n_blocks=1, # Only one block for pre-training
- embedding_dim=self.embedding_dim,
- enable_clipping=True, # Strict clipping for synthetic data
- random_state=self.random_state
- )
- # Pre-train the single block
- self.deflexformer_.fit(
- X_synthetic, y_synthetic,
- epochs=self.pretrain_epochs,
- stage="pretrain"
- )
- def _posttrain_stage(self, X: np.ndarray, y: np.ndarray) -> None:
- """Stage 2: Stack blocks and fine-tune on experimental data."""
- print("Stage 2: Post-training on experimental data...")
- # Expand to full architecture with pre-trained weights
- self.deflexformer_.expand_to_full_model(self.n_blocks)
- # Fine-tune on experimental data
- self.deflexformer_.fit(
- X, y,
- epochs=self.finetune_epochs,
- stage="posttrain"
- )
- def _symbolic_regression_stage(self, X: np.ndarray, y: np.ndarray) -> None:
- """Stage 3: Extract symbolic formulas."""
- print("Stage 3: Symbolic regression...")
- # Extract features from trained Deflexformer blocks
- block_outputs = self.deflexformer_.extract_block_outputs(X)
- # Generate additional synthetic data for symbolic regression
- X_additional, y_additional = self.lambda_regression_.generate_synthetic_data(
- n_samples=self.n_synthetic_samples // 2
- )
- # Combine experimental and synthetic data for symbolic regression
- X_combined = np.concatenate([block_outputs, X_additional], axis=0)
- y_combined = np.concatenate([y, y_additional], axis=0)
- # Perform symbolic regression
- self.lambda_regression_.fit(X_combined, y_combined)
- self.formula_ = self.lambda_regression_.get_formula()
- def predict(self, X: np.ndarray) -> np.ndarray:
- """
- Predict using the fitted model.
- Parameters
- ----------
- X : array-like of shape (n_frames, n_elements, embedding_dim)
- Input data.
- Returns
- -------
- y_pred : ndarray of shape (n_frames, n_elements, output_dim)
- Predicted values.
- """
- if not hasattr(self, 'is_fitted_'):
- raise ValueError("This DeflexEstimator instance is not fitted yet.")
- X = check_array(X, allow_nd=True)
- # Use symbolic formula if available, otherwise use Deflexformer
- if hasattr(self, 'formula_') and self.formula_:
- return self.lambda_regression_.predict(X)
- else:
- return self.deflexformer_.predict(X)
- def get_formula(self) -> str:
- """
- Get the discovered symbolic formula.
- Returns
- -------
- formula : str
- The symbolic formula in human-readable format.
- """
- if not hasattr(self, 'formula_'):
- raise ValueError("Model must be fitted before getting formula.")
- return self.formula_
- def score(self, X: np.ndarray, y: np.ndarray) -> float:
- """
- Return the coefficient of determination R^2 of the prediction.
- Parameters
- ----------
- X : array-like of shape (n_frames, n_elements, embedding_dim)
- Test samples.
- y : array-like of shape (n_frames, n_elements, output_dim)
- True values.
- Returns
- -------
- score : float
- R^2 coefficient of determination.
- """
- from sklearn.metrics import r2_score
- y_pred = self.predict(X)
- return r2_score(y, y_pred, multioutput='uniform_average')
estimator.py at commit df664b1, under MIT · at the source
Overview
- National Engineering Laboratory for Big Data Analytics, Xi’an Jiaotong University, Xi’an, Shaanxi China
- School of Mathematics and Statistics, Xi’an Jiaotong University, Xi’an, Shaanxi China
- School of Computer Science and Technology, Faculty of Electronic and Information Engineering, Xi’an Jiaotong University, Xi’an, Shaanxi China
Abstract
A fundamental problem in science is identifying underlying patterns of complex systems in the form of concise mathematical formulas. Current Artificial Intelligence (AI)-based methods have shown strong performance in single-scale systems, yet remain limited in identifying scale-specific formulas in multiscale complex systems. We present Deflex, an end-to-end AI method to automatically extract multiscale formulas with potentially different forms, including invariants and distributions, from complex systems. Deflex consists of two subsystems named Deflexformer and Deflexpressor. Deflexpressor is a lambda-calculus symbolic regression model for higher-order formulas. Deflexformer is a decomposable deep energy model for learning unified representations across scales. Deflexpressor generates synthetic data to pre-train Deflexformer, which then guides formula discovery by decoupling multiscale latent relationships. Across six representative complex systems with diverse behaviors, Deflex achieves up to 7-fold higher efficiency than the state-of-the-art methods while enabling automated multiscale discovery. Our work could be a useful tool for scientific discovery across disciplines.
Reproduced under the paper's license (CC BY), from the paper cited above.
Repository
Its files are read in the Code ↔ Paper reader above, with 5 matches between paragraphs and lines of code.
yhqjohn/deflex
df664b1d422f802543087f415eb9a5304b2c79ef, 4 August 2025Availability: 1 check, the latest on 27 September 2026: the link answers
- 27 September 2026: the link answers
45 files
- LambdaRegression/
docs/ , Julia, 29 linesmake.jl - LambdaRegression/
src/ , Julia, 23 linesLambdaRegression.jl - LambdaRegression/
src/ , Julia, 104 linesUtils.jl - LambdaRegression/
src/ , Julia, 22 linesasts/ ASTs.jl - LambdaRegression/
src/ , Julia, 106 linesasts/ contexts.jl - LambdaRegression/
src/ , Julia, 221 linesasts/ cse.jl - LambdaRegression/
src/ , Julia, 1,267 linesasts/ nodeasts.jl - LambdaRegression/
src/ , Julia, 366 linesasts/ parse.jl - LambdaRegression/
src/ , Julia, 314 linesgenerate/ GenerateTree.jl - LambdaRegression/
src/ , Julia, 92 linesgenerate/ constantctors.jl - LambdaRegression/
src/ , Julia, 399 linesreduce/ Reduce.jl - LambdaRegression/
src/ , Julia, 1,461 linestrees/ Trees.jl - LambdaRegression/
src/ , Julia, 621 linestrees/ dfstrees.jl - LambdaRegression/
src/ , Julia, 36 linestrees/ utils.jl - LambdaRegression/
src/ , Julia, 1,267 linestypesystem/ TypeSystem.jl - LambdaRegression/
test/ , Julia, 4 linesruntests.jl - conftest.py, Python, 64 lines
- deflex/
__init__.py , Python, 21 lines - deflex/
core/ , Python, 7 lines__init__.py - deflex/
core/ , Python, 231 lines, 2 matchesestimator.py - deflex/
deflexformer/ , Python, 16 lines__init__.py - deflex/
deflexformer/ , Python, 109 lines, 1 matchblocks.py - deflex/
deflexformer/ , Python, 285 lines, 1 matchestimator.py - deflex/
deflexformer/ , Python, 233 lineslayers.py - deflex/
deflexformer/ , Python, 127 lines, 1 matchmodels.py - deflex/
julia_utils.py , Python, 102 lines - deflex/
lambda_regression/ , Python, 17 lines__init__.py - deflex/
lambda_regression/ , Python, 256 linesestimator.py - deflex/
lambda_regression/ , Python, 298 linesjulia_interface.py - deflex/
lambda_regression/ , Python, 185 linesoperators.py - deflex/
utils/ , Python, 13 lines__init__.py - deflex/
utils/ , Python, 205 linesdata.py - deflex/
utils/ , Python, 212 linesvalidation.py - examples/
__init__.py , Python, 3 lines - examples/
basic_usage.py , Python, 264 lines - examples/
json_data_example.py , Python, 267 lines - setup.py, Python, 59 lines
- tests/
__init__.py , Python, 3 lines - tests/
test_core.py , Python, 148 lines - tests/
test_data_loading.py , Python, 237 lines - tests/
test_deflexformer.py , Python, 184 lines - tests/
test_lambda_regression.p , Python, 206 linesy - tests/
test_utils.py , Python, 231 lines - LICENSE, License, 21 lines
- README.md, Text, 105 lines
Code availability
The code used in the study is available in the open-source development repository at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Tracing map
Proposed by the machine: these links were found in the paper and verified at the source, without human review. The map will receive a Zenodo DOI once one of the paper's authors has validated it with their ORCID.
What the map holds:
- 1 repository of the authors' code, each at its verified commit, with its license and how the link was found in the paper;
- 43 scripts, each with its path and the digest of its content;
- 5 matches between paragraphs of the paper and lines of the code (method lexical-v1);
- neither the text of the paper nor the code itself.
Its JSON (tracing-map.json) is deposited on Zenodo with its DOI once the map is validated.
Data
No dataset and no data link were found in the paper.
Data Availability Statement
The data generated or processed in this study are available in the open-source project repository at https://
The code used in the study is available in the open-source development repository at https://
Reproduced under the paper's license (CC BY), from the paper cited above.
Versions
The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.
Version 1, 27 September 2026: the first record
Recorded: type, language, journal, volume, issue, pages, dates, 4 authors, 2 keywords, 1 funder, 21 references.
Cite
This paper
Yu, H., Yang, S., Ren, X., & Zhao, C. (2026). Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus. Nature communications, 17(1), 7639. https://
BibTeX
@article{yu2026discoveri
author = {Yu, Hanqiao and Yang, Shusen and Ren, Xuebin and Zhao, Cong},
title = {{Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus}},
journal = {Nature communications},
year = {2026},
month = jun,
volume = {17},
number = {1},
pages = {7639},
publisher = {Nature Publishing Group},
issn = {2041-1723},
doi = {10.1038/
url = {https://
pmid = {42303622},
pmcid = {PMC13434754}
}
RIS
TY - JOUR
AU - Yu, Hanqiao
AU - Yang, Shusen
AU - Ren, Xuebin
AU - Zhao, Cong
TI - Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus
T2 - Nature communications
J2 - Nat Commun
PY - 2026
DA - 2026/
VL - 17
IS - 1
SP - 7639
SN - 2041-1723
PB - Nature Publishing Group
DO - 10.1038/
UR - https://
LA - en
ER -
CSL-JSON
{
"id": "10.1038/
"type": "article-journal",
"title": "Discovering multiscale deep formulas in complex systems via neural-guided lambda calculus",
"container-title": "Nature communications",
"author": [
{
"family": "Yu",
"given": "Hanqiao"
},
{
"family": "Yang",
"given": "Shusen"
},
{
"family": "Ren",
"given": "Xuebin"
},
{
"family": "Zhao",
"given": "Cong"
}
],
"container-title-short":
"volume": "17",
"issue": "1",
"page": "7639",
"DOI": "10.1038/
"PMID": "42303622",
"PMCID": "PMC13434754",
"ISSN": "2041-1723",
"publisher": "Nature Publishing Group",
"URL": "https://
"language": "en",
"issued": {
"date-parts": [
[
2026,
6,
16
]
]
}
}
The tracing map gets a citation of its own once an author has validated it and it has a DOI.
Similar papers
The papers with a page that share the most with this one: the tools found in their code, their categories, datasets, cited references and authors, the rarest counting most.
- [1] doi:10.1016/j.molcel.2026.07.006 [code]
- DeorphaNN: Virtual screening of GPCR peptide agonists using AlphaFold-predicted active-state complexes and deep learning embeddings.Journal: Molecular cellIn common: PyTorch, scikit-learn, pandas, 3 other tools, 1 reference
- [2] doi:10.1038/s41586-026-10658-6 [code]
- An AI system to help scientists write expert-level empirical software.Journal: NatureIn common: PyTorch, scikit-learn, pandas, 3 other tools, 1 reference
- [3] doi:10.1038/s41467-026-74002-2 [code]
- Learning missing physics from legacy simulators with alternating neural integrators.Journal: Nature communicationsIn common: PyTorch, scikit-learn, pandas, 3 other tools, 1 reference
- [4] doi:10.3390/biomimetics11070516 [code]
- Evolutionary, Neural, or LLM-Driven Heuristic Generation? A Unified Ant Colony Optimization Benchmark for Nature-Inspired Routing Heuristics on the TSP and CVRP.Journal: Biomimetics (Basel, Switzerland)In common: PyTorch, scikit-learn, SciPy, 2 other tools, 1 reference
- [5] doi:10.1098/rstb.2024.0461 [code]
- Shallow recurrent decoders for neural and behavioural dynamics.Journal: Philosophical transactions of the Royal Society of London. Series B, Biological sciencesIn common: PyTorch, scikit-learn, pandas, 3 other tools, computational
- [6] doi:10.1002/hbm.70600 [code]
- A Data-Driven Closed-Loop Control Approach to Drive Neural State Transitions for Mechanistic Insight.Journal: Human brain mappingIn common: PyTorch, scikit-learn, pandas, 3 other tools, computational
- [7] doi:10.1038/s41467-026-74243-1 [code]
- Spiking neural network decoders of finger forces from high-density intramuscular microelectrode arrays.Journal: Nature communicationsIn common: PyTorch, scikit-learn, pandas, 3 other tools, computational
- [8] doi:10.1371/journal.pcbi.1014364 [code]
- A comparative study of simulation-based inference methods for epidemic models with identifiability considerations.Journal: PLoS computational biologyIn common: PyTorch, scikit-learn, pandas, 3 other tools, computational
- [9] doi:10.1371/journal.pcbi.1014337 [code]
- Fast reconstruction of degenerate populations of conductance-based neuron models from spike times.Journal: PLoS computational biologyIn common: PyTorch, scikit-learn, pandas, 3 other tools, computational
- [10] doi:10.1162/imag.a.1147 [code]
- The Virtual Brain links transcranial magnetic stimulation evoked potentials and inhibitory neurotransmitter changes in major depressive disorder.Journal: Imaging neuroscience (Cambridge, Mass.)In common: PyTorch, scikit-learn, pandas, 3 other tools, computational
Contribute
The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.
Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.
Claim this paper
Correct its record
Say what each link of this record is, remove the ones that are not the paper's, add the ones that are missing. The correction becomes a new version of the record, in its Versions section.
Validate its tracing map
You validate the map as this page shows it: 1 repository of the authors' code, each at its verified commit and with its license, 43 scripts, and 5 matches between paragraphs and code (see the Code and Map sections). It then receives a DOI on Zenodo, with you (your ORCID iD) and OSCR as its creators; the code itself is not deposited.
The map's fingerprint: sha256:1c25fbafbdd37cf3…
Add the badge to its README
The badge links the code to this page. Copy one of these into the README of the paper's code: only you decide where it goes, and nothing is changed for you.
Markdown
[, paste the snippet at the top, then “Commit changes…” and, to review it first, “Create a new branch and start a pull request”. You open the pull request; OSCR asks for no permission.
Request its removal
To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).
Discussion, reproductions, activity
Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.
Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.
Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.
