OSCR

DTLI: Distribution transformation-based lightweight learned indexing for data lakehouse query optimization.

Overview

Authors: Ye Liang1,2, Chao Xu1,2
ORCID iDs: Ye Liang
  1. School of Information Science and Technology, Beijing Foreign Studies University, Beijing, China
  2. Artificial Intelligence and Human Languages Lab, Beijing Foreign Studies University, Beijing, China
Journal: Science progress, volume 109, issue 3, article 00368504261472740
Dates: received 16 February 2026; accepted 13 July 2026; published online 1 August 2026; in print July 2026
Type: Research article · Language: English
License: CC BY-NC
Identifiers: DOI 10.1177/00368504261472740 · PMID 42541366 · PMCID PMC13428968 · OpenAlex W7172163248
Open access: gold, a free copy (OpenAlex)
Status: code on request
Categories: methods / tools (subfield)
Methods: Machine learning, Spectral & time-frequency
Keywords: learned index, distribution transformation, data lakehouse, query optimization
Topic: Advanced Database Systems and Queries (Computer Networks and Communications, Computer Science), according to OpenAlex
Funding: National Social Science Fund of China (25&ZD043)
Citations: not cited yet (Europe PMC); 29 references in the paper

Abstract

With the exponential growth of data scale, modern analytical platforms such as data lakehouses face severe query performance bottlenecks. To address the limitation of traditional indexes and existing learned indexes in balancing high-efficiency queries with lightweight structures, this paper proposes DTLI, a lightweight learned index architecture based on distribution transformation. We first systematically evaluate the distribution transformation performance of three generative models, Variational Autoencoder (VAE), Normalizing Flow, and Diffusion Model, and select the optimal Block Neural Autoregressive Flow (B-NAF) as the transformation operator to map original complex distributions to near-uniform distributions. On this basis, we propose a changepoint-based piecewise fitting algorithm for cumulative distribution functions, constructing a minimal binary tree index structure containing only model nodes and pointer nodes to achieve index lightweighting. Finally, we design local and global indexing strategies adapted to the partitioning characteristics of data lakehouses, completing the integration of DTLI on the Apache Hudi platform. Experiments demonstrate that DTLI significantly outperforms native indexes and learned indexes including RMI, PGM, and NFL across multiple datasets. Specifically, DTLI improves average throughput by 72.55%, 50.87%, 38.60%, and 5.21% over B+Tree, RMI, PGM, and NFL, respectively, while reducing 99th percentile tail latency by 57.52%, 56.44%, 29.35%, and 15.81%, with advantages amplifying as data scale increases. Ablation studies confirm that distribution transformation can improve the throughput of existing indexes by over 38%. Currently, DTLI is designed and evaluated for one-dimensional numerical keys. Its extension to multi-dimensional or non-numerical data is non-trivial and remains as future work. Additionally, the distribution transformation relies on offline training, and incremental update mechanisms for dynamic scenarios require further investigation.

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Code

The paper says that its authors' code is available on request: it was not published with the paper, so there is nothing to verify.

The paper's code and data availability statement is in the Data section.

Tracing map

A tracing map links a paper to the code its authors published: this paper has none (its code is available on request), so it has no map.

Data

No dataset and no data link were found in the paper.

Data Availability Statement

The data supporting the findings of this study are openly available. The source code for the DTLI index and associated experimental scripts developed in this study have been deposited in a GitHub repository and will be made available on reasonable request.*

Reproduced under the paper's license (CC BY-NC), from the paper cited above.

Versions

The history of this record: each version stored by the harvester or made by a correction of its authors or of the maintainers of its code, and what changed in its facts. The texts of the paper (its abstract, its availability statements) are not part of it; versions that changed only those are not listed.

Version 1, 27 September 2026: the first record

Recorded: type, language, journal, volume, issue, pages, dates, 2 authors, 4 keywords, 1 funder, 12 references.

Cite

This paper

Liang, Y., & Xu, C. (2026). DTLI: Distribution transformation-based lightweight learned indexing for data lakehouse query optimization. Science progress, 109(3), 00368504261472740. https://doi.org/10.1177/00368504261472740

BibTeX

@article{liang2026dtli,
author = {Liang, Ye and Xu, Chao},
title = {{DTLI: Distribution transformation-based lightweight learned indexing for data lakehouse query optimization}},
journal = {Science progress},
year = {2026},
month = jul,
volume = {109},
number = {3},
pages = {00368504261472740},
publisher = {SAGE Publications},
issn = {0036-8504},
doi = {10.1177/00368504261472740},
url = {https://doi.org/10.1177/00368504261472740},
pmid = {42541366},
pmcid = {PMC13428968}
}

RIS

TY - JOUR
AU - Liang, Ye
AU - Xu, Chao
TI - DTLI: Distribution transformation-based lightweight learned indexing for data lakehouse query optimization
T2 - Science progress
J2 - Sci Prog
PY - 2026
DA - 2026/07/01
VL - 109
IS - 3
SP - 00368504261472740
SN - 0036-8504
PB - SAGE Publications
DO - 10.1177/00368504261472740
UR - https://doi.org/10.1177/00368504261472740
LA - en
ER -

CSL-JSON

{
"id": "10.1177/00368504261472740",
"type": "article-journal",
"title": "DTLI: Distribution transformation-based lightweight learned indexing for data lakehouse query optimization",
"container-title": "Science progress",
"author": [
{
"family": "Liang",
"given": "Ye"
},
{
"family": "Xu",
"given": "Chao"
}
],
"container-title-short": "Sci Prog",
"volume": "109",
"issue": "3",
"page": "00368504261472740",
"DOI": "10.1177/00368504261472740",
"PMID": "42541366",
"PMCID": "PMC13428968",
"ISSN": "0036-8504",
"publisher": "SAGE Publications",
"URL": "https://doi.org/10.1177/00368504261472740",
"language": "en",
"issued": {
"date-parts": [
[
2026,
7,
1
]
]
}
}

Similar papers

No other paper with a page shares enough with this one yet: tools, categories, datasets, references or authors.

Contribute

The authors of this paper can claim it, correct its record and validate its tracing map, and the maintainers of its code (its owner, or a public member of its organization) correct what it says of their repository; anyone signed in can ask for its removal. Every request goes to OSCR's own machine, which answers it; your account page follows them.

Sign in with ORCID to claim this paper as one of its authors, correct its record or validate its tracing map: when the paper's metadata lists your ORCID iD, you are recognized at once. Maintainers of its code: sign in with GitHub, then claim the repository on your account page.

Request its removal

To ask OSCR to remove this record, the copies of its authors' scripts or its tracing map, use the removal request page: signed in, you say who you are, what to remove and why, then review and confirm the request. Published rules decide every request (how).

Discussion, reproductions, activity

Discussion: questions and error reports about this paper and its code, from signed-in readers and its authors. It opens with sign-in.

Reproductions: reports from readers who ran the authors' code: what they reproduced, with which environment, commit and data. It opens with sign-in.

Activity: what happens around this paper: new versions of its record, its map's validation, discussions and reproductions. It opens with sign-in.