OSCR

Accessing the data

Everything OSCR publishes can be downloaded and reused under its licence, without an account on this site. This guide says what is where, and how to read it.

What is open, what is not

WhatWhereAccess
The authors' scripts, under verified licencesHugging Face: OpenScientificCodeRegistry/Databasepublic
Every record the site shows (papers, authors, journals, institutions, tools, datasets, the DOI lookup)this site's JSON files (Labs)public
Tracing maps validated by their authorsZenodo, the registry's communitypublic, CC0-1.0
The source code of the harvester, the site and the WorkerGitHubpublic, Apache-2.0
The catalogue as tables, and a public SQLite databasea Hugging Face datasetprivate until the operator opens it
The authors' contact details from the papersa Hugging Face dataset and the operator's computerprivate, always

The scripts' dataset

OpenScientificCodeRegistry/Database holds the text of the authors' scripts whose repository's licence allows redistribution, confirmed by the repository's own licence file (or, for an archive without one, by its record). Each unique file is stored once, whatever the number of repositories that hold it.

blocks/NNNNN.parquet
One row per unique file: sha256 (of its text), language, size, lines, content. Rows sorted by language then size, 64 to a row group, pages compressed with zstd. A published block never changes: new files go into new blocks.
manifests/<xx>/<repository>.json
One per repository (format oscr-script-manifest/1): its address, host, commit, licence, how the licence was confirmed (license_confirmed_by), and for each file its path, sha256, language, lines, and the block and row that hold its text. <xx> is the first two hex characters of the SHA-1 of the repository's key (github.com/owner/name); the file's name is that key with every other character replaced by __.
README.md, LICENSE.md
The card, with the number of files and repositories and the repositories by licence; the note on licences.

With Python or DuckDB

Hugging Face serves each file over plain HTTPS; its libraries also read hf:// addresses.

# Python: every file of one block, with pandas and huggingface_hub installed
import pandas as pd
df = pd.read_parquet("hf://datasets/OpenScientificCodeRegistry/Database/blocks/00001.parquet")
print(df.groupby("language").size().sort_values(ascending=False).head())
-- DuckDB: the Python scripts of every block, without downloading them whole
SELECT sha256, lines FROM 'hf://datasets/OpenScientificCodeRegistry/Database/blocks/*.parquet'
WHERE language = 'Python' ORDER BY lines DESC LIMIT 10;

To find a repository's files, read its manifest (plain JSON), then the rows it names. The dataset's card also declares the blocks to Hugging Face's viewer and to the datasets library, as the configuration scripts.

In a browser, one script at a time

The blocks are laid out so that one script can be read without the rest: a reader fetches the block's footer once, then the one row group (64 rows) that holds the file — about 78 KB for a script shown — with HTTP range requests, which Hugging Face answers with the CORS headers a page needs. Labs shows how, with the hyparquet library.

Licences of the data

  • The scripts keep the licence of their repository, given in its manifest. The dataset itself adds none. A file under a non-commercial licence (CC BY-NC and its variants) may be reused only non-commercially; attribution goes to the authors, as their licence says.
  • The catalogue's metadata (the records), when released, is under CC0-1.0.
  • Tracing maps deposited on Zenodo are under CC0-1.0.
  • A paper's abstract and statements, shown only under CC BY, CC0, CC BY-SA or CC BY-NC, keep the paper's licence.

The site's own files

What the pages show comes from static JSON files anyone may fetch — the DOI lookup's 256 shards, the records of authors, journals, institutions, tools, datasets and older papers. They change every night, their layout may change with the site, and they are meant for the site's pages first: for bulk use, prefer the datasets. Their addresses and fields are listed on Labs.

Tracing maps

Each validated map is a Zenodo record with one file, tracing-map.json, found by its DOI or in the registry's Zenodo community. Its format is on Labs.

What stays private

  • The harvester's full database: the sentences of each paper that decided a link's verdict, the raw records of Europe PMC and OpenAlex, the versions with who made each correction, the scripts whose licence does not allow copying, and the off-topic papers.
  • The authors' contact details the papers publish (the privacy page).
  • The catalogue's own dataset, until the operator opens it.

Using the data politely

Hugging Face hosts public datasets for free on a best-effort basis: download what you need, once, rather than the same blocks again and again. The site's JSON files are served by Cloudflare's free plan; please do not crawl them faster than a person would read.