The public API
OSCR offers a free, keyless, read-only API over its catalogue. No account, no token: every route answers a plain GET (and HEAD) and returns JSON, with CORS open to any origin. It returns only facts: no PDF, no paper full text, and no email address. The overview is at the API page.
Quickstart
Every route hangs off /api/v1. Fetch the index, which links to every endpoint:
# curl curl https://openscicode.org/api/v1/
# Python (standard library only)
import json, urllib.request
with urllib.request.urlopen("https://openscicode.org/api/v1/") as r:
index = json.load(r)
print([e["id"] for e in index["endpoints"]])# R (base R, no package)
con <- url("https://openscicode.org/api/v1/")
index <- jsonlite::fromJSON(readLines(con, warn = FALSE))
close(con)
str(index$endpoints$id)The examples below use curl; the Python and R calls are the same everywhere, only the URL changes. In Python, requests.get(url).json() works the same as the urllib above; in R, jsonlite::fromJSON(url) reads a URL directly.
Authentication (none)
There is no authentication. Do not send a key, a token or a cookie: the API ignores them. Because it is keyless and CORS is open, you can call it from a browser page too. (A separate token-based layer, with write routes, is in development and is not part of this API: see Labs.)
The shape of an answer
Every answer is a JSON object with the API version, a canonical self link, and the payload:
{
"oscr_api": "v1",
"self": "https://openscicode.org/api/v1/stats",
"generated_at": "2026-09-30T01:40:53Z",
"figures": { "articles": 1234, "with_code": 567, ... },
"source": "/data/stats.json"
}Where a clean static file holds the same data, the answer names it (source or bulk): fetch that file directly to avoid the Worker entirely (it has no rate limit).
The endpoints
GET /api/v1/ and GET /api/v1/openapi.json
The index links to every endpoint. The OpenAPI 3.1 document describes them for tools (a static copy is at /data/openapi.json):
curl https://openscicode.org/api/v1/openapi.json
GET /api/v1/stats
The catalogue's figures (papers, papers with code, repositories, scripts, matches).
curl https://openscicode.org/api/v1/stats
GET /api/v1/paper/{doi}
One paper's record by its DOI: its code links and repositories, data links, status, tools, categories, and a matches summary. No PDF and no full text (the paper is linked by its DOI). See Giving a DOI for the slashes.
curl https://openscicode.org/api/v1/paper/10.5555%2Foscr.fixture.1 curl "https://openscicode.org/api/v1/paper?doi=10.5555/oscr.fixture.1"
# Python: a paper, by its DOI
import json, urllib.parse, urllib.request
doi = "10.5555/oscr.fixture.1"
url = "https://openscicode.org/api/v1/paper/" + urllib.parse.quote(doi, safe="")
with urllib.request.urlopen(url) as r:
paper = json.load(r)["paper"]
print(paper["title"], [c["url"] for c in paper["code"]])GET /api/v1/search
Full-text search, with the same query language and filters as the site's search. Parameters: q, page, size (1 to 50), and the site's filters. This is the only endpoint bound by a quota (see Rate and caching).
curl "https://openscicode.org/api/v1/search?q=eeg&size=20"
# R: the first page of a search
res <- jsonlite::fromJSON("https://openscicode.org/api/v1/search?q=eeg&size=20")
res$total
res$results$doiGET /api/v1/{type} and GET /api/v1/{type}/{id}
The lists (/authors, /journals, /institutions, /tools, /datasets) are paginated (see Pagination) and point to the bulk file that holds them all. One entity is fetched by its key (singular type): an ORCID iD for an author, a ROR id for an institution, a slug otherwise.
curl "https://openscicode.org/api/v1/authors?size=50&page=1" curl https://openscicode.org/api/v1/author/0000-0000-0000-0028 curl https://openscicode.org/api/v1/tool/matplotlib
GET /api/v1/repository/{host}/{owner}/{name}
A repository the registry knows: its licence, state, pinned commit, file and script counts, languages, and the papers that cite it. The code text is never here (read it in the Code to Paper reader on a paper's page, or in the scripts dataset).
curl https://openscicode.org/api/v1/repository/github.com/oscr-fixture/eeg-analysis
Giving a DOI
A DOI contains a slash (10.5555/abcd), so in the path it must be URL-encoded (the slash becomes %2F). To avoid encoding, pass it unencoded as the ?doi= query parameter instead. Both forms return the same record. A DOI is matched case-insensitively, with or without a https://doi.org/ prefix.
Pagination
A list takes page (from 1) and size (1 to 100, default 50). The answer carries total, page, size, pages and next (the URL of the next page, or null). For the whole list at once, fetch the bulk file the answer names: it is a single static download with no rate limit.
Rate and caching
The static files (/data/) and the Hugging Face dataset are unlimited: they are served by the CDN. The one exception is /api/v1/search, which reads the search database, so each call counts against the site's shared daily free budget. Call it sparingly and cache its answers. Every API answer carries a Cache-Control header (ten minutes), so a repeat of the same call is served from the cache. When the search's daily quota is spent, it answers 503 with the error quota.
Bulk downloads
To extract everything, read the static files; the API points to them so you can skip the Worker:
/data/articles.csv,repositories.csv- One row per paper, and per repository.
/data/alignments.jsonl- One JSON object per aligned paper (paragraph numbers to code lines, never the paper's text).
/data/entities/<type>.json,/data/papers/NN.json,/data/repos/NN.json- The full entity lists, and every paper's and repository's record (sharded).
# Python: download the whole paper table, then every author
import csv, io, json, urllib.request
with urllib.request.urlopen("https://openscicode.org/data/articles.csv") as r:
rows = list(csv.DictReader(io.TextIOWrapper(r, encoding="utf-8")))
print(len(rows), "papers")
with urllib.request.urlopen("https://openscicode.org/data/entities/authors.json") as r:
authors = json.load(r)
print(len(authors), "authors")# R: the same two files
papers <- read.csv("https://openscicode.org/data/articles.csv")
authors <- jsonlite::fromJSON("https://openscicode.org/data/entities/authors.json")
nrow(papers); length(authors$orcid)The authors' scripts (their text) are a separate public dataset on Hugging Face, OpenScientificCodeRegistry/Database, read with pandas or DuckDB (see the data guide). A single full catalogue dump in one file is not served here: it is too large for the hosting.
Errors
On a failure the API returns the right status code and a JSON body of a fixed shape:
{ "error": "not_found", "message": "The registry has no record of this DOI.",
"documentation_url": "https://openscicode.org/help/api/" }| Status | error | When |
|---|---|---|
| 400 | bad_doi, bad_key, bad_repository, bad_query | A parameter is malformed. |
| 404 | not_found | No such record or endpoint. |
| 405 | method_not_allowed | A method other than GET or HEAD. |
| 503 | quota, unavailable, not_configured | The search's database cannot answer (the quota is spent, or it is down). |
Versioning and stability
The API is versioned in its path (/api/v1). Within a version, fields are only added, never removed or renamed, so a reader that ignores unknown fields keeps working. A breaking change would come as a new version (/api/v2/), announced before the old one is retired. The record layout may gain fields as the catalogue grows.
Licence and citation
The catalogue's facts are open data; the authors' code stays under its own licence (that is why some files are shown from their source rather than copied: the code policy). Use of the API and the data is covered by the terms of use. To cite OSCR and a record, see Citation.
