OSCR

Importing a repository

How to move a repository from another host or version control system to your GitHub account, keeping its commit ids whenever it is already Git. Every import runs on your own computer: OSCR runs none and sees none of your credentials.

For one Git repository, the import page writes the commands for you. This guide is for the larger moves: several repositories, another version control system, a folder that becomes a repository.

Planning an import

Before moving anything, list what you have: the repositories, their sizes, which use Git LFS, which have submodules, and who pushes to them. GitHub's planning guide and itsmigration paths say which tool fits which source, and at which level: the code and its history only (what this guide does), or also issues, pull requests and wikis (GitHub Enterprise Importer, for organizations).

Trial imports

Import once into a throwaway repository (NAME-trial), check it, delete it, then import for real. A trial costs nothing and shows the size, the time it takes, and any file GitHub refuses:

git clone --mirror 'https://gitlab.com/LAB/NAME.git' 'NAME-trial.git'
cd 'NAME-trial.git'
# Branches and tags only: the source's hidden refs (pull and merge requests) stay behind.
git config --replace-all remote.origin.fetch '+refs/heads/*:refs/heads/*'
git config --add remote.origin.fetch '+refs/tags/*:refs/tags/*'
git for-each-ref --format='delete %(refname)' refs/pull refs/pull-requests refs/merge-requests refs/pipelines refs/keep-around refs/environments | git update-ref --stdin
git push --mirror 'https://github.com/OWNER/NAME-trial.git'

Sizes, before you push

GitHub refuses a file over its size limit and a push over its push limit (the limits).git-sizer lists the largest blobs, trees and histories of a repository:

git clone --mirror https://gitlab.com/LAB/NAME NAME.git
cd NAME.git
git-sizer --verbose
git count-objects -vH

A push over the limit: push in steps, one range of commits at a time, as GitHub'stroubleshooting the 2 GiB push limit shows.

Migration logs, locks and aborting

A command-line import keeps its own log: the terminal's output. Keep it (script import.log before starting records everything). GitHub's importer works on its own servers: it copies one repository over https, does not move Git LFS objects, tells you by a notification when it has finished, and can be cancelled from its page while it runs (about GitHub's importer). GitHub Enterprise Importer can lock the source repositories during a move; a command-line import locks nothing: tell the people who push to stop until the new repository is in place.

Commit attribution

GitHub links a commit to an account by the address written in the commit, if that account has it among its own addresses (why commits are linked to the wrong user). An import keeps the commits exactly as they were, addresses included: OSCR never reads them and never shows them. Commits whose address no account claims show the name only. Organizations that move with GitHub Enterprise Importer get placeholder accounts, called mannequins, for the people it could not match, which theyreclaim afterwards.

From GitLab, Bitbucket, Codeberg or a lab's server

Any Git host is imported the same way: git clone --mirror, then git push --mirror into an empty repository, every branch, tag and commit id kept. The import page writes these commands for your addresses; GitHub's command-line import explains them. The source's merge and pull request refs stay behind: GitHub refuses them on a push.

From another version control system

Converting from another system writes a new Git history: the commits get new ids. Tracing maps made on the old system cannot carry over; maps made after the import can. Each converter runs on your computer. In the commands, authors.txt maps the old system's user names to the names Git writes; you write that file yourself, one line per user:

jdoe = Jane Doe <JANE'S ADDRESS FOR GIT>
(one line per user of the old system)

Subversion

With svn2git, for a repository with the usual trunk, branches and tags:

mkdir 'NAME'
cd 'NAME'
svn2git 'https://svn.example.org/repos/NAME' --authors ../authors.txt
git remote add origin 'https://github.com/OWNER/NAME.git'
git push -u origin --all
git push origin --tags

Or with Git's own git svn:

git svn clone --stdlayout --prefix=svn/ --authors-file=authors.txt 'https://svn.example.org/repos/NAME' 'NAME'
cd 'NAME'
# If the clone stops (a long history, the network), run this until it ends:
git svn fetch
git branch -M main
git remote add origin 'https://github.com/OWNER/NAME.git'
git push -u origin main

Mercurial

With hg-fast-export:

hg clone 'https://hg.example.org/NAME' 'NAME-hg'
git init 'NAME'
cd 'NAME'
git config core.ignoreCase false
hg-fast-export.sh -r '../NAME-hg' -A ../authors.txt -M main
git checkout main
git remote add origin 'https://github.com/OWNER/NAME.git'
git push -u origin --all
git push origin --tags

TFVC

With git-tfs (Team Foundation Version Control):

git tfs clone 'https://tfs.example.org/tfs/DefaultCollection' '$/Project/Main' 'NAME' --branches=all
cd 'NAME'
git remote add origin 'https://github.com/OWNER/NAME.git'
git push -u origin --all
git push origin --tags

Perforce

With Git's own git p4:

# Sign in to Perforce first (p4 login): git p4 uses that session.
P4PORT='ssl:perforce.example.org:1666' git p4 clone '//depot/project@all' 'NAME'
cd 'NAME'
git branch -M main
git remote add origin 'https://github.com/OWNER/NAME.git'
git push -u origin main

Large files before the first push

A converted history is new, so rewriting it costs nothing: move the large files into Git LFS before the first push (git lfs migrate), or GitHub refuses the push. After the conversion, in its folder:

svn2git 'https://svn.example.org/repos/NAME'
git lfs migrate import --everything --above=50MB
git remote add origin 'https://github.com/OWNER/NAME.git'
git push -u origin --all
git push origin --tags

Large files says where data belongs before it goes into LFS.

A folder of a repository into a new repository

With git filter-repo, the folder becomes the new repository's root, with the history of its files only (splitting a subfolder). These commits get new ids.

git clone 'https://github.com/OWNER/BIG-PROJECT.git' 'NAME'
cd 'NAME'
git filter-repo --subdirectory-filter 'analysis/eeg'
git remote add origin 'https://github.com/OWNER/NAME.git'
git push -u origin --all
git push origin --tags

Another repository into a folder of yours: a subtree merge

Inside the receiving repository, the other one's history is merged under a folder, and its new commits can be brought in later (about Git subtree merges):

git remote add 'TOOLBOX' 'https://gitlab.com/LAB/TOOLBOX.git'
git fetch 'TOOLBOX'
git merge -s ours --no-commit --allow-unrelated-histories 'TOOLBOX/main'
git read-tree --prefix='vendor/toolbox/' -u 'TOOLBOX/main'
git commit -m 'Subtree merge of gitlab.com/LAB/TOOLBOX into vendor/toolbox/'
# Later, to bring in the source's new commits:
git pull -s subtree 'TOOLBOX' 'main'

A copy without a fork, and an ongoing mirror

A fork stays attached to its source on GitHub; a duplicate is a repository of its own, every commit id kept (duplicating a repository): it is the import above, from a GitHub address. When the source goes on living elsewhere, you can keep your copy up to date yourself, from the folder of the first import, by hand or from your own scheduler. OSCR runs no mirror for you.

cd 'NAME.git'
git remote set-url --push origin 'https://github.com/OWNER/NAME.git'
git fetch --prune origin
git push --mirror

Several repositories at once

List the repositories, one a line: the source's git address, then, if you like, the new name (or account/name) and the word lfs when the source uses Git LFS. The page writes a shell script (POSIX sh) that imports them one after the other, each in its own block, stopping at the first error. It is made in your browser and sent nowhere; it holds no credential: git asks GitHub for yours. Create each repository first, empty. At most 100 repositories in one script.

After the import

An import brings the history, not the repository's settings: set its description, website andtopics on GitHub (or with GitHub's command-line tool, gh), and check that it has alicence: without one, OSCR keeps no copy of its scripts. Thenlink it to its papers.

gh repo edit OWNER/NAME --description "What the code does" --homepage "https://doi.org/10.XXXX/YYYY"
gh repo edit OWNER/NAME --add-topic neuroscience --add-topic eeg

When it goes wrong

remote: error: File … is … MB; this exceeds GitHub's file size limit
A file over the limit is in the history: move it into Git LFS before pushing (above), or leave it out.
fatal: the remote end hung up unexpectedly, or a push that stops
Often a push over the push limit: push in steps (the 2 GiB push limit).
! [remote rejected] refs/pull/… (deny updating a hidden ref)
The source's pull request refs: the mirror commands leave them behind; delete them as the commands do.
! [remote rejected] … (push declined due to repository rule violations)
A ruleset of the destination refuses the push. Import into a repository without rulesets, then add them; bypass lists are set by the repository's admins (about rulesets).
Authentication failed, or a password prompt that refuses your password
GitHub takes a token as the password over https: tokens for git.

Organizations and GitHub Enterprise

Moving an organization's repositories with their issues and pull requests isGitHub Enterprise Importer's work, and live migrations between Enterprise servers are GitHub's own service: OSCR does not re-implement them. Once the repositories are on GitHub, link them here like any other.

More