Creating repositories¶
The course uses an example project throughout: a mental health in tech survey, with a funding document, a report, and a data directory. This page shows how to turn such a project into a Git repository.
What is a Git repository?¶
A Git repository (or repo) is a directory that contains:
- your files and subdirectories, the actual project content;
- Git storage, a hidden
.gitdirectory where Git keeps the full history and its own information about the project.
mental-health-workspace/ <- the repository (project directory)
├── funding.doc <- file
├── report.md <- file
├── data/ <- sub-directory
└── .git/ <- Git storage (hidden)
The files and subdirectories are what you work on. The .git directory is what makes the directory a repository: it holds every recorded version, the staging area, and the repository configuration.
Do not edit .git
Never edit, move, or delete files inside .git by hand. Git manages this directory itself, and changing it manually can corrupt the repository. Deleting .git removes the whole history: the directory becomes an ordinary folder again.
Benefits of repositories¶
- Systematically track versions of every file in the project.
- Revert to previous versions when something goes wrong.
- Compare versions at different points in time.
- Collaborate with colleagues, because everyone works from the same history.
Creating a new repository¶
To create a brand-new project, pass a name to git init. Git creates a directory with that name and initializes a repository inside it:
$ git init mental-health-workspace
Initialized empty Git repository in /home/repl/mental-health-workspace/.git/
$ cd mental-health-workspace
$ git status
On branch main
No commits yet
nothing to commit (create/copy files and use "git add" to track)
| Command | Effect |
|---|---|
git init mental-health-workspace |
Creates the mental-health-workspace directory and the .git storage inside it |
cd mental-health-workspace |
Moves into the new repository. Git commands act on the repository you are in |
git status |
Reports the state of the repository: current branch, commits, and file changes |
Here git status reports that the repository is on branch main, that nothing has been recorded yet (No commits yet), and that there are no files to record.
Converting an existing project into a repository¶
If the project directory already exists, move into it and run git init without a name. The current directory becomes the repository:
$ cd /home/repl/mental-health-workspace
$ git init
Initialized empty Git repository in /home/repl/mental-health-workspace/.git/
Your existing files are not modified: Git only adds the .git directory. According to the git init documentation, running git init again in an existing repository is safe and does not overwrite what is already there.
| Syntax | Use it when |
|---|---|
git init <directory> |
Starting a new project from scratch |
git init |
The project directory already exists and you are inside it |
What is being tracked?¶
Right after git init in a directory that already contains files, git status lists them as untracked:
$ git status
On branch main
No commits yet
Untracked files:
(use "git add <file>..." to include in what will be committed)
data/
report.md
nothing added to commit but untracked files present (use "git add" to track)
Git has noticed files in the directory that it is not tracking yet. An untracked file exists in the directory, but Git is not recording its versions. Git never starts tracking a file automatically: you have to add it explicitly, which is the subject of the next page. Git also gives a hint in parentheses: use "git add" to track.
Read git status often
git status changes nothing; it only reports. Run it whenever you are unsure what state the repository is in, and read the hints it prints: they usually tell you which command to use next.
The initial branch name: main or master¶
The course outputs show On branch main. This depends on configuration. The git init documentation states that, when no name is given, Git falls back to the default name, which is currently master. This will change to main when Git 3.0 is released. The default can be changed in two ways:
# Choose the name for one repository
git init -b main mental-health-workspace
# Or change the default for all new repositories (user-wide setting)
git config --global init.defaultBranch main
Both names work the same way; only the name differs. GitHub names the default branch of new repositories main. These notes use git init -b main in the challenges so that the outputs match the course.
Nested repositories¶
Do not create a Git repository inside another Git repository. This is called a nested repository.
project/
├── .git/ <- repository 1
└── analysis/
├── .git/ <- repository 2, nested inside repository 1
└── notebook.py
With two .git directories, it becomes unclear which repository should record changes to analysis/notebook.py. In practice, the outer repository does not track the inner repository's files as ordinary files. This is confusing and easy to get wrong.
Common mistake: running git init in your home directory or in a parent folder "to be safe", then again in each project. Before running git init, run git status. If it prints fatal: not a git repository, you are not inside a repository and can safely create one.
Note
Git does provide an official mechanism for including one repository inside another, called submodules. It is an advanced topic, listed at the end of the course as a next step, and is not covered here.
Challenge: create two repositories¶
Objective: create one repository from scratch and convert one existing project into a repository, then check what Git is tracking.
Prerequisites and initial state: Git installed. Everything happens in ~/git-practice/repos, which must not be inside an existing repository.
Setup:
mkdir -p ~/git-practice/repos
cd ~/git-practice/repos
git status # must print "fatal: not a git repository ..."
mkdir -p mental-health-project/data
echo "# Mental Health in Tech Survey" > mental-health-project/report.md
echo "age,gender,treatment" > mental-health-project/data/mental_health_survey.csv
Tasks:
- From
~/git-practice/repos, create a new, empty repository calledmental-health-workspacewhose first branch is namedmain, and check its status. - Convert the existing
mental-health-projectdirectory into a repository, withmainas its first branch. - Check which files Git sees in
mental-health-project, and whether they are tracked. - Confirm that each repository has its own
.gitdirectory and that the two repositories are not nested.
Expected result and verification:
git statusinmental-health-workspaceshowsNo commits yetandnothing to commit.git statusinmental-health-projectlistsdata/andreport.mdunderUntracked files.ls -ashows a.gitdirectory in each project, and~/git-practice/repositself is not a repository (git statusthere still fails).
Solution
cd ~/git-practice/repos
# 1. New repository from scratch
git init -b main mental-health-workspace
cd mental-health-workspace
git status # On branch main / No commits yet / nothing to commit
cd ..
# 2. Convert an existing directory
cd mental-health-project
git init -b main # no name: the current directory becomes the repository
# 3. What is tracked?
git status # data/ and report.md are listed as untracked
cd ..
# 4. Checks
ls -a mental-health-workspace mental-health-project # both contain .git
git status # still "fatal: not a git repository" in ~/git-practice/repos
git init <name>creates the directory, whilegit initalone uses the current directory.-b mainsets the name of the first branch. Without it, Git usesmasterunlessinit.defaultBranchis configured.- The files in
mental-health-projectare untracked: Git sees them but will not record them until they are added withgit add. - The repositories are siblings, not nested: neither one is inside the other, and their parent directory is not a repository.
Clean up when you are done: rm -rf ~/git-practice/repos.