Aller au contenu

Staging and committing files

The Git workflow

Recording a change in Git is a three-step workflow:

  1. Edit and save files on your computer, as usual, in the working directory (the project files you see and edit).
  2. Add the files to the staging area. The staging area holds the changes you have selected for the next record. It tracks what has been modified and is ready to be recorded.
  3. Commit the files. Git takes a snapshot of the staged files at that point in time and stores it in the repository. This snapshot is called a commit. Commits are what let you later compare files and revert them.
 working directory  ──git add──▶  staging area  ──git commit──▶  repository (.git)
   (edit & save)                 (next snapshot)                 (recorded history)

Staging versus committing

The course uses a postal analogy:

Step Analogy Meaning
Staging (git add) Putting a letter in an envelope You prepare what will be sent. You can still add or remove items.
Committing (git commit) Dropping the envelope in a postbox The content is sent and recorded. That snapshot is now part of the history.

The staging area lets you choose exactly what goes into each commit. For example, if you modified three files but only two belong to the same logical change, stage and commit those two, then commit the third one separately.

Adding files to the staging area

Command Effect
git add README.md Stages a single file
git add . Stages all new and modified files in the current directory and its subdirectories
git add README.md   # stage one file
git add .           # stage everything under the current directory

The . is a path: it means "the current directory". If you run git add . from a subdirectory such as data/, only the changes under data/ are staged.

git add also turns an untracked file into a tracked one. After the first git add and commit, Git tracks the file and reports any later change to it.

$ git add report.md
$ git status
On branch main

No commits yet

Changes to be committed:
  (use "git rm --cached <file>..." to unstage)
    new file:   report.md

Untracked files:
  (use "git add <file>..." to include in what will be committed)
    data/

report.md is now listed under Changes to be committed (staged), while data/ is still untracked.

Common mistake: git add . stages everything

git add . is convenient but also stages files you may not want to record, such as temporary files or large outputs. Run git status before committing to check what is staged.

Making a commit

$ git commit -m "Adding a README."
[main cb33c18] Adding a README.
 1 file changed, 1 insertion(+)
 create mode 100644 README.md
Part of the command Role
git commit Records the staged changes as a new commit
-m "Adding a README." Provides the log message directly on the command line, without opening a text editor

Reading the output:

  • [main cb33c18]: the branch (main) and the beginning of the new commit's identifier (its hash, explained in Viewing the version history).
  • 1 file changed, 1 insertion(+): a summary of the recorded changes.
  • create mode 100644 README.md: README.md is a new file in the repository. 100644 is the file mode Git records for a regular, non-executable file.

For the very first commit of a repository, Git writes (root-commit) after the branch name, for example [main (root-commit) 366c3f7].

The log message

  • The log message is useful for reference: it explains why the change was made, for anyone reading the history later, including you.
  • Best practice: keep it short and concise, and describe the change. For example "Add 2024 survey responses" instead of "update".
  • If you omit -m, Git opens a text editor so that you can write the message. Saving an empty message aborts the commit.

Only staged changes are committed

git commit records the content of the staging area, not the working directory. A file modified but not added with git add is not part of the commit. Run git status afterwards: it still lists those changes as Changes not staged for commit.

Configure your identity once

Every commit records an author name and email. If they are not configured, git commit fails with Author identity unknown and asks you to set them:

git config --global user.name "Your Name"
git config --global user.email "[email protected]"

In the challenges, the identity is set without --global, so only the practice repository is affected.

Challenge: your first commits

Objective: use the add/commit workflow to record a project in two separate commits, choosing what goes into each one.

Prerequisites and initial state: Git installed. The setup creates a fresh repository in ~/git-practice/staging with two untracked files.

Setup:

mkdir -p ~/git-practice/staging && cd ~/git-practice/staging
git init -b main
git config user.name "Practice User"
git config user.email "[email protected]"
printf '# Mental Health in Tech Survey\nTODO: write executive summary.\n' > report.md
mkdir data && echo "age,gender,treatment" > data/mental_health_survey.csv

Tasks:

  1. Display the status: which files are untracked?
  2. Stage only report.md, check the status, then commit it with the message Add report.
  3. Stage everything that remains with a single command, and commit it with the message Add survey data.
  4. Check that nothing is left to commit.

Expected result and verification:

  • After step 2, git status shows data/ still untracked, and the first commit output contains (root-commit) and 1 file changed.
  • After step 4, git status prints nothing to commit, working tree clean.
  • git log --oneline (detailed on the next pages) lists exactly two commits: Add survey data above Add report.
Solution
git status                       # data/ and report.md are untracked

git add report.md                # stage a single file
git status                       # report.md under "Changes to be committed", data/ still untracked
git commit -m "Add report"       # [main (root-commit) xxxxxxx] Add report

git add .                        # stage everything under the current directory
git commit -m "Add survey data"  # 1 file changed ... data/mental_health_survey.csv

git status                       # nothing to commit, working tree clean
git log --oneline                # two commits, newest first
  • Staging report.md alone keeps data/ out of the first commit: the staging area defines the content of each commit.
  • git add . stages all remaining changes from the current directory down, which here means data/mental_health_survey.csv.
  • The commit hashes on your machine differ from any example: they depend on content, author, and date.

Clean up when you are done: rm -rf ~/git-practice/staging.