Lately I have been working a lot with HPC clusters and one thing I keep noticing from my teammates and others around is the doubts that arise when you are starting out, writing code seems like a part of the challenge. Where should you edit the code? How do you share changes with teammates? Where should training run, and what happens if your SSH connection drops?
This post walks through a practical workflow for your first deep learning project on a high-performance computing (HPC) cluster. We will use a video-classification project as an example, but the same ideas apply to other deep-learning projects.
You should already be able to open a terminal and connect to your cluster with SSH. The examples assume Git, uv, and tmux are available, and that the cluster uses Slurm to schedule work. Your university or lab’s documentation should explain how to access these tools and which storage locations to use.
Here is how the pieces fit together:
| Tool or location | What it does in this workflow |
|---|---|
| Git | Tracks code changes and lets teammates exchange them. |
| Cluster storage | Holds datasets, checkpoints, and experiment outputs. |
uv |
Creates a Python environment from the project’s dependency files. |
| Slurm | Allocates compute resources and runs training jobs. |
tmux |
Keeps your terminal workspace available when SSH disconnects. |
I use a bare Git repository on the cluster as the shared remote in this example. GitHub or GitLab can serve that role too, if your project and cluster allow it. This is one possible setup that you can adapt to your team.
The situation
Okay so the project we would be dealing with at hand is a video-classification problem. Here the raw videos and extracted frames are large, while the useful source files are relatively small. The paths below use placeholders. Replace /cluster/projects/video-classification and /cluster/git/video-classification.git with appropriate paths that you might encounter.
Keeping everything in a single Git repository would be inconvenient and risky:
- Git is not designed to version large datasets for every code change.
- model checkpoints, metrics, and logs can grow quickly.
- a Python virtual environment contains generated files that should be recreated rather than committed.
- absolute paths from one machine will not work on another.
At the same time, developing by manually copying files between machines makes it easy to lose track of which code and configuration produced a given experiment. The solution is to separate concerns clearly:
Git history
source + configuration
|
v
Laptop/editor <--> cluster working checkout <--> cluster Git origin
|
+--> dataset storage
+--> checkpoints and logs
+--> scheduled training jobs
The working checkout is where code is edited. The cluster-local bare repository acts as the Git synchronization point for the source tree. Large datasets and experiment outputs live next to the checkout, but outside version control.
Repository layout
A useful project layout in general:
video-classification/
├── src/ # importable project code
├── scripts/ # dataset checks and experiment entry points
├── jobs/ # scheduler submission scripts
├── configs/ # explicit experiment settings
├── tests/ # lightweight tests
├── pyproject.toml # project metadata and dependencies
├── uv.lock # locked dependency resolution
├── .gitignore # files Git must not track
├── README.md # project orientation
└── notebooks/ # optional local exploration notebooks
The dataset, virtual environment, checkpoints, and generated artifacts are kept outside Git on purpose:
/cluster/datasets/ucf101/
/cluster/projects/video-classification/.venv/
/cluster/projects/video-classification/artifacts/
/cluster/projects/video-classification/checkpoints/
The exact directories may differ. The key idea is that source code should refer to the dataset through configuration, not through a machine-specific absolute path.
The Git arrangement on the cluster
In this setup, there are two different Git repositories to keep in mind.
1. The working repository
This is the directory where files are checked out and edited:
/cluster/projects/video-classification/
It has a working tree, a current branch, and a .git directory. Commands such as git status, git diff, and git commit are run here.
2. The bare repository
The central source-of-truth for code is stored in a separate bare repository:
/cluster/git/video-classification.git
A bare repository contains Git history and references but no checked-out working tree. That makes it suitable as a private remote for one or more working checkouts.
In this project, origin is cluster-local rather than GitHub or GitLab:
git remote -vThis typically shows something like:
origin /cluster/git/video-classification.git (fetch)
origin /cluster/git/video-classification.git (push)
The branch arrangement is intentionally simple:
main ---> origin/main
The local main branch tracks origin/main, so ordinary pulls and pushes know which remote branch to use.
This is useful even without a public hosting service. The bare repository provides a clean synchronization point for the source tree while keeping unpublished research on the cluster.
Creating the setup from scratch
The following is a generalized example. Adapt it to your cluster’s storage policy and account layout.
Create the bare repository
mkdir -p /cluster/git
git init --bare --initial-branch=main /cluster/git/video-classification.gitThe --bare flag matters: this repo is intended to receive pushes, not to be edited directly. --initial-branch=main sets the branch that new clones expect. For a team, put this repository in an approved project directory where collaborators have access. Ask your cluster support team about shared group permissions; creating a repository does not automatically give teammates permission to push.
Clone a working checkout
mkdir -p /cluster/projects
git clone /cluster/git/video-classification.git /cluster/projects/video-classification
cd /cluster/projects/video-classificationIf the bare repository is empty, Git may not yet have a branch. After adding the initial files:
git branch -M main
git add .
git commit -m "Initial project setup"
git push -u origin mainThe -u option establishes the upstream relationship. After that, git push and git pull usually work without repeating the remote and branch names.
Verify the connection
git remote -vv
git branch
git status
git log --oneline --decorate -5An expected result looks conceptually like:
* main ash007 [origin/main] latest commit
The [origin/main] part confirms that the local branch is tracking the expected remote branch.
Protecting the repository with .gitignore
Before adding data or running experiments, define what should stay outside Git. A typical .gitignore includes:
# Python-generated files
__pycache__/
*.py[cod]
# Virtual environments
.venv/
# Generated experiment output
artifacts/
checkpoints/
logs/
# Local datasets
datasets/
The ignore file acts as a guardrail. To understand why a path is ignored:
git check-ignore -v path/to/fileIf a large file was accidentally staged, remove it from the index without deleting the local copy:
git restore --staged path/to/fileReview the result with git status before doing anything more destructive.
For deep-learning projects, generated outputs are often broader than this example. You may also want to exclude directories such as wandb/, mlruns/, .pytest_cache/, .ruff_cache/, or notebook checkpoints if they appear in your workflow.
The daily working loop
Once this setup exists, a normal development session is deliberately simple:
cd /cluster/projects/video-classification
git status --short
git pull --ff-onlyThe --ff-only flag prevents Git from silently creating a merge commit when the local branch and remote branch have diverged. If it refuses to pull, inspect the situation:
git branch
git log --oneline --decorate --graph --all -10
git diffThen edit the code and run a lightweight validation step. For this project which utilizes uv:
uv sync
uv run python -m compileall srcBefore committing, review both the file list and the actual patch:
git status --short
git diffThen create a small, meaningful checkpoint:
git add src scripts
git commit -m "Add video classification baseline pipeline"
git pushKeep each commit focused on one coherent change. That makes it easier for a teammate to review and for you to revert if an experiment breaks. Stage configuration, dependency, and job-script changes too when they belong to the same change.
Collaborating without overwriting each other
Each teammate should have their own working checkout and Python environment. Share the bare remote and datasets, but use separate output directories for experiments. Editing the same checkout can mix uncommitted changes, and switching its branch can change files another person’s job is using.
For a small team, agree on who merges changes into main. Start a piece of work from an up-to-date branch:
git switch main
git pull --ff-only
git switch -c feature/video-loader
# Edit and check your changes.
git add src scripts
git commit -m "Add video dataset loader"
git push -u origin feature/video-loaderA teammate can inspect that branch from their own checkout:
git fetch origin
git diff origin/main...origin/feature/video-loaderAfter review, the person integrating the change can run:
git switch main
git pull --ff-only
git merge origin/feature/video-loader
# Run the project’s checks before pushing.
git push origin mainIf Git reports conflicts, resolve them together before pushing. A bare remote stores branches but has no pull-request interface, so agree on a review process with your team. GitHub or GitLab provide that interface if you use a hosted remote.
Before submitting a training job, record git rev-parse HEAD with your experiment configuration. Use a committed checkout whose files you leave unchanged while the job is queued or running: Slurm does not take a snapshot of your source code when you submit it.
What uv contributes
Git records the project’s dependency recipe, not the installed virtual environment.
pyproject.tomldescribes the project and its direct dependencies.uv.lockrecords the resolved dependency versions..venv/is generated locally and remains ignored.
On a fresh checkout, the environment is recreated intentionally:
cd /cluster/projects/video-classification
uv sync
uv run python -c "import torch; print(torch.__version__)"This is more reproducible than relying on whatever happens to be installed globally on the login node. It also avoids transferring a virtual environment between machines.
A note of caution: uv.lock helps with dependency reproducibility, but it does not solve cluster-specific runtime constraints such as GPU availability, CUDA compatibility, or system-level library differences. A reproducible Python environment and a compatible HPC environment are related, but not identical concerns.
Login nodes versus compute jobs
An SSH login node is a coordination and development environment. It is appropriate for:
- Git commands;
- inspecting files and logs;
- editing code;
- checking metadata;
- small imports and syntax checks;
- submitting jobs.
Long training runs, GPU workloads, and dataset-scale processing should be submitted through the cluster scheduler. This example uses Slurm, which is common on HPC systems. A representative job script might look like:
#!/bin/bash
#SBATCH --job-name=video-baseline
#SBATCH --cpus-per-task=4
#SBATCH --mem=4G
#SBATCH --time=01:00:00
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
#SBATCH --gres=gpu:1
cd /cluster/projects/video-classification
export DATASET_ROOT=/cluster/datasets/ucf101
srun uv run --locked python scripts/train_fusion.py --mode late --epochs 10This assumes your project already provides scripts/train_fusion.py; replace it with your own training entry point and save the script as jobs/video-baseline.sh. Install dependencies before submitting. Here, uv run --locked requires the committed lockfile to match the dependency configuration rather than updating it during training.
The exact partition, GPU option, memory limit, and time limit depend on the cluster. Check your cluster documentation before copying this script. Create the log directory before submitting, because Slurm opens the output files when the job starts:
cd /cluster/projects/video-classification
mkdir -p logs
sbatch jobs/video-baseline.sh
squeue --meThe sbatch command submits the script and returns a job ID. Use that ID to inspect a completed job and its logs:
squeue --job JOB_ID
sacct --jobs JOB_ID --format=JobID,State,Elapsed,MaxRSS
tail -n 200 logs/video-baseline-JOB_ID.out
tail -n 200 logs/video-baseline-JOB_ID.errFor a short interactive test on an allocated compute node, use srun instead of running the workload directly on the login node:
srun --pty --time=00:15:00 --cpus-per-task=2 --mem=2G bashThe key principle is simple: Git tracks job scripts and configuration, while Slurm executes the workload, and the filesystem stores logs and checkpoints.
If Slurm is new to you, these are useful starting points:
- HPC Carpentries: Introduction to high-performance computing, which introduces clusters, schedulers, and job submission.
- Slurm Quick Start User Guide, the official overview of jobs and commands such as
sbatch,srun, andsqueue.
Keeping SSH sessions persistent with tmux
Imagine you are editing a job script and watching its logs when your Wi-Fi drops. A normal SSH terminal can disappear along with the programs attached to it. tmux keeps a terminal session on the remote machine so you can reconnect to that workspace. Start it after connecting to the cluster; starting it on your laptop only preserves the local terminal.
Start a named session
# From your laptop; replace cluster-host with your cluster’s SSH address.
ssh YOUR_USERNAME@cluster-host
# On the cluster.
hostname
tmux new -s video-projectNote the hostname so you know where this session lives. Inside tmux, work as usual:
cd /cluster/projects/video-classification
mkdir -p logs
sbatch jobs/video-baseline.sh
squeue --me
# Replace JOB_ID with the number returned by sbatch.
tail -f logs/video-baseline-JOB_ID.outOnce the job starts and its log exists, tail -f displays new output as it arrives. Press Ctrl+C to stop following the log; the submitted job keeps running.
Detach and return later
Press Ctrl+b, release both keys, then press d. This detaches your terminal while leaving the session running. You can then close SSH.
After reconnecting to the same login node, list and attach to your session:
tmux ls
tmux attach -t video-projectYour terminal workspace should still be there. If the cluster routes SSH connections to different login nodes, follow its documentation for reconnecting to the original node: a session on one node is not available on another.
A few useful shortcuts
The default prefix is Ctrl+b. Release it before pressing the next key.
| Shortcut | Action |
|---|---|
Ctrl+b, then d |
Detach from the session. |
Ctrl+b, then c |
Create another terminal window. |
Ctrl+b, then n |
Move to the next window. |
Ctrl+b, then % |
Split into left and right panes. |
Ctrl+b, then " |
Split into top and bottom panes. |
Ctrl+b, then an arrow key |
Move between panes. |
For example, keep your editor in one pane and job logs in another. Type exit in a shell when you are finished with it; closing the last remaining pane ends the session. The official tmux getting-started guide explains sessions, windows, panes, and key bindings in more detail.
How tmux and Slurm fit together
Use tmux for your terminal workspace and Slurm for compute resources. A job submitted with sbatch is managed by Slurm and continues independently of SSH or tmux. You do not need tmux to keep a batch training job running.
For interactive debugging, start your cluster’s approved srun --pty allocation from inside tmux. Reattaching can let you return to that interactive shell while the allocation remains active. Its time limit still applies, and tmux does not reserve GPUs or make training on a login node appropriate.
Session persistence covers disconnections, not node reboots, administrative cleanup, or expired allocations. Keep saving your code, committing changes, and writing model checkpoints. Check whether your cluster permits tmux on login nodes.
Moving between a laptop and the cluster
There are two patterns that are worth distinguishing, because they serve different purposes.
Git-based synchronization
If the laptop and cluster can both reach the same Git hosting service, commits can move through that shared remote. You can also reach the cluster-local bare repository over SSH, if your cluster permits it:
# From your laptop; use your username and the cluster’s SSH address.
git clone YOUR_USERNAME@cluster-host:/cluster/git/video-classification.gitThe laptop remote uses an SSH address, while the cluster checkout can use a filesystem path. Both point to the same bare repository. Commit and push from one checkout, then pull from the other before continuing work.
rsync for file transfer
Use Git to exchange tracked source changes. For datasets or other untracked files, rsync copies only the changes needed on subsequent transfers. For example:
rsync -avh -P local-dataset/ YOUR_USERNAME@cluster-host:/cluster/datasets/ucf101/The trailing slash is significant: it controls whether the contents of the directory or the directory itself are copied. Before a large transfer, use a dry run:
rsync -avhn SOURCE/ DEST/Verify the destination and available storage before transferring a large dataset. For project-directory transfers, exclude .git/, virtual environments, credentials, and caches. Use --delete only when you intend to remove destination files that are absent from the source.
Common failure modes
“I committed the dataset”
Check .gitignore before the first git add. If a file is staged but not committed, unstage it with git restore --staged and confirm that the local dataset still exists. If it is already committed, adding it to .gitignore does not remove it from tracking or history. Use git rm --cached path/to/file to stop tracking the file while retaining the local copy, then commit that change. Earlier commits still contain the file and always coordinate with your team before rewriting shared history to remove large data.
“My push went somewhere unexpected”
Run:
git remote -v
git branch -vvNever assume that origin means GitHub. In this above setup it means the cluster-local bare repository.
“The code works in one checkout but not another”
Check the current directory, branch, commit, and environment:
pwd
git status --short
git log -1 --oneline
which uv
uv run python -c "import sys; print(sys.executable)"“Training is slow or the login node is overloaded”
Stop the long-running process and move the workload into a Slurm job. Login-node convenience should not become a resource-management problem for other users.
“The experiment cannot be reproduced”
Commit the source code, pyproject.toml, uv.lock, configuration, and job script. Record the dataset location, random seed, model settings, and output path. Do not rely on shell history as the experiment record.
Conclusion
The working checkout is where code changes are made. The bare repository is the versioned synchronization point for the source tree. This is not a universal blueprint for setting up projects and working with your teammates. Honestly i was just fed up :( , of constantly mentioning what to do and what not to. This is a practical pattern that works well for research projects with large data, environment-driven execution. With this setup, you can spend more time on the model and less time figuring out which files, environment, or terminal session you were using.
Further reading
- HPC Carpentries: Using High-Performance Computing Systems — a beginner course covering the shell, remote access, and scheduling work on a cluster.
- Princeton Research Computing: Your first Slurm job — a worked job-submission tutorial. Adapt its cluster-specific settings to your own system.
- Pro Git: Distributed workflows — how teams organize work around a shared repository, with options beyond the small-team workflow here.
- tmux: Getting started — the project’s guide to sessions, panes, navigation, and customization.
- Astral blog: uv — Unified Python packaging — background on uv’s project and environment tools. Also take a look at locking and syncing guide when setting up your project.