Linux for Bioinformatics on Zaratan
This lab introduces the Linux skills you will use in BIOI611. Work through the main lesson in order, using your own Zaratan account. By the end, you should be able to find and manage files, filter a small biological dataset, check your resource usage, and submit and inspect a Slurm job.
Where to run commands: use the Bash terminal opened through Zaratan Shell Access. Copy commands inside a bash block, without copying a terminal prompt or the expected output. Lines starting with # are comments. Instructions explicitly labeled local computer are the exception.
Reading this as a notebook? The cells are lesson text, not Python programs. Copy the Bash examples into a terminal. In a Python-kernel notebook, %%bash runs a cell in a separate Bash process; a cd or variable assignment there does not persist into the next cell. Do not type %%bash into the portal terminal.
1. Open your Zaratan terminal
- Sign in to Zaratan OnDemand with your UMD account and complete any authentication prompts.
- Select Clusters → Zaratan Shell Access.
- Wait for a Linux command prompt, then run:
whoami
hostname
pwd
id
whoami should print your Directory ID. hostname identifies the machine, and pwd prints your current directory. id shows your Unix account and groups; BIOI611 allocation members should have the zt-bioi611 group.
Portal login, terminal access, and permission to submit a job are separate checks. If the terminal or BIOI611 group is missing, send your instructor your Directory ID and the exact error. Never include your password or an authentication code. Account provisioning problems may require UMD HPC support.
Login nodes and compute nodes
The shell opens on a login node, where you can navigate, edit scripts, transfer files, and run the tiny examples in this lesson. Submit substantial analyses, such as read alignment, to compute nodes through Slurm. Opening a terminal does not allocate compute resources. Follow the UMD HPC policies.
2. Choose the right storage and check quotas
| Location | Intended use | Quota and backup |
|---|---|---|
Home: ~ or $HOME |
Small scripts and configuration files | Personal 10 GB soft quota; backed up. A temporary increase to 20 GB is allowed for up to seven days. |
~/scratch.bioi611 |
Data and outputs for active analyses | Shared BIOI611 project quota; not backed up. Use this for job input/output. |
~/SHELL.bioi611 |
Medium-term storage between analyses | Shared BIOI611 project quota; not backed up. Unavailable on compute nodes. |
Move retained data off scratch when it is no longer needed for active work. SHELL is not a backup: keep a separate copy of irreplaceable data. These rules are described in UMD's storage guide and storage policies.
The similarly named Bash shell is the program interpreting your commands. $SHELL is an environment variable that normally names your login shell, such as /bin/bash. SHELL storage is a filesystem; it is unrelated to that variable.
Inspect the shortcuts
ls -ld ~/scratch ~/scratch.bioi611 ~/SHELL ~/SHELL.bioi611
readlink -f ~/scratch.bioi611
readlink -f ~/SHELL.bioi611
The -> in the listing marks a symbolic link, a shortcut to another path. The BIOI611 scratch link normally resolves to /scratch/zt1/project/bioi611/user/YOUR_DIRECTORY_ID. The generic scratch and SHELL links may point to another project if you belong to several projects. Use the explicit BIOI611 links in this course. If a link is absent or inaccessible, stop and ask the instructor rather than creating a replacement in your home directory.
Storage quota versus compute allocation
home_quota
scratch_quota
shell_quota
sbalance
The first three commands report storage usage and limits. sbalance reports computational allocation usage in service units (SU). A personal scratch quota displayed as unlimited does not override the shared project quota. Everyone in the project contributes to that limit. See the UMD command reference.
du -sh DIRECTORY measures the disk space used by a directory; it does not report your allocation quota. Avoid scanning the entire shared project tree just to check your own usage.
3. Prepare a small practice workspace
Use these synthetic teaching files, which contain no real study or student data:
- sequences.fasta: three short sequences.
- counts.tsv: four genes, two samples, and a header row. Fields are separated by tabs.
- hello_slurm.sh: a minimal batch job.
In your Zaratan terminal, first enter your personal BIOI611 scratch directory:
cd ~/scratch.bioi611
pwd -P
If cd fails, resolve the access problem before continuing. Create a dedicated workspace, then create the two small datasets directly. This setup does not require downloading files or waiting for the course website to be published:
mkdir -p linux_practice/{data,results,scripts,logs}
cd linux_practice
cat > data/sequences.fasta <<'FASTA'
>seq1 synthetic teaching sequence
ACGTACGTACGT
>seq2 synthetic teaching sequence
GGGAAATTTCCC
>seq3 synthetic teaching sequence
ATATCGCGATAT
FASTA
printf '%s\t%s\t%s\n' \
gene_id sample_1 sample_2 \
geneA 600 700 \
geneB 100 900 \
geneC 800 200 \
geneD 550 650 > data/counts.tsv
ls -lh data
Paste the complete block, including the final FASTA marker on a line by itself. The lines between the markers are file contents, not commands; printf creates real tab separators in the count table. These commands overwrite only the two named practice files when repeated, including an empty file left by a failed download. We will create the job script in section 12. The links above provide equivalent downloadable copies if you prefer uploading files through the portal's Files menu.
Unless an example explicitly changes directories, run the rest of the main lesson from ~/scratch.bioi611/linux_practice. On returning to a new terminal, enter that directory again.
4. Navigate the filesystem
The filesystem starts at /, the root directory. An absolute path starts with /; a relative path is interpreted from your current directory. For example, data/counts.tsv is relative to linux_practice.
| Expression | Meaning |
|---|---|
. |
Current directory |
.. |
Parent directory |
~ or $HOME |
Your home directory |
cd - |
Return to the previous working directory |
cd with no argument |
Return home |
pwd
ls
ls -lh data
ls -la
ls -ld data
cd data
pwd
cd ..
readlink -f data/counts.tsv
cd ~
cd -
ls -l gives a detailed listing, -h uses readable sizes, and -a includes hidden names beginning with .. ls -ld data describes the directory itself. Options can be combined: ls -la means ls -l -a.
Linux names are case-sensitive: Data and data are different. Spaces separate command arguments. Quote a path containing spaces, and quote variables containing paths:
mkdir -p "results/practice notes"
ls -ld "results/practice notes"
printf '%s\n' "$HOME"
Prefer simple filenames without spaces in analysis projects. The shell expands wildcards before running a command: data/*.tsv means matching filenames, not a regular expression for their contents.
5. Terminal essentials and getting help
| Key or command | Use |
|---|---|
| Tab; Tab twice | Complete a command/path; list possible completions |
| Up/Down arrows | Recall previous commands; inspect them before pressing Enter |
| Ctrl+R | Search command history |
| Ctrl+C | Interrupt a foreground command or cancel an unfinished command line |
history |
Show recent commands |
ls --help |
Read a command's usage and options |
man ls |
Open its manual; press q to leave |
exit |
Close this shell session |
A continuation prompt such as > often means a quote or command is unfinished. Ctrl+C returns to a fresh prompt. A successful command may produce no output. echo $? immediately after a command shows its exit status: zero usually means success.
6. Copy, rename, create, and remove files
Run the following in your practice workspace:
cp -i data/counts.tsv results/counts_copy.tsv
mkdir -p results/backup
cp -i data/sequences.fasta results/backup/
cp -r data results/data_copy
mv -i results/counts_copy.tsv results/counts_renamed.tsv
touch results/notes.txt
ls -lh results
cp copies a file; cp -r copies a directory recursively. mv moves or renames files. -i asks before overwriting an existing destination. mkdir -p creates missing parents. touch creates an empty file if it does not exist, or updates its timestamp if it does. If you repeat the directory-copy example with an existing destination, cp may nest the source inside it; inspect the result with ls.
Practice deletion only on disposable files
rm does not send files to a desktop recycle bin. Check pwd and ls before deleting anything.
mkdir -p results/empty_folder
rmdir results/empty_folder
cp data/counts.tsv results/delete_me.tsv
rm -i results/delete_me.tsv
mkdir -p results/disposable
touch results/disposable/test.txt
rm -ri results/disposable
Answer the deletion prompts only for these practice files. rmdir removes an empty directory. rm -r removes a directory and its contents. A nonempty directory does not require -f. The force option suppresses prompts and ignores missing files; it is unnecessary here.
7. View files and redirect output
cat data/sequences.fasta
head -n 3 data/counts.tsv
tail -n 2 data/counts.tsv
less data/counts.tsv
Use cat for small files and less for larger text. Inside less, use Space to advance, b to go back, /geneC then Enter to search, n for the next match, and q to quit. head and tail print ten lines by default; -n changes that count. tail -n +2 starts at line two and is useful for skipping a header.
Standard input, output, and error
A program reads standard input, writes results to standard output, and writes diagnostics to standard error. By default, output and errors appear in the terminal.
| Operator | Meaning |
|---|---|
> |
Write standard output to a file, overwriting it |
>> |
Append standard output to a file |
2> |
Write standard error to a file |
2>&1 |
Send standard error to the current standard-output destination |
\| |
Pass standard output to the next command's standard input |
printf 'First observation\n' > results/notes.txt
printf 'Second observation\n' >> results/notes.txt
cat results/notes.txt
ls data/does_not_exist 2> logs/expected_error.txt
cat logs/expected_error.txt
The failed ls is intentional. Never redirect output onto the input file: sort file > file truncates the input before sort reads it. Write to a new filename instead. The order of redirections matters: command > log.txt 2>&1 combines both streams in the log. A plain pipe carries standard output only.
8. Search and filter biological data
Find files and count records
find data -type f -name '*.fasta'
wc -l data/counts.tsv
grep '^>' data/sequences.fasta
grep -c '^>' data/sequences.fasta
Quote the pattern passed to find so the shell does not expand it first. The table has 5 lines: one header and four genes. The FASTA header count is 3. ^> means a > at the start of a line; counting every line of a FASTA file does not count sequences.
A pipe connects commands:
grep '^>' data/sequences.fasta | wc -l
tail -n +2 data/counts.tsv | wc -l
Expected counts are 3 sequences and 4 genes, respectively.
Select columns and filter rows
cut -f1,2 data/counts.tsv
awk -F '\t' 'NR > 1 && $2 > 500 && $3 > 500' data/counts.tsv
awk -F '\t' 'NR > 1 && $2 > 500 && $3 > 500' data/counts.tsv | wc -l
cut uses tabs by default. In awk, -F '\t' selects a tab separator, $2 and $3 mean columns two and three, and NR > 1 skips the header. This selects geneA and geneD, so the final count is 2. Without explicitly handling the header, a numeric-looking filter can accidentally include text.
Preserve the header when saving a filtered table:
awk -F '\t' 'NR == 1 || ($2 > 500 && $3 > 500)' data/counts.tsv > results/high_counts.tsv
cat results/high_counts.tsv
Expected output, displayed here with tabs between columns:
gene_id sample_1 sample_2
geneA 600 700
geneD 550 650
Extract sequence identifiers
grep '^>' data/sequences.fasta | sed 's/^>//; s/ .*//' > results/sequence_ids.txt
cat results/sequence_ids.txt
sort results/sequence_ids.txt | uniq -c
sed removes the leading > and the description after the first space. The IDs are seq1, seq2, and seq3. uniq combines adjacent repeated lines, so sort first when counting duplicates across a file.
Wildcards versus regular expressions
Use grep -F for literal text and grep -E for extended regular expressions:
grep -F 'geneA' data/counts.tsv
grep -E '^gene(A|D)[[:space:]]' data/counts.tsv
grep -E '^>seq[0-9]+' data/sequences.fasta
In extended regular expressions, ^ anchors the start, $ anchors the end, . matches a character, [0-9] matches a digit, * repeats zero or more times, + repeats one or more times, ? makes the preceding item optional, | means alternatives, and {n} requests a repetition count. Use [[:digit:]], [[:space:]], and [[:alnum:]_] for common character classes. The \d shorthand belongs to some other regex dialects; do not assume it works with ordinary grep or grep -E.
Compare the filename wildcard *.fasta with the regex .*\.fasta$. The latter escapes the literal dot and anchors the filename ending. See the GNU grep manual for syntax details.
9. Permissions and executable scripts
id
ls -ld data
ls -l data/counts.tsv
In -rw-r--r--, the first character describes the file type; the next three groups describe owner, group, and others. The permission letters have different meanings for files and directories:
| Permission | File | Directory |
|---|---|---|
r |
Read contents | List entry names |
w |
Change contents | Create, remove, or rename entries, normally together with x |
x |
Execute as a program | Traverse/access entries by name |
Deleting or renaming a file is governed primarily by its parent directory permissions, not by the file's write bit. ACLs and sticky-directory rules can impose additional restrictions. SHELL uses AFS access controls; the simple scratch examples here are not instructions for managing SHELL permissions.
Run a small Python script
Load the cluster's Python environment, then create a script. The lines between the two PY markers are file contents; the closing PY must be on a line by itself.
module load python
cat > scripts/hello.py <<'PY'
#!/usr/bin/env python3
print("Hello, BIOI611!")
PY
chmod u-x scripts/hello.py
python3 scripts/hello.py
This works without execute permission because Python reads the file. To run the file itself, add execute permission:
chmod u+x scripts/hello.py
./scripts/hello.py
ls -l scripts/hello.py
Both commands print Hello, BIOI611!. The first line, called a shebang, selects an interpreter. /usr/bin/env python3 finds Python through PATH. A missing execute bit normally causes Permission denied; a nonexistent interpreter path can cause a different error even when the script exists.
Symbolic permissions include u+x (add owner execute), g-w (remove group write), and o=r (set others to read only). Numeric modes add read = 4, write = 2, and execute = 1; for example, chmod 750 scripts/hello.py gives the owner all three permissions, the group read/execute, and others none. Change permissions only on your own practice files.
10. Edit, save, and inspect a script
You can edit a file using the portal's file editor or use vi/vim in the terminal. Vim supports the basic vi commands used in this exercise.
How vi/vim modes work
In vi/vim, what a key does depends on the current mode. Start with the three modes shown in this diagram:
In the diagram, “vi mode” means normal mode, “Input mode” means insert mode, and “Command mode” means the command-line prompt opened with :.
| Mode | What you do there | How to enter or leave |
|---|---|---|
| Normal mode (“vi mode”) | Move around and use editing commands. This is the starting mode. | Press Esc to return from insert mode or cancel the command-line prompt. |
| Insert mode (“Input mode”) | Type text into the file. | From normal mode, press i to insert before the cursor, a to append after it, or o to open a new line below. Press Esc when finished typing. |
| Command-line mode (“Command mode”) | Enter commands such as :w (save), :q (quit), or :wq (save and quit). |
From normal mode, type :, then the command, and press Enter to execute it. |
A typical sequence is i → type text → Esc → :wq → Enter. The diagram abbreviates the exit command as wq; when starting in normal mode, include the colon: :wq. Commands such as :w return to normal mode after running; :wq saves and exits the editor.
Create and save a script
Open your practice file:
vi scripts/editor_practice.py
- Press
ito enter insert mode. - Type
print("I saved a script on Zaratan."). - Press Esc to leave insert mode.
- Type
:wq, then press Enter, to save and quit.
If you want to abandon unsaved edits, press Esc, type :q!, and press Enter. :w saves without quitting. :q quits if there are no unsaved changes.
cat scripts/editor_practice.py
python3 scripts/editor_practice.py
Expected output: I saved a script on Zaratan. If you create a shell script on Windows, save it with LF (Unix) line endings, not CRLF. Slurm can reject scripts containing DOS line breaks.
11. Environment variables and available software
printf '%s\n' "$HOME"
printf '%s\n' "$SHELL"
printf '%s\n' "$PATH" | tr ':' '\n'
command -v python3
module list
module avail python
PATH is a colon-separated list of directories searched for commands. To temporarily add your own program directory while preserving the existing search path:
mkdir -p "$HOME/bin"
export PATH="$HOME/bin:$PATH"
This change lasts for the current shell and processes launched from it. If you need it every session, add that exact line once to ~/.bashrc using an editor; do not repeatedly append it or replace the whole file. Batch scripts should explicitly initialize the environment they need.
Use module avail to discover software before installing another copy. module load makes a package available and module list shows what is loaded. Record the version used for an analysis. Personal Miniforge installation, containers, and source compilation are covered in the optional extensions.
12. Submit your first Slurm job
This first job uses only standard Linux commands and the supplied tiny FASTA file. No Conda environment or RNA-seq reference is required.
Check the allocation and create the script
sbalance
BIOI611 students should see the bioi611-class account. If it is absent or has no available resources, contact the instructor before submitting. Create scripts/hello_slurm.sh by pasting the entire block below into your terminal, including the final SLURM_SCRIPT marker. The quoted marker preserves the dollar signs in the script for Slurm to expand when the job runs. This saves the script; it does not submit or execute the job:
cat > scripts/hello_slurm.sh <<'SLURM_SCRIPT'
#!/bin/bash
#SBATCH --job-name=bioi611_hello
#SBATCH --account=bioi611-class
#SBATCH --partition=standard
#SBATCH --time=00:02:00
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=1
#SBATCH --mem=1G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
source /etc/profile
set -euo pipefail
cd "$SLURM_SUBMIT_DIR"
printf 'Job ID: %s\n' "$SLURM_JOB_ID"
printf 'Host: '
hostname
date
printf 'Working directory: '
pwd
grep -c '^>' data/sequences.fasta > results/sequence_count.txt
printf 'Sequences: '
cat results/sequence_count.txt
SLURM_SCRIPT
Check the saved script before submitting it:
bash -n scripts/hello_slurm.sh
cat scripts/hello_slurm.sh
bash -n checks syntax without running the script; no output normally means the syntax check passed.
All #SBATCH directives come before executable commands. This requests one task, one CPU, 1 GB of memory, and at most two minutes of running time. The time limit does not include waiting in the queue. source /etc/profile initializes the cluster environment; set -euo pipefail makes this script stop on failures, unset variables, or failed pipeline stages.
Submit from the practice directory
cd ~/scratch.bioi611/linux_practice
mkdir -p logs results
sbatch scripts/hello_slurm.sh
Slurm prints Submitted batch job followed by a job ID. Record that number. The logs directory must exist before submission because Slurm opens the log files before running the script. %x becomes the job name and %j the job ID. Do not run the script using bash scripts/hello_slurm.sh: that would bypass Slurm.
Monitor and check completion
squeue -u "$USER"
PD means pending and R means running. A pending job may be waiting for resources; check the reason shown in the queue. A short job may finish before you see it in squeue. Disappearing from the queue does not itself prove success.
In the following examples, replace 123456 with your actual job ID:
sacct -j 123456 --format=JobID,JobName,State,ExitCode,Elapsed,AllocCPUS,MaxRSS
cat logs/bioi611_hello-123456.out
cat logs/bioi611_hello-123456.err
cat results/sequence_count.txt
Look for COMPLETED, exit code 0:0, and a sequence count of 3. Accounting information can take a short time to appear. MaxRSS may be reported on the .batch step. FAILED, OUT_OF_MEMORY, and TIMEOUT require checking the error log and the requested resources.
To cancel one of your own pending or running jobs, substitute its ID:
scancel 123456
Use UMD's job-submission guide for resource options. The optional extensions show how the same ideas apply to STAR after its software and real input files are ready.
13. Compress, archive, and transfer files
Compress a copy, keeping the original data
cp data/counts.tsv results/counts_for_compression.tsv
gzip results/counts_for_compression.tsv
zcat results/counts_for_compression.tsv.gz | head -n 3
gunzip results/counts_for_compression.tsv.gz
By default, gzip replaces its input with a .gz file, and gunzip reverses that operation. zcat reads compressed text without permanently decompressing it. These tiny files may not become smaller after compression because of format overhead.
Create, inspect, then extract an archive
tar -czvf results/practice_data.tar.gz data/sequences.fasta data/counts.tsv
tar -tzvf results/practice_data.tar.gz
mkdir -p results/extracted
tar -xzvf results/practice_data.tar.gz -C results/extracted
ls results/extracted/data
du -sh data results
c creates, t lists, and x extracts an archive; z uses gzip, v prints names, and f introduces the archive filename. Only after checking the extracted files, optionally remove the practice archive with rm -i results/practice_data.tar.gz. Do not delete the archive before the extraction step.
Transfer between your computer and Zaratan
The portal's Files menu can upload and download files. An alternative is scp. Run the following on your local computer, not inside the Zaratan terminal, and replace YOUR_DIRECTORY_ID:
scp local_notes.txt YOUR_DIRECTORY_ID@login.zaratan.umd.edu:/scratch/zt1/project/bioi611/user/YOUR_DIRECTORY_ID/linux_practice/results/
scp YOUR_DIRECTORY_ID@login.zaratan.umd.edu:/scratch/zt1/project/bioi611/user/YOUR_DIRECTORY_ID/linux_practice/results/high_counts.tsv .
The first command uploads a local file you have created; the second downloads your filtered table to the current local directory. Connections may require UMD authentication or VPN depending on your setup; follow UMD's connection instructions. For repeated directory transfers, rsync is another option; preview with its --dry-run option before a large transfer and avoid --delete in beginner exercises.
14. Independent exercise and submission
Using your own account and practice files:
- Report your Directory ID, login hostname, and the resolved path of
~/scratch.bioi611. - Copy the FASTA file to a new name under
results/, then count its sequence headers and save the count in a text file. - Save a table containing the header and genes with counts greater than 500 in both samples. Explain why a header is not a gene record.
- Submit the small Slurm job, record its job ID, and find its completion state and output log.
- Explain which command checks home storage, which checks scratch storage, and which checks computational allocation usage. State where you would keep active input data and whether that location is backed up.
Submit a short text file containing your commands, the sequence and filtered-gene counts, your job ID, completion state, and log filename. Use actual output from your own account. Successful completion means 3 sequences, 2 selected genes, and a completed job with exit code 0:0.
Troubleshooting checklist
| Symptom | First checks |
|---|---|
No such file or directory |
Run pwd and ls; check spelling, capitalization, and whether setup/downloads completed. |
Permission denied |
Inspect id, ls -l FILE, and ls -ld PARENT; do not try to repair shared course permissions yourself. |
command not found |
Check command -v COMMAND, module list, and the required environment. |
| Quota exceeded | Run all three quota commands; identify the filesystem before moving or deleting your own files. |
| Invalid Slurm account | Check sbalance and confirm BIOI611 allocation membership with the instructor. |
| Job failed or produced no output | Check sacct, .err, input paths, environment setup, and whether logs/ existed at submission. |
| Shell shows a continuation prompt | Complete the quoted command or press Ctrl+C and start again. |
Continue with optional Linux, software, and HPC extensions when you are comfortable with the main workflow.