Tutorial Goals

By the end you’ll be able to:

  1. Look at and search through files without opening an editor
  2. Reach for a small existing tool instead of writing C
  3. Handle archives and do basic text surgery with sed and awk

Getting Started

Open the T6: The Command Line container, then change directory into tutorial:

cd ~/tutorial

Motivation

Right now, you solve problems in two ways: you write a C program, or you click around in a file explorer. They have their usecases, but there are some times when it’s harder.

Consider the following problem: You have a file of names with duplicates, and you want the unique ones, sorted. You may consider writing a C program: read the file, store the names, sort, remove duplicates, print. All that work just to process a single file.

Imagine if you could do this:

sort -u dupes.txt

That’s what this tutorial is about: the system is full of small programs, and you can command them instead of writing your own.

Viewing files

Often you just want to see what’s in a file, not edit it. Opening Vim or Nano for that is overkill, and risks accidental changes.

ProgramDescription
catPrint the whole file to the terminal
headPrint the first few lines
tailPrint the last few lines
lessScroll through a file, page by page

cat, short for “concatenate”, prints a file to the terminal. For a short file that’s all you need.

Note

Printing a file isn’t really what cat is for. Its actual job is to concatenate: join several files together (e.g. cat a.txt b.txt). Dumping one file to the terminal is just the most common side use.

View a file with cat:

  1. cat foo.c
  2. Try head and tail on a longer file (e.g. long.txt). What’s different?

Reading long files with less

For a long file, cat dumps everything at once and you lose the top. less lets you scroll and search, without the risk of editing.

Note

less is an improvement over an older program called more. In other words, less is more, and more is… less.

If you’ve used man, you’ve already used less; it powers that scrollable interface.

Tip

If you know Vim, you already know most of less.

KeyActionNotes
j, kMove down / up one linesame as vim
d, uScroll down / up half a screenalmost same as vim. See tip box below.
GGoes to last linesame as vim
{number}GGo to line {number}same as vim
/{pattern}, ?{pattern}Search forward / search backwardsame as vim
n, NNext / previous search resultsame as vim
Esc, followed by uClear search highlightingu for “unhighlight”
qQuit

Tip

In Vim, scrolling is done with Ctrl-D and Ctrl-U. In less, it’s just d and u. You can still hold Ctrl if that makes you happy.

Open the man page for any command (e.g. man kill) and try out the keybindings above.

Compressing data

Sometimes many files are bundled into one, often smaller. That bundle is called an archive. Two common ways:

  1. zip / unzip
  2. tar + gzip (seen as .tar.gz)

Your friend John sent you a .tar.gz. Extract it:

tar xf john_doe.tar.gz

Figure out what these flags mean, and make up a mnemonic:

  • czf on tar (create): tar czf out.tar.gz dir/
  • xf on tar (extract)
  • -r on zip: zip -r out.zip dir/

What a file really is

Here’s something that might rewire how you see your computer.

Unlike some operating systems, Linux doesn’t trust file extensions. It looks at the first few bytes of a file to decide what it is (see: magic bytes). The tool that does this is file.

Run file on john_doe.tar.gz. Is the type it reports the one you expected?

What exactly is a pptx file?

Run file -k slides.pptx. What type does it report? Using the appropriate tool, view its contents.

PS: it really is a normal PowerPoint file. That’s the point.

Finding files

This is the classic file-explorer job, except the explorer is slow when the file is buried deep, or when you don’t remember the name, only something inside it. Two tools cover both cases.

Finding by name (find)

find searches for files and folders by name.

find <path> <options>
OptionMeaning
-name "*.c"anything ending .c
-type ffiles only
-type ddirectories only

Examples

find . -name "*.txt"          # every text file here
find . -type f -name "*lab*"  # files with "lab" in the name
find ~ -type d -name "cs*"    # folders under home starting with "cs"

Tip

* is a wildcard that matches zero or more characters, so "*.txt" matches any file ending in .txt (e.g. .txt, arst.txt, a.txt).

Use find to answer these:

  1. There’s a directory with the word “find” in its name, but John doesn’t remember the exact name. Find it
  2. John claims there are 108 files in the extracted project directory. Do you agree?

Hint: check man find, and look for the -type option.

Hint: too many to count by eye? You may want to consider saving the list to a file with > (from a previous tutorial), then numbering the lines with cat -n.

Searching contents (grep)

grep searches the contents of a file. It’s so common that “grep” is now a word in the Oxford English Dictionary.

On its own it reads the one file you name, so searching a whole project means pointing it at a directory and asking it to walk down with -r.

grep <options> <pattern> <path>
OptionMeaning
-itodo also matches TODO
-revery file below <path>
-lnames, not lines

Examples

grep "hello" hello.c  # lines with "hello" in one file
grep -r "hello" .     # the same, but every file below here
grep -ril "todo" .    # names of files mentioning "todo", any case

Tip

Think of grep like Ctrl-F, but across all your files at once.

Tip

grep is case-sensitive by default.

For each task, should you use find or grep? How do they differ?

  • Find a file by its name
  • Find the file that contains the word “hi”

There’s some unfinished code lurking in the project. John left some “TODO” comments, though they might be in any mix of upper or lower case. He thinks that project/part-3/notes/old/find-me/file-037.c is the only file with a “TODO”. What do you think?

  1. Check man grep for the options that ignore case, and that search every file in a folder
  2. Search every file in project for “TODO”, ignoring case

Note: we’ll worry about fixing that “TODO” later.

Editing text

A UNIX habit is to store data as plain text, which makes it easy to manipulate. Two heavier tools for that are sed and awk. Both are really little programming languages of their own, so they edge toward “writing code”. Treat this as a quick preview, not something to learn in depth.

Rewriting text (sed)

sed is a stream editor. It reads a file line by line, runs your script on each line, and prints what comes out.

sed <options> '/<search>/<flags>d'           <path>  # delete every line that matches
sed <options> 's/<search>/<replace>/<flags>' <path>  # find and replace
FlagMeaning
Iignore case
gevery match on the line, substitute only
OptionMeaning
-iwrite the file

Examples

sed '/malloc/d' hello.c          # print without the matching lines
sed 's/int/long/g' hello.c       # print with every "int" swapped
sed -i 's/todo/TODO/gI' hello.c  # write the change into the file

It prints to the screen by default. Use -i to edit the file in place.

Note

Does this remind you of Vim?

  1. Delete the line containing “TODO”, in any case: sed -i '/todo/Id' project/part-3/notes/old/find-me/file-037.c
  2. Check that the line is gone: cat project/part-3/notes/old/find-me/file-037.c

Tables (awk) (Extra)

If you want to learn more, here’s a good tutorial: https://www.grymoire.com/Unix/Awk.html

The UNIX philosophy

The tools you’ve used share a design principle: each does one thing. cat prints. find finds. grep searches. None of them try to do everything.

This is the UNIX philosophy: programs should be small and do one thing well, in contrast to monolithic applications (e.g. Microsoft Word) that try to do everything.

Tip

“Do one thing well” is only one part of the UNIX philosophy. You can read the rest here: https://en.wikipedia.org/wiki/Unix_philosophy#Origin

Here are a few more of these single-purpose tools:

ProgramDescription
echoPrint text to the terminal
sortSort lines of text
uniqReport or omit repeated lines
wcCount lines, words, and characters

Try out the commands above. Use man to help you.

With the exception of echo, most commands follow <cmd> <filename> (e.g. sort dupes.txt, wc dupes.txt).

Think back to counting the files in project:

find project -type f > files.txt
cat -n files.txt

It took two commands and a file in between, which is quite troublesome. Every extra tool in the chain needs another file in between, so a problem that needs five tools leaves you with four files to name and clean up.

Getting tools to work together directly is for a future tutorial.

Hint

If you did the optional exercise in a previous tutorial, you’ve already seen a way around this:

./run < test/test_cases/1.in | diff test/test_cases/1.out -

| sent the output of ./run straight into diff, with no file in between.

Graded Task(s)

Bob needs your help! He doesn’t manage his files well, and he has so many of them. He remembers that he left a TODO comment in exactly one of the .c files, and it’s full of leftover [debug] print lines. He was hoping you would help him look through all the files. He “tar”-ed the file (that’s what he claims), but you realise tar doesn’t work on it.

Enter the Graded Task Folder

Go to ~/tutorial-graded before starting

Help Bob make the following changes:

  1. The archive won’t extract with tar. Find out what it really is, then extract it

  2. Locate the one .c file that contains a TODO

  3. Delete every line containing [debug] (in any case) from that file

  4. Run verify . inside robert to check your work and get the flag:

    abc@tut6:~/tutorial-graded/robert$ verify .
    

Info

The verify program runs on the entire directory, because Bob wants to make sure that you didn’t mess up the other files.

Before you leave: verify your progress

Run this command inside your Tutorial environment:

progress

It shows whether you have successfully completed the tutorial and which individual tasks have been recorded as complete.

The progress command is available only inside a Tutorial environment, not in the general CS1010 WebTop environment.