Look at and search through files without opening an editor
Reach for a small existing tool instead of writing C
Handle archives and do basic text surgery with sed and awk
Getting Started
Open the T6: The Command Line container,
then change directory into tutorial:
cd ~/tutorial
Motivation
Right now, you solve problems in two ways:
you write a C program, or you click around in a file explorer.
They have their usecases, but there are some times when it’s harder.
Consider the following problem:
You have a file of names with duplicates, and you want the unique ones, sorted.
You may consider writing a C program:
read the file, store the names, sort, remove duplicates, print.
All that work just to process a single file.
Imagine if you could do this:
sort -u dupes.txt
That’s what this tutorial is about:
the system is full of small programs,
and you can command them instead of writing your own.
Viewing files
Often you just want to see what’s in a file, not edit it.
Opening Vim or Nano for that is overkill,
and risks accidental changes.
Program
Description
cat
Print the whole file to the terminal
head
Print the first few lines
tail
Print the last few lines
less
Scroll through a file, page by page
cat, short for “concatenate”, prints a file to the terminal.
For a short file that’s all you need.
Note
Printing a file isn’t really what cat is for.
Its actual job is to concatenate:
join several files together (e.g. cat a.txt b.txt).
Dumping one file to the terminal is just the most common side use.
View a file with cat:
cat foo.c
Try head and tail on a longer file (e.g. long.txt). What’s different?
Reading long files with less
For a long file, cat dumps everything at once and you lose the top.
less lets you scroll and search, without the risk of editing.
Note
less is an improvement over an older program called more.
In other words, less is more, and more is… less.
If you’ve used man, you’ve already used less; it powers that scrollable interface.
Tip
If you know Vim, you already know most of less.
Key
Action
Notes
j, k
Move down / up one line
same as vim
d, u
Scroll down / up half a screen
almost same as vim. See tip box below.
G
Goes to last line
same as vim
{number}G
Go to line {number}
same as vim
/{pattern}, ?{pattern}
Search forward / search backward
same as vim
n, N
Next / previous search result
same as vim
Esc, followed by u
Clear search highlighting
u for “unhighlight”
q
Quit
Tip
In Vim, scrolling is done with Ctrl-D and Ctrl-U. In less, it’s just d and u.
You can still hold Ctrl if that makes you happy.
Open the man page for any command (e.g. man kill)
and try out the keybindings above.
Compressing data
Sometimes many files are bundled into one, often smaller. That bundle is called an archive.
Two common ways:
zip / unzip
tar + gzip (seen as .tar.gz)
Your friend John sent you a .tar.gz. Extract it:
tar xf john_doe.tar.gz
Figure out what these flags mean, and make up a mnemonic:
czf on tar (create): tar czf out.tar.gz dir/
xf on tar (extract)
-r on zip: zip -r out.zip dir/
What a file really is
Here’s something that might rewire how you see your computer.
Unlike some operating systems, Linux doesn’t trust file extensions.
It looks at the first few bytes of a file to decide what it is (see: magic bytes).
The tool that does this is file.
Run file on john_doe.tar.gz.
Is the type it reports the one you expected?
What exactly is a pptx file?
Run file -k slides.pptx. What type does it report?
Using the appropriate tool, view its contents.
PS: it really is a normal PowerPoint file. That’s the point.
Finding files
This is the classic file-explorer job, except the explorer is slow when the file is buried deep, or when you don’t remember the name, only something inside it.
Two tools cover both cases.
Finding by name (find)
find searches for files and folders by name.
find <path> <options>
Option
Meaning
-name "*.c"
anything ending .c
-type f
files only
-type d
directories only
Examples
find . -name "*.txt" # every text file herefind . -type f -name "*lab*" # files with "lab" in the namefind ~ -type d -name "cs*" # folders under home starting with "cs"
Tip
* is a wildcard that matches zero or more characters, so "*.txt" matches any file ending in .txt (e.g. .txt, arst.txt, a.txt).
Use find to answer these:
There’s a directory with the word “find” in its name, but John doesn’t remember the exact name. Find it
John claims there are 108 files in the extracted project directory. Do you agree?
Hint: check man find, and look for the -type option.
Hint: too many to count by eye?
You may want to consider saving the list to a file with >
(from a previous tutorial),
then numbering the lines with cat -n.
Solution
find . -type d -name "*find*" # the directory is project/part-3/notes/old/find-mefind project -type f > files.txt # save the list of filescat -n files.txt # the last line is numbered 108
Searching contents (grep)
grep searches the contents of a file.
It’s so common that “grep” is now a word in the Oxford English Dictionary.
On its own it reads the one file you name, so searching a whole project means
pointing it at a directory and asking it to walk down with -r.
grep <options> <pattern> <path>
Option
Meaning
-i
todo also matches TODO
-r
every file below <path>
-l
names, not lines
Examples
grep "hello" hello.c # lines with "hello" in one filegrep -r "hello" . # the same, but every file below heregrep -ril "todo" . # names of files mentioning "todo", any case
Tip
Think of grep like Ctrl-F, but across all your files at once.
Tip
grep is case-sensitive by default.
For each task, should you use find or grep? How do they differ?
Find a file by its name
Find the file that contains the word “hi”
There’s some unfinished code lurking in the project.
John left some “TODO” comments, though they might be in any mix of upper or lower case.
He thinks that project/part-3/notes/old/find-me/file-037.c is the only file with a “TODO”.
What do you think?
Check man grep for the options that ignore case,
and that search every file in a folder
Search every file in project for “TODO”, ignoring case
Note: we’ll worry about fixing that “TODO” later.
Editing text
A UNIX habit is to store data as plain text, which makes it easy to manipulate.
Two heavier tools for that are sed and awk.
Both are really little programming languages of their own, so they edge toward “writing code”. Treat this as a quick preview, not something to learn in depth.
Rewriting text (sed)
sed is a stream editor.
It reads a file line by line, runs your script on each line, and prints what comes out.
sed <options> '/<search>/<flags>d' <path> # delete every line that matchessed <options> 's/<search>/<replace>/<flags>' <path> # find and replace
Flag
Meaning
I
ignore case
g
every match on the line, substitute only
Option
Meaning
-i
write the file
Examples
sed '/malloc/d' hello.c # print without the matching linessed 's/int/long/g' hello.c # print with every "int" swappedsed -i 's/todo/TODO/gI' hello.c # write the change into the file
It prints to the screen by default. Use -i to edit the file in place.
Note
Does this remind you of Vim?
Delete the line containing “TODO”, in any case:
sed -i '/todo/Id' project/part-3/notes/old/find-me/file-037.c
Check that the line is gone: cat project/part-3/notes/old/find-me/file-037.c
The tools you’ve used share a design principle: each does one thing.
cat prints. find finds. grep searches. None of them try to do everything.
This is the UNIX philosophy: programs should be small and do one thing well, in contrast to monolithic applications (e.g. Microsoft Word) that try to do everything.
Here are a few more of these single-purpose tools:
Program
Description
echo
Print text to the terminal
sort
Sort lines of text
uniq
Report or omit repeated lines
wc
Count lines, words, and characters
Try out the commands above. Use man to help you.
With the exception of echo, most commands follow <cmd> <filename> (e.g. sort dupes.txt, wc dupes.txt).
Think back to counting the files in project:
find project -type f > files.txtcat -n files.txt
It took two commands and a file in between, which is quite troublesome.
Every extra tool in the chain needs another file in between,
so a problem that needs five tools leaves you with four files to name and clean up.
Getting tools to work together directly is for a future tutorial.
Hint
If you did the optional exercise in a previous tutorial,
you’ve already seen a way around this:
| sent the output of ./run straight into diff, with no file in between.
Graded Task(s)
Bob needs your help!
He doesn’t manage his files well,
and he has so many of them.
He remembers that he left a TODO comment in exactly one of the .c files,
and it’s full of leftover [debug] print lines.
He was hoping you would help him look through all the files.
He “tar”-ed the file (that’s what he claims),
but you realise tar doesn’t work on it.
Enter the Graded Task Folder
Go to ~/tutorial-graded before starting
Help Bob make the following changes:
The archive won’t extract with tar.
Find out what it really is, then extract it
Locate the one .c file that contains a TODO
Delete every line containing [debug] (in any case) from that file
Run verify . inside robert to check your work and get the flag:
abc@tut6:~/tutorial-graded/robert$ verify .
Info
The verify program runs on the entire directory,
because Bob wants to make sure that you didn’t mess up the other files.
Before you leave: verify your progress
Run this command inside your Tutorial environment:
progress
It shows whether you have successfully completed the tutorial and which
individual tasks have been recorded as complete.
The progress command is available only inside a Tutorial environment, not in
the general CS1010 WebTop environment.