Lecture 13: File Reading 1
Watch on YouTube →
Overview
Professor Rosen introduces Python file handling through text and binary file examples, then demonstrates how file locations, relative paths, and operating-system conventions affect whether Python can open a file. Using Project Gutenberg’s Treasure Island text and the Lab 7 Hurricane Irma tracker, the lecture covers `open()`, `read()`, `readlines()`, line-by-line iteration, UTF-8 encoding, and the `data/` subfolder paths needed to load CSV data.
Key takeaways
- A bare filename such as `pg120.txt` is a relative path, so it works only when the file is in the directory Python is searching; checking the working folder is a first step in diagnosing `FileNotFoundError`.
- Use `read()` for one string containing the whole file, `readlines()` for a list of lines, or `for line in file` to process lines without storing the whole file in a separate collection.
- Text lines returned by `readlines()` retain their newline characters, so printing each line with `print()` can create apparent double spacing.
- On Windows, Project Gutenberg’s UTF-8 text may require `open('pg120.txt', encoding='utf8')` to avoid a codec decoding error.
- The Hurricane Irma lab relies on preserved relative paths: CSV files are under `data/` and map assets are under `images/`, both relative to the starter project.
- The `w` file mode overwrites existing content, whereas `a` appends; selecting a mode deliberately prevents accidental loss of a file’s contents.
Chapters
- The lecture focuses on reading files from a computer’s file system rather than relying on textbook files already available in an environment.
- Downloaded data and the Python script that reads it must be placed where the program can locate both files.
- Cloud-based workflows can obscure file organization, but file paths still matter when writing programs.
- A text file contains human-readable characters, while binary files such as PDFs, Word documents, and compiled programs generally need specialized tools or libraries.
- CSV means comma-separated values: rows are separated by line breaks, and columns by commas, making it useful for moving tables between programs such as Excel and Python.
- The lecture identifies `.txt`, `.csv`, `.md`, and `.py` as useful file types; Markdown adds formatting conventions such as bold text and hyperlinks to plain text.
- Lab 7 provides a ZIP archive containing `Irma.py`, data files, and other resources; students must unzip it and preserve its folder structure.
- The assignment plots Hurricane Irma’s path across the Gulf of Mexico using Python’s `turtle` graphics, with CSV records supplying location and storm information.
- The starter code handles the map and turtle setup; latitude and longitude are easy to reverse, which can make the plotted turtle appear in the wrong place.
- The demonstration creates `hello.py` as a small target text file and a separate file-reading script, saving both in the Documents folder.
- Keeping the target file beside the script allows a simple filename such as `hello.py` to work as a relative path.
- On Windows, enabling filename extensions in File Explorer helps distinguish files such as `.py`, `.pdf`, and `.txt`.
- The Java example uses a `Scanner`, a file input, a loop over lines, and exception handling to illustrate the extra setup involved.
- Python can open the text file with `open('hello.py')`; the returned file object can then be read and printed.
- A simple filename works only when the file is in the directory Python is using; otherwise, the path or working directory must be corrected.
- Windows paths use drive letters such as `C:`, while macOS and Linux use a root directory written as `/`; user files commonly live under a user’s Documents folder.
- An absolute path identifies a file from the file-system root, but hard-coding a username-specific path makes code less portable.
- For the reading demonstration, Project Gutenberg supplies a public-domain plain-text edition of Robert Louis Stevenson’s Treasure Island.
- On Project Gutenberg, choose the plain-text format rather than saving the web page as HTML.
- The lecture uses `pg120.txt`; right-clicking the plain-text link and choosing “Save link as” downloads the text file.
- If the browser saves the file in Downloads, move it to the working folder so the Python script can find it using a relative filename.
- A `FileNotFoundError` usually means the filename is wrong or the file is not in the directory Python is searching.
- In VS Code, opening the containing folder helps align the workspace and file paths; opening a single script may leave related files undiscovered.
- The file modes discussed are `r` for reading, `w` for writing and overwriting, and `a` for appending; the Gutenberg example on Windows may require `encoding='utf8'`.
- Calling `file.read()` loads the file’s contents into one string, which is convenient when the whole text is needed at once.
- The Treasure Island example prints `text[:3000]` to inspect the first 3,000 characters without dumping the entire book.
- Reading a large file all at once can produce unwieldy output, so this method is less suitable when processing data record by record.
- `file.readlines()` returns a list whose elements correspond to lines in the source file.
- Each line retains its newline character, so printing each element can produce double spacing: the stored newline is followed by `print()`’s own line break.
- A list of lines is useful for later processing, such as examining CSV records row by row.
- The third approach, `for line in file`, processes a file line by line without first storing the entire file in a string or list.
- The Treasure Island example counts occurrences with `line.count('pirate')`, then checks the related term `buccaneer` as well.
- The example finds 23 lowercase instances of “pirate” and raises the combined count to 53 after counting “buccaneer” and “buccaneers.”
- The Lab 7 ZIP file should be extracted with the operating system’s unzip or extract command rather than treated as an ordinary working folder inside the archive.
- Archives such as ZIP, 7z, and `.tar.gz` bundle files and preserve directory structure; compression reduces transfer size.
- Keep the supplied folder arrangement intact so `Irma.py` can find its data and image resources.
- The starter project places the hurricane CSV files in a `data` folder and map assets in an `images` folder beside `Irma.py`.
- A path such as `images/hurricane.gif` shows how to navigate from the script’s location into a subfolder; CSV files use the same pattern, for example `data/<filename>.csv`.
- The next step in Lab 7 is to read the selected CSV and use its latitude and longitude values to move the turtle along Irma’s recorded path.
Summary, takeaways, and chapters were generated by AI from the video's transcript and may contain errors. The video belongs to its creator, Professor Rosen.