"The Lazy River" — Iterators, Generators and Context Managers
A 40 GB transaction export will not fit in memory. The iterator protocol, generators and yield, lazy generator pipelines, itertools, exceptions and EAFP, pathlib and csv, and with-statements as Python's try-with-resources.
Story Opening
The out-of-memory crash was embarrassing, mostly because Arjun knew better. In Java he would never Files.readAllLines() a 40 GB file; he’d use Files.lines() and a lazy Stream, or a BufferedReader in a try-with-resources block.
He just hadn’t known what the Python equivalents were. Priya showed him in four lines:
from pathlib import Path
Path("txns.csv").write_text("id,amount\nT1,120.0\nT2,45.5\nT3,3000.0\n") # tiny sample
with open("txns.csv", encoding="utf-8") as f: # try-with-resources next(f) # skip the header line total = sum(float(line.split(",")[1]) for line in f) # streams line by line
print(total) # -> 3165.5A file object is an iterator that yields one line at a time. The generator expression pulls lines through lazily. At no point does more than one line live in memory. That one pattern — lazy iteration — runs through the entire Python data stack, from file reading to PyTorch’s DataLoader.
Java → Python: The Quick Map
| Java | Python |
|---|---|
Iterable<T> / Iterator<T> | Iterable (__iter__) / iterator (__next__) |
hasNext() + next() | next() until StopIteration is raised |
Lazy Stream pipeline | Chained generators / generator expressions |
Custom Spliterator | A generator function with yield |
Stream.limit, concat, iterate | itertools.islice, chain, count |
try-with-resources / AutoCloseable | with / context managers (__enter__, __exit__) |
| Checked exceptions | None — all exceptions are unchecked |
Files.lines(path) | open(path) — the file object iterates lines |
java.nio.file.Path | pathlib.Path |
Deep Dive: The Iterator Protocol
Every for loop in Python runs on two tiny methods:
- An iterable has
__iter__(), which returns an iterator. Lists, dicts, strings, files, ranges are iterables. - An iterator has
__next__(), which returns the next item or raisesStopIterationwhen exhausted. (Iterators also have__iter__returning themselves.)
for x in xs: is shorthand for this:
amounts = [120.0, 45.5, 3000.0]
it = iter(amounts) # calls amounts.__iter__()while True: try: x = next(it) # calls it.__next__() except StopIteration: # end of data is signalled by an exception, not hasNext() break print(x)
# next() accepts a default instead of raising:print(next(iter([]), "empty")) # -> emptyIterables vs iterators: the one-shot trap
A list can be iterated many times — each for asks for a fresh iterator. An iterator (a generator, a file, a map object, a zip) can be consumed once. Java Streams behave the same way, but Java at least throws IllegalStateException on reuse. Python silently gives you nothing:
squares = (x * x for x in range(4)) # generator: an ITERATOR
print(sum(squares)) # -> 14print(sum(squares)) # -> 0 (already exhausted — no error!)
# Same trap with map/zip/filter objects:pairs = zip(["a", "b"], [1, 2])print(list(pairs)) # -> [('a', 1), ('b', 2)]print(list(pairs)) # -> []
# If you need multiple passes, materialise once:data = list(x * x for x in range(4))print(sum(data), max(data)) # -> 14 9Gotcha — This bites hardest in functions that iterate their argument twice (e.g. compute a mean, then a variance). Pass a list and it works; pass a generator and the second pass sees nothing, producing a silently wrong answer. Either document that you need a sequence, or call
list()at the top.
Generators: Iterators You Write Like Functions
Writing an iterator class by hand (__iter__ + __next__ + state fields) is tedious — exactly like implementing Iterator<T> in Java. A generator function does it for you: any function containing yield returns a generator object when called. Each next() runs the body until the next yield, then pauses with all local state preserved.
def countdown(n): print("starting") # runs on the FIRST next(), not when countdown() is called while n > 0: yield n # hand back a value and pause here n -= 1 print("done") # runs when the loop ends, just before StopIteration
gen = countdown(3) # nothing printed yet: the body hasn't startedprint(type(gen).__name__) # -> generatorprint(next(gen)) # prints "starting", then the value 3print(list(gen)) # prints "done" after consuming 2 and 1 -> [2, 1]Why it matters: memory
import sys
eager = [i * 2 for i in range(1_000_000)] # one million ints materialisedlazy = (i * 2 for i in range(1_000_000)) # a recipe for producing them
print(sys.getsizeof(eager) > 8_000_000) # -> True (~8 MB just for the pointers)print(sys.getsizeof(lazy) < 500) # -> True (a couple of hundred bytes, regardless of size)print(sum(lazy) == sum(eager)) # -> TrueInfinite sequences
Generators can be infinite, because they only compute what’s asked for:
from itertools import islice
def transaction_ids(prefix="TXN"): n = 1 while True: # never ends — that's fine yield f"{prefix}-{n:06d}" n += 1
print(list(islice(transaction_ids(), 3))) # -> ['TXN-000001', 'TXN-000002', 'TXN-000003']Deep Dive: Generator Pipelines
The real power appears when you chain generators into stages — each stage pulls items from the previous one, one at a time. It is a Java Stream pipeline built from plain functions, and it processes a 40 GB file with constant memory.
import csvfrom pathlib import Pathfrom collections import defaultdict
# --- create a small sample file so the example is runnable ------------------Path("txns.csv").write_text( "txn_id,merchant,amount,country\n" "T1,AcmeMart,120.0,IN\n" "T2,ZipFuel,not_a_number,IN\n" "T3,AcmeMart,15000.0,US\n" "T4,BookNook,80.0,IN\n" "T5,ZipFuel,22000.0,SG\n", encoding="utf-8",)
# --- Stage 1: source. Yields dict rows lazily. ------------------------------def read_rows(path): with open(path, newline="", encoding="utf-8") as f: # file stays open while iterating yield from csv.DictReader(f) # delegate to another iterator
# --- Stage 2: parse + clean. Drops bad rows instead of crashing. ------------def parse(rows): for row in rows: try: row["amount"] = float(row["amount"]) except ValueError: print(f"skipping bad row {row['txn_id']}") continue yield row
# --- Stage 3: filter. -------------------------------------------------------def foreign_only(rows, home="IN"): return (r for r in rows if r["country"] != home) # generator expression stage
# --- Stage 4: terminal operation (the 'collect'). ---------------------------pipeline = foreign_only(parse(read_rows("txns.csv"))) # NOTHING has executed yet
totals = defaultdict(float)for row in pipeline: # pulling drives every stage totals[row["merchant"]] += row["amount"]
print(dict(totals))# skipping bad row T2# {'AcmeMart': 15000.0, 'ZipFuel': 22000.0}terminal] T -.->|next| FO FO -.->|next| P P -.->|next| R
Each next() request travels up the chain; each item travels down. Exactly one row is in flight at a time.
Tip —
yield from iterabledelegates to a sub-iterator: it yields every item ofiterablein turn. It replacesfor x in iterable: yield xand is how you compose generators.
itertools: The Stream Operators You Were Missing
from itertools import islice, chain, batched, accumulate, takewhile, groupby, pairwise
amounts = [120, 45, 3000, 80, 15000, 60]
print(list(islice(amounts, 2, 5))) # -> [3000, 80, 15000] (skip 2, limit 3)print(list(chain([1, 2], (3, 4), range(5, 7)))) # -> [1, 2, 3, 4, 5, 6] (concat)print(list(batched(amounts, 4))) # -> [(120, 45, 3000, 80), (15000, 60)] (3.12+)print(list(accumulate(amounts))) # -> [120, 165, 3165, 3245, 18245, 18305] (running total)print(list(takewhile(lambda a: a < 1000, amounts))) # -> [120, 45]print(list(pairwise([10, 15, 12]))) # -> [(10, 15), (15, 12)] (sliding pairs)
# groupby groups CONSECUTIVE keys — sort first for SQL-like grouping.events = sorted([("IN", 10), ("US", 5), ("IN", 7)], key=lambda e: e[0])print({k: sum(v for _, v in grp) for k, grp in groupby(events, key=lambda e: e[0])}) # -> {'IN': 17, 'US': 5}Tip —
batchedis your micro-batcher. Sending rows to a model or an API in chunks of 512 is one line:for chunk in batched(rows, 512): score(chunk). On Python < 3.12, useislicein a loop.
Exceptions: Unchecked, Hierarchical, and Used for Control Flow
Python has no checked exceptions — every exception behaves like a RuntimeException. The structure is familiar, with two additions: else and exception chaining with from.
def parse_amount(raw: str) -> float: try: value = float(raw) except ValueError as e: # 'raise ... from e' chains the cause (like new X(msg, cause) in Java). raise InvalidTransaction(f"bad amount {raw!r}") from e else: # runs only if the try block raised NOTHING — keeps the try block minimal if value < 0: raise InvalidTransaction("negative amount") return value finally: pass # always runs: cleanup goes here (but prefer 'with' — see below)
class SentinelError(Exception): # project base exception """Base class for all Sentinel errors."""
class InvalidTransaction(SentinelError): # specific subtype pass
for raw in ["120.5", "abc", "-3"]: try: print(parse_amount(raw)) except InvalidTransaction as e: cause = type(e.__cause__).__name__ if e.__cause__ else None print(f"rejected: {e} (cause: {cause})")# 120.5# rejected: bad amount 'abc' (cause: ValueError)# rejected: negative amount (cause: None)EAFP vs LBYL
Java culture is mostly LBYL — Look Before You Leap: check map.containsKey(k) then map.get(k). Python culture prefers EAFP — Easier to Ask Forgiveness than Permission: just try it and handle the exception. Exceptions are cheap enough in Python, and EAFP avoids race conditions between the check and the action.
config = {"threshold": "0.8"}
# LBYL (works, but two lookups and a check-then-act gap)if "threshold" in config: threshold = float(config["threshold"])
# EAFP (idiomatic)try: timeout = float(config["timeout"])except KeyError: timeout = 30.0print(threshold, timeout) # -> 0.8 30.0Gotcha — never write a bare
except:. It catches everything, includingKeyboardInterrupt(Ctrl-C) andSystemExit. Evenexcept Exception:should be rare and should log the error. Catch the narrowest exception that you can actually handle.
Tip — Common built-in exceptions map neatly:
ValueError≈IllegalArgumentException,TypeError≈ClassCastException,KeyError/IndexError≈NoSuchElementException/IndexOutOfBoundsException,AttributeError≈NullPointerException(usually you called a method onNone),NotImplementedError≈UnsupportedOperationException.
Context Managers: with Is try-with-resources
with guarantees cleanup. Any object with __enter__ and __exit__ works, exactly as any AutoCloseable works in try-with-resources.
from pathlib import Path
path = Path("scores.txt")
# File automatically closed when the block exits — even if an exception is raised.with path.open("w", encoding="utf-8") as f: f.write("T1,0.91\nT2,0.12\n")
print(f.closed) # -> True
# Multiple resources in one statement (parenthesised form, 3.10+):with ( open("scores.txt", encoding="utf-8") as src, open("high.txt", "w", encoding="utf-8") as dst,): for line in src: if float(line.split(",")[1]) > 0.5: dst.write(line)
print(Path("high.txt").read_text().strip()) # -> T1,0.91Writing your own: the class way and the generator way
import timefrom contextlib import contextmanager
class Timer: """Class-based context manager: __enter__ / __exit__."""
def __enter__(self): self.start = time.perf_counter() return self # bound to the 'as' target
def __exit__(self, exc_type, exc, tb): self.elapsed = time.perf_counter() - self.start return False # False = don't swallow exceptions
@contextmanagerdef model_mode(state: dict, mode: str): """Generator-based: code before 'yield' is __enter__, after it is __exit__.""" previous = state["mode"] state["mode"] = mode try: yield state # the body of the 'with' block runs here finally: state["mode"] = previous # restored even if the block raised
with Timer() as t: sum(range(100_000))print(t.elapsed > 0) # -> True
model = {"mode": "train"}with model_mode(model, "eval"): print(model["mode"]) # -> evalprint(model["mode"]) # -> trainThat second pattern — temporarily switch a mode and guarantee it’s restored — is exactly what torch.no_grad() does in Part 10.
from contextlib import suppress
# suppress: "ignore this specific exception" — cleaner than try/except/pass.cache = {}with suppress(KeyError): del cache["missing"]print("still running") # -> still runningFiles the Modern Way: pathlib, csv, json
import csvimport jsonfrom pathlib import Path
data_dir = Path("data") / "raw" # '/' joins paths — OS-independentdata_dir.mkdir(parents=True, exist_ok=True)
# JSON: dicts and lists map directly to objects and arrays.config = {"model": "logreg", "threshold": 0.8, "features": ["amount", "hour"]}cfg_path = data_dir / "config.json"cfg_path.write_text(json.dumps(config, indent=2), encoding="utf-8")loaded = json.loads(cfg_path.read_text(encoding="utf-8"))print(loaded["features"]) # -> ['amount', 'hour']
# CSV writing with DictWriter (newline="" avoids blank lines on Windows).rows = [{"txn_id": "T1", "amount": 120.0}, {"txn_id": "T2", "amount": 45.5}]csv_path = data_dir / "txns.csv"with csv_path.open("w", newline="", encoding="utf-8") as f: writer = csv.DictWriter(f, fieldnames=["txn_id", "amount"]) writer.writeheader() writer.writerows(rows)
print(sorted(p.name for p in data_dir.glob("*.*"))) # -> ['config.json', 'txns.csv']print(csv_path.suffix, csv_path.stem) # -> .csv txnsTip — always pass
encoding="utf-8". Python’s default file encoding is platform-dependent (historically cp1252 on Windows), which breaks on merchant names like “Café”. Python 3.15 makes UTF-8 the default everywhere; until you’re on it, be explicit.
Tip — For real tabular data you’ll use
pandas.read_csv(path, chunksize=100_000), which gives you an iterator of DataFrames — the same lazy idea, vectorised (Part 8). Plaincsvremains useful for quick scripts and for streaming transforms.
Tips, Tricks & Gotchas
Tip —
enumerate,zip,map,filter,reversed,dict.items()are lazy too. They return iterators or views, not lists. Wrap them inlist()when you need to print or reuse them.
Gotcha — generators delay errors. A bug inside a generator doesn’t surface when you create the pipeline, only when something consumes it. If a pipeline “does nothing”, check that something actually iterates it.
Gotcha — closing over files in generators. If you return a generator from inside a
with open(...)block in a regular function, the file closes when the function returns, and iteration fails withValueError: I/O operation on closed file. Put thewithinside the generator function, asread_rowsdoes above.
Tip —
sum,min,max,any,all,sorted,"".join,dict(),set()all accept any iterable, so you can feed them a generator expression directly without building a list.
Key Takeaways
| Concept | Remember |
|---|---|
| Iterator protocol | iter() gives an iterator; next() until StopIteration |
| One-shot iterators | Generators, files, zip, map exhaust silently — materialise if reused |
| Generators | yield pauses the function; constant memory; can be infinite |
| Pipelines | Chain generator stages like Stream operations; yield from to delegate |
itertools | islice, chain, batched, accumulate, groupby (sort first) |
| Exceptions | All unchecked; raise ... from for causes; else for the success path |
| EAFP | Try and handle, rather than check then act |
| Context managers | with = try-with-resources; write with a class or @contextmanager |
Story Closing
The rewritten loader streamed the 40 GB export through a four-stage generator pipeline in constant memory, logging and skipping the 0.3% of rows with corrupt amounts. Arjun’s laptop fan barely noticed.
It wasn’t fast, though. Forty minutes for a single pass. And when Arjun opened the shared features.py module to plug his loader in, he hit a different kind of problem: a function called build(data, cfg, mode=None) with no documentation and no types. What was data? A list? A DataFrame? What keys did cfg need?
He missed his compiler.
In Part 6, Arjun brings types back — type hints, mypy, Pydantic — and sets up the tooling that makes Python feel as safe as Maven and JUnit.
This is Part 5 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”