"Collections Without Ceremony" — Lists, Dicts, Sets and Comprehensions
The Streams API, translated. Lists, tuples, dicts and sets; slicing and unpacking; comprehensions as the idiomatic map/filter/collect; sorting with keys; and the collections module — Counter, defaultdict, deque — applied to transaction data.
Story Opening
Priya’s notebook was the densest code Arjun had ever read. One line computed the total spend per merchant for flagged transactions. In Java that would have been a stream().filter().collect(groupingBy(..., summingDouble(...))) chain spread over six lines and two static imports.
txns = [ {"id": "T1", "merchant": "AcmeMart", "amount": 120.0, "flagged": True}, {"id": "T2", "merchant": "AcmeMart", "amount": 80.0, "flagged": False}, {"id": "T3", "merchant": "ZipFuel", "amount": 45.0, "flagged": True},]
flagged_merchants = {t["merchant"] for t in txns if t["flagged"]}print(sorted(flagged_merchants)) # -> ['AcmeMart', 'ZipFuel']“It’s not clever,” Priya said. “It’s just the normal way to write it. Python’s built-in collections are so central that the language gives them their own syntax.”
This part covers the four core containers, the syntax that makes them pleasant, and the collections module that fills in the rest. Everything here is pre-NumPy: plain Python is what you use for configuration, metadata, small lookups and glue code — which is a lot of ML code.
Java → Python: The Quick Map
| Java | Python | Literal | Mutable? |
|---|---|---|---|
ArrayList<T> | list | [1, 2, 3] | Yes |
Immutable List.of(...) / record | tuple | (1, 2, 3) | No |
HashMap<K, V> / LinkedHashMap | dict (insertion-ordered) | {"a": 1} | Yes |
HashSet<T> | set | {1, 2, 3} | Yes |
Set.of(...) | frozenset | frozenset({1, 2}) | No |
ArrayDeque<T> | collections.deque | — | Yes |
groupingBy(..., counting()) | collections.Counter | — | Yes |
computeIfAbsent(k, ArrayList::new) | collections.defaultdict(list) | — | Yes |
stream().map().filter().collect() | Comprehensions | [f(x) for x in xs if p(x)] | — |
Lists: The Workhorse
A list is a dynamic array of object references — ArrayList<Object> with batteries. It can hold mixed types, but in practice you keep them homogeneous.
amounts = [120.0, 80.0, 45.0]
amounts.append(300.0) # add to end (ArrayList.add)amounts.insert(0, 10.0) # insert at index (add(index, e))amounts.extend([5.0, 7.5]) # add many (addAll)print(amounts) # -> [10.0, 120.0, 80.0, 45.0, 300.0, 5.0, 7.5]
last = amounts.pop() # remove & return last (O(1))first = amounts.pop(0) # remove at index (O(n) — shifts elements)print(first, last) # -> 10.0 7.5
amounts.remove(45.0) # remove first matching VALUE (ValueError if missing)print(80.0 in amounts) # -> True (contains — O(n) for lists)print(amounts.index(300.0)) # -> 2print(len(amounts), min(amounts), max(amounts), sum(amounts)) # -> 4 5.0 300.0 505.0Negative indices and slicing
Slicing is one of Python’s best features and it carries straight into NumPy and pandas, so learn it well. The syntax is seq[start:stop:step] — stop is exclusive, like subList(from, to).
hours = [0, 1, 2, 3, 4, 5, 6, 7, 8, 9]
print(hours[-1]) # -> 9 (last element — no hours.get(size()-1))print(hours[-3:]) # -> [7, 8, 9] (last three)print(hours[2:5]) # -> [2, 3, 4] (index 2 up to, not including, 5)print(hours[:3]) # -> [0, 1, 2] (omitted start = 0)print(hours[::2]) # -> [0, 2, 4, 6, 8] (every second element)print(hours[::-1]) # -> [9, 8, 7, 6, 5, 4, 3, 2, 1, 0] (reversed copy)
# Slices return NEW lists (shallow copies) — unlike Java's subList view.window = hours[0:3]window[0] = 99print(hours[0]) # -> 0
# Slice assignment replaces a range in place.hours[0:3] = ["night"]print(hours[:3]) # -> ['night', 3, 4]
# Out-of-range slices never throw; out-of-range indexes do.print(hours[100:]) # -> []Tuples: Lightweight Immutable Records
Tuples are fixed-size, immutable sequences. Use them for small, heterogeneous groupings — “this function returns a pair” — where Java would need a record or Map.Entry.
point = (19.07, 72.87) # lat, lonsingle = (42,) # one-element tuple needs a trailing comma!not_a_tuple = (42) # just the int 42 in parentheses
print(type(single).__name__, type(not_a_tuple).__name__) # -> tuple int
# Tuples are hashable (if their contents are), so they can be dict keys.# Java would need a composite key class with equals/hashCode.route_risk = {("IN", "US"): 0.4, ("IN", "IN"): 0.05}print(route_risk[("IN", "US")]) # -> 0.4
try: point[0] = 0.0except TypeError as e: print(e) # -> 'tuple' object does not support item assignmentUnpacking — destructuring everywhere
lat, lon = (19.07, 72.87) # tuple unpackingprint(lat) # -> 19.07
def min_max(values): return min(values), max(values) # returns a tuple — no wrapper class
lo, hi = min_max([5, 3, 9])print(lo, hi) # -> 3 9
# Star-unpacking collects the rest (like varargs, in reverse).first, *middle, last = [1, 2, 3, 4, 5]print(first, middle, last) # -> 1 [2, 3, 4] 5
# '_' is the conventional "I don't care" name._, second, *_ = ["a", "b", "c", "d"]print(second) # -> b
# Unpacking in for-loops — very common with pairs.for merchant, total in [("AcmeMart", 200.0), ("ZipFuel", 45.0)]: print(f"{merchant}={total}")For readable records, use typing.NamedTuple (or a dataclass — Part 4):
from typing import NamedTuple
class Txn(NamedTuple): id: str amount: float currency: str = "INR"
t = Txn("T9", 250.0)print(t.amount, t[1]) # -> 250.0 250.0 (attribute AND index access)print(t) # -> Txn(id='T9', amount=250.0, currency='INR')txn_id, amount, ccy = t # still unpacks like a tupleDicts: The Most Important Data Structure in Python
Dicts are hash maps that preserve insertion order (guaranteed since 3.7) — think LinkedHashMap by default. Python itself is built on them: module namespaces, object attributes and keyword arguments are all dicts.
txn = {"id": "T1", "amount": 120.0, "merchant": "AcmeMart"}
print(txn["amount"]) # -> 120.0txn["country"] = "IN" # putprint("country" in txn) # -> True (containsKey — O(1))
# txn["mcc"] would raise KeyError. Use .get() for optional keys.print(txn.get("mcc")) # -> Noneprint(txn.get("mcc", "5411")) # -> 5411 (getOrDefault)
# setdefault = putIfAbsent that also returns the value.txn.setdefault("tags", []).append("first_purchase")print(txn["tags"]) # -> ['first_purchase']
del txn["country"] # remove (KeyError if missing)removed = txn.pop("tags", None) # remove with default — never throws
# Iteration: keys by default; .items() for entrySet()for key, value in txn.items(): print(f"{key} -> {value}")
# Merging (3.9+): right side wins on conflicts.defaults = {"currency": "INR", "channel": "card"}merged = defaults | {"channel": "upi"}print(merged) # -> {'currency': 'INR', 'channel': 'upi'}
# Build from pairsprint(dict(zip(["a", "b"], [1, 2]))) # -> {'a': 1, 'b': 2}Gotcha — Dict keys must be hashable — effectively, immutable. Strings, numbers and tuples of those work. Lists and dicts don’t:
{[1, 2]: "x"}raisesTypeError: unhashable type: 'list'. Convert to a tuple first.
Sets
seen_devices = {"dev-a", "dev-b"}new_devices = {"dev-b", "dev-c"}
print(seen_devices & new_devices) # -> {'dev-b'} (retainAll / intersection)print(sorted(seen_devices | new_devices)) # -> ['dev-a', 'dev-b', 'dev-c'] (union)print(new_devices - seen_devices) # -> {'dev-c'} (removeAll / difference)print(sorted(seen_devices ^ new_devices)) # -> ['dev-a', 'dev-c'] (symmetric difference)
empty = set() # NOT {} — that's an empty dict!print(type({}).__name__) # -> dict
# Fast dedup of a list (order lost); dict.fromkeys keeps first-seen order.ids = ["T3", "T1", "T3", "T2", "T1"]print(list(dict.fromkeys(ids))) # -> ['T3', 'T1', 'T2']Deep Dive: Comprehensions Are Your Streams API
A comprehension builds a new collection from an iterable in a single expression. The shape is always:
[ <expression> for <name> in <iterable> if <condition> ] └─ map ─┘ └──── source ────┘ └─ filter ─┘Here is the same computation in Java and Python, side by side:
// Java: amounts (in rupees) of large flagged transactions, converted to paiseList<Long> paise = txns.stream() .filter(t -> t.flagged() && t.amount() > 100) .map(t -> Math.round(t.amount() * 100)) .collect(Collectors.toList());txns = [ {"id": "T1", "amount": 120.0, "flagged": True}, {"id": "T2", "amount": 80.0, "flagged": True}, {"id": "T3", "amount": 450.0, "flagged": False}, {"id": "T4", "amount": 999.99, "flagged": True},]
paise = [round(t["amount"] * 100) for t in txns if t["flagged"] and t["amount"] > 100]print(paise) # -> [12000, 99999]Four flavours
txns = [("T1", "AcmeMart", 120.0), ("T2", "ZipFuel", 45.0), ("T3", "AcmeMart", 80.0)]
# 1. List comprehension -> Collectors.toList()amounts = [amt for _, _, amt in txns]print(amounts) # -> [120.0, 45.0, 80.0]
# 2. Set comprehension -> Collectors.toSet()merchants = {m for _, m, _ in txns}print(sorted(merchants)) # -> ['AcmeMart', 'ZipFuel']
# 3. Dict comprehension -> Collectors.toMap()by_id = {tid: amt for tid, _, amt in txns}print(by_id) # -> {'T1': 120.0, 'T2': 45.0, 'T3': 80.0}
# 4. Generator expression (lazy!) -> an un-terminated Streamtotal = sum(amt for _, _, amt in txns) # no brackets needed inside a callprint(total) # -> 245.0Conditional expressions vs filters
The if at the end filters. An if/else at the front maps conditionally. Mixing them up is a classic first-week mistake:
amounts = [50, 150, 900, 20]
# Filter: keep only large onesprint([a for a in amounts if a > 100]) # -> [150, 900]
# Map: label every one (if/else goes BEFORE 'for')print(["high" if a > 100 else "low" for a in amounts]) # -> ['low', 'high', 'high', 'low']Nested loops — read left to right
merchants = ["AcmeMart", "ZipFuel"]channels = ["card", "upi"]
# Equivalent to: for m in merchants: for c in channels: ...combos = [(m, c) for m in merchants for c in channels]print(combos) # -> [('AcmeMart', 'card'), ('AcmeMart', 'upi'), ('ZipFuel', 'card'), ('ZipFuel', 'upi')]
# Flattening a list of lists (flatMap):batches = [[1, 2], [3], [4, 5]]print([x for batch in batches for x in batch]) # -> [1, 2, 3, 4, 5]
# Building a 2-D grid correctly (each row a NEW list):grid = [[0] * 3 for _ in range(2)]grid[0][0] = 1print(grid) # -> [[1, 0, 0], [0, 0, 0]]When not to use a comprehension
- Side effects.
[print(x) for x in xs]builds a useless list. Use aforloop. - More than two
fors or a complex condition. Readability beats cleverness; extract a function or use a loop. - Big numeric data. A comprehension over a million floats is a Python-level loop. That is NumPy’s job (Part 7) — 50–100× faster.
Sorting: key Functions Instead of Comparators
Java uses Comparator.comparing(...).thenComparing(...). Python uses a key function that maps each element to a sort key; tuples compare element by element, which gives you multi-level sorting for free. Python’s sort (Timsort — also what Java uses for objects) is stable.
txns = [ {"id": "T1", "merchant": "ZipFuel", "amount": 45.0}, {"id": "T2", "merchant": "AcmeMart", "amount": 120.0}, {"id": "T3", "merchant": "AcmeMart", "amount": 300.0},]
# sorted() returns a NEW list; list.sort() sorts in place and returns None.by_amount = sorted(txns, key=lambda t: t["amount"], reverse=True)print([t["id"] for t in by_amount]) # -> ['T3', 'T2', 'T1']
# Multi-key: merchant ascending, then amount descending (negate numbers).multi = sorted(txns, key=lambda t: (t["merchant"], -t["amount"]))print([t["id"] for t in multi]) # -> ['T3', 'T2', 'T1']
# operator.itemgetter is a faster, cleaner key for dict/tuple fields.from operator import itemgetterprint([t["id"] for t in sorted(txns, key=itemgetter("merchant", "amount"))]) # -> ['T2', 'T3', 'T1']
# Top-N without sorting everythingimport heapqprint([t["id"] for t in heapq.nlargest(2, txns, key=itemgetter("amount"))]) # -> ['T3', 'T2']Gotcha —
result = my_list.sort()setsresulttoNone. In-place methods returnNoneby convention so you can’t accidentally chain them.
Built-ins That Replace Loops
merchants = ["AcmeMart", "ZipFuel", "BookNook"]amounts = [120.0, 45.0, 80.0]
# enumerate: index + value (no manual counter)for i, m in enumerate(merchants, start=1): print(i, m)
# zip: walk sequences in parallel (stops at the shortest; strict=True to enforce equal length)pairs = list(zip(merchants, amounts, strict=True))print(pairs[0]) # -> ('AcmeMart', 120.0)
# any / all: short-circuiting anyMatch / allMatchprint(any(a > 100 for a in amounts)) # -> Trueprint(all(a > 10 for a in amounts)) # -> True
# min/max with key — returns the ELEMENT, not the keyprint(max(zip(merchants, amounts), key=lambda p: p[1])) # -> ('AcmeMart', 120.0)
# reversed and rangeprint(list(reversed(merchants))) # -> ['BookNook', 'ZipFuel', 'AcmeMart']print(list(range(0, 10, 3))) # -> [0, 3, 6, 9]The collections Module
Counter — frequency counting
from collections import Counter
decline_codes = ["51", "05", "51", "14", "51", "05"]counts = Counter(decline_codes)
print(counts) # -> Counter({'51': 3, '05': 2, '14': 1})print(counts["51"]) # -> 3print(counts["99"]) # -> 0 (missing keys count as zero — no KeyError)print(counts.most_common(2)) # -> [('51', 3), ('05', 2)]
# Counters support arithmetic — handy for comparing distributions.yesterday = Counter({"51": 1, "05": 4})print(counts - yesterday) # -> Counter({'51': 2, '14': 1})defaultdict — grouping without boilerplate
from collections import defaultdict
txns = [("AcmeMart", 120.0), ("ZipFuel", 45.0), ("AcmeMart", 80.0)]
# Java: map.computeIfAbsent(merchant, k -> new ArrayList<>()).add(amount)by_merchant = defaultdict(list)for merchant, amount in txns: by_merchant[merchant].append(amount) # missing key -> list() created on demand
print(dict(by_merchant)) # -> {'AcmeMart': [120.0, 80.0], 'ZipFuel': [45.0]}
# Summing per keytotals = defaultdict(float)for merchant, amount in txns: totals[merchant] += amountprint(dict(totals)) # -> {'AcmeMart': 200.0, 'ZipFuel': 45.0}Gotcha — A
defaultdictcreates entries on read too.if by_merchant["Unknown"]:silently inserts"Unknown": []. Use"Unknown" in by_merchantfor membership checks.
deque — sliding windows
from collections import deque
# Keep the last 3 transaction amounts for a velocity check.recent = deque(maxlen=3) # oldest items fall off automaticallyfor amount in [10, 20, 30, 40, 50]: recent.append(amount)
print(list(recent)) # -> [30, 40, 50]print(sum(recent) / len(recent)) # -> 40.0
recent.appendleft(5) # O(1) at both ends (list.insert(0, x) is O(n))print(list(recent)) # -> [5, 30, 40]Tips, Tricks & Gotchas
Tip — Complexity cheat sheet:
x in listis O(n);x in setandk in dictare O(1). If you test membership in a loop, convert the list to a set first. This one change fixes a surprising number of slow scripts.
Tip —
itertools.groupbyis not SQLGROUP BY: it only groups consecutive equal keys. Sort by the key first, or just usedefaultdict. (For real grouping, pandas — Part 8.)
Gotcha — modifying a collection while iterating: removing from a list inside
for x in lstskips elements; modifying a dict’s size raisesRuntimeError. Iterate over a copy (for x in lst[:]) or build a new collection with a comprehension.
amounts = [10, 200, 300, 20]
# WRONG: skips elements because indices shift during removalbuggy = amounts.copy()for a in buggy: if a > 100: buggy.remove(a)print(buggy) # -> [10, 300, 20]
# RIGHT: build a new listprint([a for a in amounts if a <= 100]) # -> [10, 20]Tip — Use
jsonfor quick inspection:print(json.dumps(obj, indent=2))pretty-prints nested dicts and lists — Python’stoString()for data structures.pprint.pprintworks for non-JSON-serialisable objects.
Key Takeaways
| Concept | Remember |
|---|---|
list | Dynamic array; O(1) append/pop at end; O(n) membership |
tuple | Immutable, hashable record; use for multi-returns and composite keys |
dict | Insertion-ordered hash map; .get(), .items(), | to merge |
set | O(1) membership; &, |, - operators; set() not {} |
| Slicing | [start:stop:step], stop exclusive, negatives from the end, returns copies |
| Comprehensions | [expr for x in xs if cond] — the idiomatic map/filter/collect |
| Sorting | key= functions; tuple keys for multi-level; stable |
collections | Counter, defaultdict, deque replace most boilerplate |
Story Closing
By Friday, Arjun had rewritten one of his own Java-style helper modules. Ninety lines became thirty-one. Priya approved the PR with a single emoji, which he decided to interpret as respect.
But one thing still bothered him. Priya’s code passed functions around like candy — key=lambda t: ..., sorted(..., key=itemgetter(...)) — and one of her modules had an @retry(times=3) line sitting above a function definition. It looked a lot like a Spring annotation. It wasn’t.
In Part 3, Arjun discovers that in Python, functions are just objects — and that decorators are AOP you can write in ten lines.
This is Part 2 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”