Story Opening

This was Arjun’s Transaction class — forty-five lines of honest Java instinct:

class Transaction:
def __init__(self, txn_id, amount):
self.__txn_id = txn_id
self.__amount = amount
def getTxnId(self):
return self.__txn_id
def getAmount(self):
return self.__amount
def setAmount(self, amount):
if amount < 0:
raise ValueError("negative amount")
self.__amount = amount
# ...plus equals, hashCode and toString equivalents, written by hand
t = Transaction("T1", 120.0)
print(t) # prints something like <__main__.Transaction object at 0x10e3c2d50>

Priya replaced it with this:

from dataclasses import dataclass
@dataclass(frozen=True)
class Transaction:
txn_id: str
amount: float
t = Transaction("T1", 120.0)
print(t) # -> Transaction(txn_id='T1', amount=120.0)
print(t == Transaction("T1", 120.0)) # -> True

“Constructor, repr, equality, hashing, immutability,” she said. “Five lines. And in Python we don’t write getters until we actually need one.”

Python is deeply object-oriented — every value is an object — but its culture is the opposite of enterprise Java: start simple, keep attributes public, and add machinery only when a real need appears.


Java → Python: The Quick Map

JavaPython
Constructor__init__(self, ...)
this (implicit)self (explicit first parameter)
Fields declared in class bodyAttributes assigned in __init__ (or declared in a dataclass)
private / getters / settersPublic attributes; _internal convention; @property when needed
toString()__repr__ (developers) and __str__ (end users)
equals() / hashCode()__eq__ / __hash__
Comparable.compareTo__lt__ (+ functools.total_ordering)
record@dataclass(frozen=True)
static method@staticmethod; alternative constructors use @classmethod
interfacetyping.Protocol (structural) or abc.ABC (nominal)
enumenum.Enum / StrEnum
Operator overloadingNot in Java! Python: __add__, __getitem__, …

Classes: The Basics

class Account:
# Class attribute — shared by ALL instances (like a static field).
currency = "INR"
def __init__(self, account_id: str, balance: float = 0.0):
# Instance attributes are created by assignment. There is no field declaration.
self.account_id = account_id
self.balance = balance
def deposit(self, amount: float) -> None:
# 'self' is passed explicitly — acc.deposit(5) is sugar for Account.deposit(acc, 5)
self.balance += amount
def __repr__(self) -> str:
return f"Account({self.account_id!r}, balance={self.balance})"
acc = Account("ACC-1")
acc.deposit(500)
print(acc) # -> Account('ACC-1', balance=500.0)
print(Account.deposit(acc, 1) is None, acc.balance) # -> True 501.0
print(acc.currency) # -> INR (found on the class, not the instance)

Gotcha — mutable class attributes are shared. A tags = [] in the class body is one list shared by every instance — the same trap as mutable default arguments (Part 3). Initialise mutable state in __init__.

class Bad:
flags = [] # ONE list, attached to the class
a, b = Bad(), Bad()
a.flags.append("FRAUD")
print(b.flags) # -> ['FRAUD']

Deep Dive: Why Python Doesn’t Need Getters and Setters

In Java, you write getters from day one because changing a public field to a method later breaks every caller. Python solves that problem differently: @property lets you turn an attribute into a method without changing the calling syntax. So you start with a plain public attribute and upgrade it only when you need validation or computation.

class Transaction:
def __init__(self, txn_id: str, amount: float):
self.txn_id = txn_id
self.amount = amount # goes through the property setter below!
@property
def amount(self) -> float: # read: txn.amount
return self._amount
@amount.setter
def amount(self, value: float) -> None: # write: txn.amount = 42
if value < 0:
raise ValueError(f"negative amount: {value}")
self._amount = value
@property
def amount_paise(self) -> int: # computed, read-only property
return round(self._amount * 100)
t = Transaction("T1", 120.5)
print(t.amount, t.amount_paise) # -> 120.5 12050 (no parentheses — looks like a field)
try:
t.amount = -5
except ValueError as e:
print(e) # -> negative amount: -5
try:
t.amount_paise = 1
except AttributeError as e:
print("read-only:", type(e).__name__) # -> read-only: AttributeError

Callers wrote t.amount before validation existed and still write t.amount after. No API break, so no reason for defensive getters.

What about private?

  • _name (single underscore): “internal, please don’t touch”. Not enforced, but linters, IDEs and from module import * respect it.
  • __name (double underscore, no trailing underscores): triggers name mangling to _ClassName__name. Its purpose is avoiding accidental clashes in subclasses, not security. Arjun’s self.__amount was legal but un-idiomatic.
class Vault:
def __init__(self):
self.__secret = 42
v = Vault()
print(hasattr(v, "__secret")) # -> False
print(v._Vault__secret) # -> 42 (mangled, not hidden)

Tip — Python’s motto here is “we’re all consenting adults.” Use _internal names to communicate intent and rely on code review, not the compiler.


Deep Dive: The Data Model — Dunder Methods

“Dunder” = double underscore. These special methods are hooks the interpreter calls for built-in syntax. Implement them and your objects behave like native types: len(x) calls x.__len__(), x[i] calls x.__getitem__(i), a + b calls a.__add__(b), for calls __iter__, and so on.

This isn’t a curiosity. It is exactly how NumPy makes array * 2 multiply every element and how pandas makes df[df.amount > 100] filter rows. Understanding dunders makes those libraries stop feeling like magic.

from functools import total_ordering
@total_ordering # derive <=, >, >= from __eq__ and __lt__
class Money:
def __init__(self, amount: float, currency: str = "INR"):
self.amount = amount
self.currency = currency
# repr: unambiguous, ideally valid Python. Used by the REPL, debuggers, containers.
def __repr__(self) -> str:
return f"Money({self.amount!r}, {self.currency!r})"
# str: friendly display for end users. Falls back to __repr__ if not defined.
def __str__(self) -> str:
return f"{self.currency} {self.amount:,.2f}"
# equality — must agree with __hash__ (same contract as Java's equals/hashCode)
def __eq__(self, other: object) -> bool:
if not isinstance(other, Money):
return NotImplemented # lets Python try other.__eq__ / fall back sensibly
return (self.amount, self.currency) == (other.amount, other.currency)
def __hash__(self) -> int:
return hash((self.amount, self.currency))
def __lt__(self, other: "Money") -> bool:
self._check(other)
return self.amount < other.amount
# arithmetic operator overloading
def __add__(self, other: "Money") -> "Money":
self._check(other)
return Money(self.amount + other.amount, self.currency)
def __bool__(self) -> bool: # truthiness: zero money is falsy
return self.amount != 0
def _check(self, other: "Money") -> None:
if self.currency != other.currency:
raise ValueError("currency mismatch")
a, b = Money(100), Money(250.5)
print(a + b) # -> INR 350.50
print(repr(a + b)) # -> Money(350.5, 'INR')
print(a < b, a >= b) # -> True False
print(a == Money(100), {a, Money(100)} == {a}) # -> True True
print(bool(Money(0))) # -> False
print(sorted([b, a])) # -> [Money(100, 'INR'), Money(250.5, 'INR')]

Gotcha — defining __eq__ removes __hash__. If a class defines __eq__ but not __hash__, Python sets __hash__ = None, making instances unhashable (can’t go in sets or be dict keys). Java lets you silently break the contract; Python refuses. Define both, or use a frozen dataclass.

Making a container

class TransactionBatch:
"""A sequence of transaction amounts that behaves like a built-in list."""
def __init__(self, amounts):
self._amounts = list(amounts)
def __len__(self): # len(batch)
return len(self._amounts)
def __getitem__(self, index): # batch[0], batch[-1], batch[1:3]
if isinstance(index, slice):
return TransactionBatch(self._amounts[index])
return self._amounts[index]
def __iter__(self): # for amt in batch
return iter(self._amounts)
def __contains__(self, amount): # 120.0 in batch
return amount in self._amounts
def __repr__(self):
return f"TransactionBatch({self._amounts})"
batch = TransactionBatch([120.0, 45.0, 300.0, 80.0])
print(len(batch), batch[-1], 45.0 in batch) # -> 4 80.0 True
print(batch[1:3]) # -> TransactionBatch([45.0, 300.0])
print(sum(batch), max(batch)) # -> 545.0 300.0 (built-ins just work)
SyntaxDunder called
repr(x), str(x)__repr__, __str__
x == y, x < y__eq__, __lt__ (and friends)
hash(x)__hash__
len(x), bool(x)__len__, __bool__
x[k], x[k] = v__getitem__, __setitem__
for i in x, y in x__iter__, __contains__
x + y, x * y, x @ y__add__, __mul__, __matmul__
x(...)__call__
with x:__enter__, __exit__ (Part 5)

Dataclasses: Records, and Then Some

@dataclass generates __init__, __repr__ and __eq__ from type-annotated class attributes. Options add ordering, immutability, hashing and memory savings.

from dataclasses import dataclass, field, asdict, replace
from datetime import datetime
@dataclass(frozen=True, slots=True) # immutable + no per-instance __dict__ (less memory)
class Transaction:
txn_id: str
amount: float
merchant: str
currency: str = "INR" # fields with defaults must come after those without
tags: tuple[str, ...] = () # immutable default is safe
created_at: datetime = field(default_factory=datetime.now) # computed per instance
raw: dict = field(default_factory=dict, repr=False, compare=False)
def __post_init__(self): # validation hook, runs after generated __init__
if self.amount < 0:
raise ValueError("amount must be non-negative")
t1 = Transaction("T1", 120.0, "AcmeMart", created_at=datetime(2026, 10, 2))
print(t1)
# -> Transaction(txn_id='T1', amount=120.0, merchant='AcmeMart', currency='INR', tags=(), created_at=datetime.datetime(2026, 10, 2, 0, 0))
# Frozen: "modify" by creating a changed copy (like a record 'wither').
t2 = replace(t1, amount=99.0)
print(t2.amount, t1.amount) # -> 99.0 120.0
print(asdict(t1)["merchant"]) # -> AcmeMart (to dict — handy for JSON)
print(t1 in {t1}) # -> True (frozen + eq => hashable)
try:
t1.amount = 1
except Exception as e:
print(type(e).__name__) # -> FrozenInstanceError

Gotcha — A mutable default like tags: list = [] raises ValueError: mutable default ... use default_factory at class definition time. Dataclasses actively protect you from the Part 3 trap.

NeedUse
Simple immutable record@dataclass(frozen=True) or NamedTuple
Mutable data holder@dataclass
Validation of untrusted input (JSON, APIs, configs)Pydantic (Part 6)
Millions of rowsNot objects at all — a pandas DataFrame (Part 8)

Class Methods, Static Methods and Alternative Constructors

Without overloading, Python uses @classmethod factories for alternative constructors — like static factory methods (Optional.of, LocalDate.parse) in Java.

from dataclasses import dataclass
@dataclass
class Merchant:
merchant_id: str
name: str
mcc: str
@classmethod
def from_csv_row(cls, row: str) -> "Merchant":
# 'cls' is the class itself — subclasses calling this get a subclass instance.
merchant_id, name, mcc = row.split(",")
return cls(merchant_id.strip(), name.strip(), mcc.strip())
@staticmethod
def is_valid_mcc(mcc: str) -> bool: # no self/cls: just a namespaced function
return len(mcc) == 4 and mcc.isdigit()
m = Merchant.from_csv_row("M1, AcmeMart, 5411")
print(m) # -> Merchant(merchant_id='M1', name='AcmeMart', mcc='5411')
print(Merchant.is_valid_mcc("54A1")) # -> False

Tip — Module-level functions are often better than @staticmethod. Python doesn’t force everything into a class; a module already is a namespace.


Inheritance and Abstract Base Classes

from abc import ABC, abstractmethod
class Scorer(ABC): # cannot be instantiated: like an abstract class
def __init__(self, name: str):
self.name = name
@abstractmethod
def score(self, txn: dict) -> float: ...
def describe(self) -> str: # concrete template method
return f"{self.name} ({type(self).__name__})"
class RuleScorer(Scorer):
def __init__(self, name: str, limit: float):
super().__init__(name) # call the parent constructor explicitly
self.limit = limit
def score(self, txn: dict) -> float:
return 1.0 if txn["amount"] > self.limit else 0.0
s = RuleScorer("big-ticket", 5_000)
print(s.describe(), s.score({"amount": 9_000})) # -> big-ticket (RuleScorer) 1.0
print(isinstance(s, Scorer)) # -> True
try:
Scorer("abstract")
except TypeError as e:
print("cannot instantiate:", "abstract" in str(e)) # -> cannot instantiate: True

Python supports multiple inheritance, resolved via the C3 method resolution order (ClassName.__mro__). In practice it’s used for small mixins — classes that add one capability (e.g. JsonMixin). Keep hierarchies shallow; prefer composition, as you would in Java.


Protocols: Duck Typing With a Contract

“If it walks like a duck and quacks like a duck…” Python code traditionally doesn’t check types at all — it just calls the method and lets it fail if missing. typing.Protocol (3.8+) adds a contract for type checkers without requiring inheritance: any class with matching methods satisfies it. It’s structural typing, like Go interfaces or TypeScript.

This is the key to understanding the ML ecosystem: scikit-learn never asks your model to extend a base class — anything with fit and predict works in its pipelines.

from typing import Protocol, runtime_checkable
@runtime_checkable # also allow isinstance() checks at runtime
class Model(Protocol):
def predict(self, features: list[float]) -> float: ...
class ThresholdModel: # does NOT inherit from Model
def predict(self, features: list[float]) -> float:
return 1.0 if sum(features) > 1.0 else 0.0
class MeanModel:
def predict(self, features: list[float]) -> float:
return sum(features) / len(features)
def batch_predict(model: Model, rows: list[list[float]]) -> list[float]:
# A type checker verifies the argument has a compatible predict(); nothing else needed.
return [model.predict(r) for r in rows]
rows = [[0.2, 0.3], [0.9, 0.8]]
print(batch_predict(ThresholdModel(), rows)) # -> [0.0, 1.0]
print(batch_predict(MeanModel(), rows)) # -> [0.25, 0.8500000000000001]
print(isinstance(MeanModel(), Model)) # -> True
UseWhen
ProtocolDescribing what you need from an argument; third-party classes you can’t modify
ABCYou own the hierarchy and want shared implementation plus enforced overrides

Enums

from enum import Enum, StrEnum, auto
class Decision(StrEnum): # 3.11+: members ARE strings — great for JSON and DataFrames
APPROVE = auto() # auto() on a StrEnum gives the lowercase member name
REVIEW = auto()
DECLINE = auto()
class Channel(Enum):
CARD = 1
UPI = 2
d = Decision.REVIEW
print(d, d == "review") # -> review True
print(Decision("decline").name) # -> DECLINE
print([c.name for c in Channel]) # -> ['CARD', 'UPI']
print(Channel.UPI.value) # -> 2

Tips, Tricks & Gotchas

Tip — Always write __repr__ (or use a dataclass). You’ll spend a lot of time looking at objects in notebooks and debuggers; <object at 0x...> helps nobody.

Tip — vars(obj) and obj.__dict__ show an instance’s attributes as a dict — great for debugging objects from unfamiliar libraries.

Gotcha — attributes can be added anywhere. acc.balnce = 10 (typo) silently creates a new attribute. slots=True on a dataclass (or __slots__ on a class) turns that into an AttributeError, and type checkers catch it too.

Gotcha — super() in multiple inheritance follows the MRO, not “the parent”. Always call super().__init__(...) so cooperative mixins work.

Tip — Don’t build deep class hierarchies for data. In data work, records live in DataFrames and behaviour lives in functions. Classes earn their place for models, pipelines, clients and resources.


Key Takeaways

ConceptRemember
selfExplicit first parameter; obj.m() is Class.m(obj)
EncapsulationPublic by default; _internal by convention; @property when needed
DundersHooks for built-in syntax — how NumPy and pandas overload operators
__eq__ / __hash__Same contract as Java; defining __eq__ alone makes objects unhashable
DataclassesGenerated init/repr/eq; frozen, slots, default_factory, __post_init__
Factories@classmethod alternative constructors replace overloading
InterfacesProtocol for structural typing; ABC for nominal hierarchies
EnumsStrEnum for string-valued categories

Story Closing

Arjun’s Transaction became a frozen, slotted dataclass. His Money class got __add__ and __lt__, and for the first time he wrote sorted(payments) without a comparator. It felt like cheating.

Then the real data arrived. The fraud team’s historical export was a 40 GB CSV — three years of transactions. Arjun wrote rows = open("txns.csv").readlines() and watched his laptop’s memory graph climb vertically until the kernel killed Python.

“Don’t load it,” said Priya. “Stream it.”

In Part 5, Arjun learns iterators, generators and context managers — Python’s lazy, memory-safe answer to Java Streams and try-with-resources.


This is Part 4 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”