"Types Strike Back" — Type Hints, Pydantic and Modern Tooling
Arjun misses his compiler. Type hints and what they do (and do not) do at runtime, generics, TypedDict, mypy, runtime validation with Pydantic, project layout with pyproject.toml, ruff, pytest and logging — the Maven-and-JUnit toolkit for Python.
Story Opening
The function in features.py looked like this:
def build(data, cfg, mode=None): ...No types. No docstring. Arjun spent forty minutes reading call sites to discover that data was a list of dicts, cfg needed a "window_days" key (an int — except one caller passed the string "7", which “worked” until it didn’t), and mode could be "train", "serve" or None, with None meaning "train".
In Java, the signature alone would have told him most of that, and javac would have rejected the string.
“Python has types now,” Priya said. “We just never enforced them in this repo. Let’s fix that.”
Modern Python is a gradually typed language. You can annotate as much or as little as you like, a static checker verifies the annotations, and libraries like Pydantic use the very same annotations to validate data at runtime. Combined with a handful of fast tools, you get most of the safety net you’re used to — without giving up the REPL.
Java → Python: The Tooling Map
| Java ecosystem | Python ecosystem | Role |
|---|---|---|
javac type checking | mypy or pyright | Static type checking |
| Bean Validation / Jackson | Pydantic | Runtime validation and (de)serialisation |
| Maven / Gradle | uv + pyproject.toml | Dependencies, builds, environments |
| Checkstyle, Spotless, PMD | ruff | Linting and formatting (very fast) |
| JUnit 5 | pytest | Testing |
| Mockito | unittest.mock | Mocking |
| JaCoCo | coverage / pytest-cov | Coverage |
| SLF4J + Logback | logging | Logging |
| Javadoc | Docstrings (+ MkDocs / Sphinx) | API documentation |
Deep Dive: Type Hints Are Not Enforced at Runtime
This is the critical thing to understand: annotations are metadata. The interpreter stores them and otherwise ignores them. Nothing stops a wrong type from flowing through — until a static checker (or a library like Pydantic) reads the annotations.
def window_sum(amounts: list[float], window_days: int) -> float: return sum(amounts[-window_days:])
# CPython happily runs a call that violates the hints...try: window_sum([1.0, 2.0, 3.0], "2") # str instead of intexcept TypeError as e: print("runtime failure:", e) # -> runtime failure: bad operand type for unary -: 'str'
# ...and the annotations are just data you can inspect.print(window_sum.__annotations__["window_days"]) # -> <class 'int'>The error came from deep inside the function, not at the call site. A static checker catches it before the code runs:
$ uv run mypy features.pyfeatures.py:6: error: Argument 2 to "window_sum" has incompatible type "str"; expected "int" [arg-type]Found 1 error in 1 file (checked 1 source file)So the mental model is: types are checked in CI and in your editor, not by the interpreter. Run mypy (or pyright, which powers VS Code’s Pylance) as part of the build, the way javac is part of yours.
Annotation Syntax: A Field Guide
from collections.abc import Callable, Iterable, Mapping, Sequencefrom typing import Any, Final, Literal
# Built-in generics use the plain built-in names (3.9+).amounts: list[float] = [120.0, 45.5]by_merchant: dict[str, list[float]] = {}point: tuple[float, float] = (19.07, 72.87) # fixed sizeids: tuple[str, ...] = ("T1", "T2", "T3") # variable length, homogeneous
# Optional / union: 'X | None' (3.10+). Older code: Optional[X], Union[X, Y].def find_merchant(merchant_id: str) -> str | None: return {"M1": "AcmeMart"}.get(merchant_id)
# Literal: a closed set of values — a lightweight enum for parameters.Mode = Literal["train", "serve"]
def build(rows: Sequence[Mapping[str, Any]], window_days: int, mode: Mode = "train") -> list[float]: # Accept ABSTRACT types (Sequence, Mapping, Iterable), return CONCRETE ones (list, dict). # Same advice as "accept List, not ArrayList" in Java — but even more general. return [float(r["amount"]) for r in rows][-window_days:]
# Callables: Callable[[ArgTypes...], ReturnType] ≈ Function<T, R>Rule = Callable[[dict], bool]
def apply_rules(txn: dict, rules: Iterable[Rule]) -> list[str]: return [r.__name__ for r in rules if r(txn)]
MAX_AMOUNT: Final = 1_000_000 # Final: type checker forbids reassignment
print(build([{"amount": "10"}, {"amount": 20}], window_days=1)) # -> [20.0]print(find_merchant("M9")) # -> NoneTip —
Anyvsobject.Anyswitches type checking off for that value (anything goes in both directions).objectis the top type: anything can be passed in, but you can’t call methods on it without narrowing first. Preferobjectfor “I accept anything”; useAnyas a last resort.
Generics (3.12+ syntax)
Python 3.12 introduced a concise generics syntax that will feel natural coming from Java:
from collections.abc import Callablefrom dataclasses import dataclass
# Generic function: <T> T first(List<T> items)def first[T](items: list[T]) -> T: return items[0]
# Generic class: class Page<T> { List<T> items; int total; }@dataclassclass Page[T]: items: list[T] total: int
def map[R](self, fn: Callable[[T], R]) -> "Page[R]": # quotes: Page is not defined yet return Page([fn(x) for x in self.items], self.total)
# Bounded type variables: <N extends Number>def largest[N: (int, float)](values: list[N]) -> N: return max(values)
# Type aliases with the 'type' statementtype FeatureVector = list[float]
page = Page(items=["T1", "T2"], total=2)print(first(page.items)) # -> T1print(page.map(len)) # -> Page(items=[2, 2], total=2)print(largest([3, 9, 4])) # -> 9You’ll still see the pre-3.12 style everywhere — T = TypeVar("T") and class Page(Generic[T]). It means the same thing.
Gotcha — no type erasure surprises, but no reified generics either.
list[int]andlist[str]are the same class at runtime, just like Java.isinstance(x, list[int])raisesTypeError; checkisinstance(x, list)and inspect elements if you must.
TypedDict: typing the dict-shaped data you can’t change
Lots of real data arrives as dicts — JSON payloads, CSV rows, library return values. TypedDict describes their shape for the checker without changing the runtime object (it’s still a plain dict).
from typing import NotRequired, TypedDict
class TxnRow(TypedDict): txn_id: str amount: float merchant: str mcc: NotRequired[str] # key may be absent
def risk(row: TxnRow) -> float: return min(row["amount"] / 10_000, 1.0) # editor autocompletes keys; typos are errors
row: TxnRow = {"txn_id": "T1", "amount": 2_500.0, "merchant": "AcmeMart"}print(risk(row), type(row).__name__) # -> 0.25 dictNarrowing
Type checkers understand control flow. After an isinstance or is None check, the type is narrowed — like Java’s pattern matching for instanceof.
def normalise(value: str | float | None) -> float: if value is None: return 0.0 # here: None if isinstance(value, str): return float(value.replace(",", "")) # here: str return value # here: float
print(normalise("1,250.5"), normalise(None), normalise(3.0)) # -> 1250.5 0.0 3.0Deep Dive: Pydantic — Types That Validate at Runtime
Static checking protects you from your own code. It can’t protect you from a JSON payload, a YAML config, a CSV row or an LLM’s output. For data crossing a boundary you need runtime validation — and Pydantic does it using ordinary type hints.
Pydantic (v2, with a Rust core) is the de facto standard: it powers FastAPI, LangChain, the OpenAI and Anthropic SDKs, and countless ML config systems. Think Jackson + Bean Validation in one class definition.
from datetime import datetimefrom typing import Literal
from pydantic import BaseModel, Field, ValidationError, field_validator
class Transaction(BaseModel): txn_id: str = Field(min_length=3, pattern=r"^T\d+$") amount: float = Field(gt=0, le=1_000_000) currency: Literal["INR", "USD", "EUR"] = "INR" merchant: str created_at: datetime tags: list[str] = [] # Pydantic copies defaults per instance — safe here
@field_validator("merchant") @classmethod def strip_merchant(cls, v: str) -> str: return v.strip().title()
# Parsing COERCES compatible types: "120.5" -> 120.5, ISO string -> datetime.raw = { "txn_id": "T1001", "amount": "120.5", "merchant": " acme mart ", "created_at": "2026-10-02T09:15:00",}t = Transaction.model_validate(raw)print(t.amount, t.merchant, t.created_at.hour) # -> 120.5 Acme Mart 9print(t.model_dump_json())# -> {"txn_id":"T1001","amount":120.5,"currency":"INR","merchant":"Acme Mart","created_at":"2026-10-02T09:15:00","tags":[]}
# Invalid input produces ONE exception listing EVERY problem, with locations.bad = {"txn_id": "X1", "amount": -5, "currency": "GBP", "created_at": "yesterday"}try: Transaction.model_validate(bad)except ValidationError as e: print(e.error_count()) # -> 5 for err in e.errors(): print(err["loc"], "-", err["type"])# ('txn_id',) - string_too_short# ('amount',) - greater_than# ('currency',) - literal_error# ('merchant',) - missing# ('created_at',) - datetime_from_date_parsing| Feature | @dataclass | Pydantic BaseModel |
|---|---|---|
Generated __init__, repr, eq | Yes | Yes |
| Runtime type validation | No | Yes |
Type coercion ("1" → 1) | No | Yes (strict mode available) |
| JSON (de)serialisation | Manual | model_validate_json / model_dump_json |
| JSON Schema generation | No | model_json_schema() |
| Speed / overhead | Minimal | Small, but non-zero |
| Use for | Internal data structures | Data at boundaries: APIs, configs, files, LLM outputs |
Tip — LLM structured output.
Transaction.model_json_schema()produces a JSON Schema you can hand to an LLM API as a response format, thenmodel_validate_json()the reply. This “schema in, validated object out” loop is the backbone of reliable LLM applications (Part 10).
Project Layout and pyproject.toml
A Python project that a Java developer would recognise:
sentinel/├── pyproject.toml # ≈ pom.xml: metadata, dependencies, tool config├── uv.lock # locked dependency tree — commit it├── src/│ └── sentinel/ # the importable package (≈ src/main/java/com/ledgerline/sentinel)│ ├── __init__.py│ ├── features.py│ └── models.py└── tests/ # ≈ src/test/java ├── conftest.py # shared pytest fixtures └── test_features.pyThe src layout forces tests to import the installed package rather than whatever happens to be in the current directory — it prevents a whole class of “works on my machine” import bugs.
[project]name = "sentinel"version = "0.1.0"requires-python = ">=3.12"dependencies = [ "numpy>=2.0", "pandas>=2.2", "pydantic>=2.7", "scikit-learn>=1.5",]
[dependency-groups]dev = ["mypy", "pytest", "pytest-cov", "ruff"]
[build-system]requires = ["hatchling"]build-backend = "hatchling.build"
[tool.ruff]line-length = 100
[tool.ruff.lint]select = ["E", "F", "I", "B", "UP", "SIM"] # pycodestyle, pyflakes, isort, bugbear, pyupgrade, simplify
[tool.mypy]strict = true # the closest thing to javac's strictnessplugins = ["pydantic.mypy"]
[tool.pytest.ini_options]testpaths = ["tests"]addopts = "-q --strict-markers"The everyday commands, which map neatly onto a Maven lifecycle:
uv sync # mvn dependency:resolveuv run ruff format . # spotless:applyuv run ruff check . --fix # checkstyle / PMD, with autofixuv run mypy src # compile-time type checkinguv run pytest --cov=sentinel # mvn test + JaCoCouv build # mvn package -> wheel (.whl ≈ .jar) and sdistTip — Start with
strict = trueon new code and use per-module overrides ([[tool.mypy.overrides]]) to relax it for legacy modules. Retrofitting strictness later is far more painful.
Testing with pytest
pytest is to JUnit what Python is to Java: less ceremony, more power. Tests are plain functions; assertions are plain assert statements (pytest rewrites them to show detailed diffs on failure); fixtures replace @BeforeEach through dependency injection by parameter name.
import pytest
# --- code under test (would normally be: from sentinel.features import ...) ---def velocity(amounts: list[float], window: int) -> float: """Average of the last `window` amounts.""" if window <= 0: raise ValueError("window must be positive") recent = amounts[-window:] return sum(recent) / len(recent)
# --- fixtures: dependency-injected by parameter name (≈ @BeforeEach + fields) ---@pytest.fixturedef history() -> list[float]: return [100.0, 200.0, 300.0, 400.0]
def test_velocity_uses_last_window(history): assert velocity(history, 2) == 350.0
def test_velocity_float_comparison(history): # pytest.approx: tolerance-based float equality (≈ assertEquals(a, b, delta)) assert velocity([0.1, 0.2], 2) == pytest.approx(0.15)
# --- parametrize: one test, many cases (≈ @ParameterizedTest + @CsvSource) ---@pytest.mark.parametrize( ("window", "expected"), [(1, 400.0), (3, 300.0), (10, 250.0)], # window larger than history uses everything)def test_velocity_windows(history, window, expected): assert velocity(history, window) == expected
# --- expected exceptions (≈ assertThrows) ---def test_velocity_rejects_bad_window(history): with pytest.raises(ValueError, match="positive"): velocity(history, 0)$ uv run pytest -q...... [100%]6 passed in 0.02sTip — fixtures compose. A fixture can request other fixtures, can have
scope="session"(≈@BeforeAll), and canyielda value to run teardown code afterwards. Put shared fixtures intests/conftest.pyand they’re available everywhere without imports.
Tip — mocking:
unittest.mock.patch("sentinel.features.fetch_rates")replaces a name where it is looked up, for the duration of a test. The classic mistake is patching where a function is defined instead of where it’s imported.
Logging Instead of print
import logging
# Configure ONCE, at application entry point (≈ logback.xml).logging.basicConfig( level=logging.INFO, format="%(asctime)s %(levelname)s %(name)s - %(message)s",)
# One logger per module, named after the module (≈ LoggerFactory.getLogger(X.class)).log = logging.getLogger(__name__)
def score(txn_id: str, amount: float) -> float: log.debug("scoring %s", txn_id) # lazy %-formatting: skipped when DEBUG is off risk = min(amount / 10_000, 1.0) if risk > 0.8: log.warning("high risk txn=%s risk=%.2f", txn_id, risk) return risk
score("T1", 9_500) # logs: ... WARNING __main__ - high risk txn=T1 risk=0.95Gotcha — Libraries should never call
logging.basicConfig(); only applications should. And preferlog.info("x=%s", x)over f-strings in log calls, so formatting is skipped when the level is disabled — the same reasoning as SLF4J’s{}placeholders.
Tips, Tricks & Gotchas
Tip — Forward references: a class can’t name itself in annotations during its own definition on Python ≤ 3.13 without quotes (
"Page[R]") orfrom __future__ import annotationsat the top of the module. Python 3.14 evaluates annotations lazily by default (PEP 649/749), so the quotes become unnecessary.
Tip —
reveal_type(x)inside code makes mypy print the inferred type ofx— the fastest way to debug confusing type errors. Remove it before running the code (on Python 3.11+ it’s also importable fromtyping, where it prints the runtime type).
Gotcha — untyped libraries. Some packages ship without type information. mypy then treats them as
Any. Look for stub packages (pandas-stubs,types-requests) or useignore_missing_importsper module.
Gotcha — data science notebooks are rarely typed. That’s normal: exploration and production have different needs. The pattern that works is “explore in notebooks, then move stable code into a typed, tested package that notebooks import.”
Key Takeaways
| Concept | Remember |
|---|---|
| Type hints | Metadata only — enforced by mypy/pyright, not the interpreter |
| Syntax | list[int], X | None, Literal, Callable, Final |
| Generics | def f[T](x: T) -> T and class Box[T] (3.12+); TypeVar in older code |
| Abstract inputs | Accept Sequence/Mapping/Iterable, return concrete types |
TypedDict | Types for dict-shaped data without changing the runtime object |
| Pydantic | Runtime validation and coercion at boundaries; JSON and JSON Schema |
| Tooling | uv + ruff + mypy + pytest ≈ Maven + Checkstyle + javac + JUnit |
| Logging | Module-level logging.getLogger(__name__); configure once |
Story Closing
Within a week, features.py had types, a Pydantic FeatureConfig that rejected "7" with a precise error message, and forty pytest tests running in under a second in CI. Arjun even caught a real bug: a function annotated to return float that returned None on one branch. mypy found it in three seconds.
But the forty-minute pipeline still bothered him. He profiled it. Ninety-four percent of the time was spent in one innocent-looking function — a pure-Python loop computing a z-score for every transaction amount against the customer’s history.
Priya looked at the loop and shook her head. “Every iteration goes through the interpreter. You need to stop thinking in elements and start thinking in arrays.”
In Part 7, Arjun meets NumPy — the foundation of the entire scientific Python stack — and makes the pipeline two orders of magnitude faster.
This is Part 6 of a 10-part series: “Python for Java Developers: From Streams to Tensors.”