This guide provides ideals for production-grade Python code, code that humans are likely to read, code that will likely be edited and/or re-executed at least once, and code that I will review in detail. Code whose correctness, reproducibility, and reliability matter should follow these guidelines.

This guide partly exists because Python is so flexible. There are enough ways of using Python to warrant a clean separation between “scripts that will be executed once, never reviewed by another person, or used again” from “services or ‘important’ tasks whose correctness matters and for which I might be called in to debug at midnight.”

Containers

Cont.1. Prefer immutable collections and data structures to mutable ones especially to represent “facts known at a particular instant”.

Mutable collections are great for accreting information over time. Though when possible, accrete that state in one scope so understanding “how did a collection get to be how it is” only requires local analysis and reasoning.

Generally we agree with P.10 of the C++ Core Guidelines and its reasoning.

Reason: It is easier to reason about constants than about variables. Something immutable cannot change unexpectedly. Sometimes immutability enables better optimization. You can’t have a data race on a constant.

Cont.2. Prefer NamedTuples and (frozen) dataclasses for most “plain-old-data” containers over bare classes

Reason: typing.NamedTuples and frozen dataclasses.dataclasses signal to experienced programmers that a “struct” is a “fact known at a particular point in time.” They can be more easily and reliably passed across async and picked across interprocess boundaries than many kinds of classes. Even unfrozen/“normal” dataclasses signal to experienced developers that “this container is mainly for holding data.” Bare/raw classes are better for state machines or crafting a Pythonic API with instance-scoped encapsulated implementation details. Ultimately an experienced developer can skim and fully understand a NamedTuple or (frozen) dataclass more quickly than a raw/bare class.

Example: A dataframe should be a class because it typically consists of a crafted “Pythonic” API over instance-scoped encapsulated/hidden “struct-of-arrays”-oriented numpy arrays or arrow record batches. A struct used to represent a validated collection of RPC request parameters that will be passed to a service might be better as a dataclass.

Cont.3. Use TypedDict for externally-enforced fixed-structure existing dictionaries.

Reason: typing.TypedDict was designed for adding type annotations/static analyzability for existing untyped code bases without sweeping structural changes.

PEP-589 that introduced typing.TypedDict states the following.

This PEP proposes a type constructor typing.TypedDict to support the use case where a dictionary object has a specific set of string keys, each with a value of a specific type…Dataclasses are a more recent alternative to solve this use case, but there is still a lot of existing code that was written before dataclasses became available, especially in large existing codebases where type hinting and checking has proven to be helpful. Unlike dictionary objects, dataclasses don’t directly support JSON serialization

typing.TypedDict is a more descriptive and less prescriptive kind of documentation than a typing.NamedTuple or (frozen) dataclasses.dataclass.

typing.TypedDict is extremely useful for adding static analysis to “legacy” applications where something outside the scope of the Python program creates the dictionary. For example, consider a proprietary service layer where highly tuned C++ written 30 years ago handles all schema enforcement and (de)serialization, and the choice was made long ago to expose these to Python using dictionaries. typing.TypedDict can be a useful code generation target because a system outside the scope of the Python program reliably handles schema enforcement and the decision to use Python dictionaries as the interface with the C++ layer is outside the scope of the Python program.

typing.TypedDict can also be useful for adding type annotations to an existing untyped code base that passes untyped dictionaries across function boundaries – or worse, module boundaries. Consider a function in such a code base that returns some dictionary where precise key names and value types can be inferred from one function or module scope. Changing to a dataclass or NamedTuple may prove difficult because one would need to change all call sites of the function – a rather dangerous endeavor without strong static analysis and tests. One can create a typing.TypedDict for the return type of the function and at least achieve some type inference and static analysis around the call sites within one commit/atomic change. Then after strengthening the test suite and type annotations for the code base (e.g. passing mypy strict mode), one can more confidently refactor to use dataclasses or NamedTuples.

Input-Output

IO.1. Validate untested/external data at runtime as close to the external boundary as possible

Parse HTTP payloads, queue messages, configuration files, environment variables, plugin inputs, and third-party responses into validated domain values as near to the external boundary as possible. Use pydantic or dataclass-based adapters, or explicit parsing depending on schema complexity and project dependencies. Static types typically do not validate runtime values. They complement, but do not replace, boundary validation.

Reason: Detecting unmet expectations early maximizes the chances of good error reporting with sufficient context for debugging. Avoid scenarios where a program uses external data that doesn’t meet expectations for much of the program and not knowing whether a KeyError or AttributeError is caused by garbage data or a programmer bug.

The industry has largely settled on pydantic for strict parsing of complex/nested data from outside the scope of the Python program, such as API responses, user/client-provided configurations and forms, and queue messages because it is fast and uses good idioms for avoiding possibly hundreds or thousands of lines of parsing code. Pydantic, as long as one parses external data as soon as possible, can help detect violations of contracts/expectations can be reliably detected and handled.

The key here is that data that is generated within the Python program should be able to use dataclasses and NamedTuples with static type checking to get the often formidable “enforcement” afforded by well-tuned static analysis and CI. There is generally no need to parse nested data structures entirely generated and consumed within the same Python program because type checkers detect expectation/contract violations at struct-creation time and access-time.

Exceptions: Small scripts that choose to be single-file and zero-dependency can parse hand-written, trusted, and controlled configuration files manually. For a few key-value pairs, especially ones that are mostly strings and integers, keep the parsing to a single function. pydantic makes more sense once the cost of packaging is already being paid but it is typically not worth it if bringing in pydantic creates the cost of packaging.

from dataclasses import dataclass
from pathlib import Path


@dataclass(frozen=True)
class _BackupRestoreValidationConfig: ...


def _parse_toml_config_file(config_path: Path) -> _BackupRestoreValidationConfig: ...

Control Flow

CF.1. Prefer context managers for scoped resource access and ‘computational contexts’ over try-finally

Reason: Context managers idiomatically enable setup-teardown patterns and are a reliable and simple way to encapsulate the process of releasing resources/tearing down execution context. This applies to many external resources, such as file handles, database connections and transactions, locks, and sockets/TCP connections.

Bad Example

import sqlite3
from dataclasses import dataclass


@dataclass(frozen=True)
class _DBConfig:  # stand in for something akin to a connection pool
    db_uri: str


def _get_username(user_id: int, config: _DBConfig) -> str | None:
    conn = sqlite3.connect(config.db_uri)
    cursor = conn.cursor()
    try:
        maybe_user_name = cursor.execute(
            "SELECT username FROM users WHERE user_id=?",
            (user_id,),
        ).fetchone()
        return maybe_user_name[0] if maybe_user_name else None
    finally:
        cursor.close()
        conn.close()

This is bad because there are situations where the connection and/or cursor will not be reliably closed. For example, if cursor.close() fails, the connection will leak. Moreover, failing to encapsulate the resource freeing in a context manager forces more functions to also implement the clean-up. The try-finally also complicates the swift closing/releasing of the connection.

Example

from contextlib import closing
import sqlite3
from dataclasses import dataclass


@dataclass(frozen=True)
class _DBConfig:
    db_uri: str


def _get_username(user_id: int, config: _DBConfig) -> str | None:
    with closing(sqlite3.connect(config.db_uri)) as conn:
        maybe_user_name = conn.execute(
            "SELECT username FROM users WHERE user_id=?",
            (user_id,),
        ).fetchone()
    return maybe_user_name[0] if maybe_user_name else None

A subtle difference here is that context managers help release the resources sooner, which can matter in certain demanding applications. While holding the connection for an extra ternary expression likely doesn’t matter, closing the connection early can matter with more demanding data processing.

Database client libraries, especially DBAPI-compatible ones, sometimes differ in their interpretation of context-manager usage. Some, such as sqlite3 and pyodbc, use the Connection.__exit__ method for transaction management and roll back transactions on errors, while using the Connection.close method (usually via with contextlib.closing(connect(....)) as conn) for closing the connection. Others, such as psycopg3, use the Connection.__exit__ method to roll back transactions on errors AND close the connection, making the invocation of the Connection.close method via contextlib.closing unnecessary. Connection pools, such as those provided by sqlalchemy have Connection.__exit__ return the connection back to the connection pool. Pay careful attention to how database libraries implement their context managers. Typically Cursor.__exit__ closes the cursor/stops receiving rows from the database and prevents further use, though this is underspecified by the dbapi2 spec, leaving implementations to decide the desired behavior of the cursor context manager.

Further, note that for long-running production services, the better recommendation is usually to use a database driver/pool whose lifecycle and transaction semantics are clearly understood, and manage it at application startup/shutdown rather than opening a new connection per request. One or a couple of isolated connection-per-SELECT queries is fine for scripts.

Bad Example

import json


def read_config_file(file_name: str) -> dict[str, str]:
    fob = open(file_name)
    try:
        return json.load(fob)
    finally:
        fob.close()

Example

import json


def read_config_file(file_name: str) -> dict[str, str]:
    with open(file_name) as fob:
        return json.load(fob)

The second block more concisely expresses and guarantees (under most, though not all scenarios) that the file handle must be closed. Especially when the ‘computational context’ extends across more statements, context managers often more reliably tear down the context/release the resource because the “setup” and “teardown” logic is adjacent – in the class definition __enter__ and __exit__ method or in a contextlib.contextmanager-decorated function – rather than separated by potentially many lines of code. This pattern applies to many kinds of computational context apart from just resource management.

Bad Example

import shutil
from logging import Logger
import time


def copy_directory_and_log_timing(src: str, dest: str, logger: Logger) -> None:
    start = time.monotonic()
    shutil.copytree(src, dest, dirs_exist_ok=True)
    end = time.monotonic()
    logger.info(
        "Copied %s to %s",
        src,
        dest,
        extra={"copy_duration_ms": int((end - start) * 1000)},
    )

Example

import shutil
from logging import Logger
from collections.abc import Iterator
from contextlib import contextmanager
import time


@contextmanager
def log_timing(
    message: str, logger: Logger, duration_ms_extra_key: str
) -> Iterator[None]:
    start = time.monotonic()
    yield
    end = time.monotonic()
    logger.info(
        message,
        extra={duration_ms_extra_key: int((end - start) * 1000)},
    )


def copy_directory_and_log_timing(src: str, dest: str, logger: Logger) -> None:
    with log_timing(
        f"Copied {src} to {dest}", logger, duration_ms_extra_key="copy_duration_ms"
    ):
        shutil.copytree(src, dest, dirs_exist_ok=True)

This change simplifies the process of instrumenting/logging timings of other operations.

With the contextlib.contextmanager decorator and its async sister contextlib.asynccontextmanager, we almost never need to write explicit __enter__ and __exit__ methods with their complex function signatures, as we have a very simple pattern for separating the setup and teardown with a yield statement.

Error Handling

Error states abound in production, especially for long-running processes. External services/dependencies can have downtime, performance regressions, and backwards-incompatible schema changes. Writing production-grade Python programs involves deliberately handling many failure modes.

  • For errors arising from a programmer bug or invalid application state (e.g. deadlock), there is no need to try-except because such situations should crash the program in many cases, or at least fail loudly.
  • Some errors can be handled within the “responsibility scope” of the function or module. For example, we might retry ephemerally failing HTTP/service requests with backoff (and possibly jitter), especially if we know that the API being requested is particularly flaky.
  • Some errors are better handled one or more layers up the call stack, in which case, we might use an errors-as-values pattern and propagate the stable/well-understood/controlled domain errors up the call stack via return values and let the caller (attempt to) recover or retry.

Do not continue when process-wide correctness is uncertain. Do not terminate healthy work unnecessarily when the failure is safely isolated.

In general, catch an exception only when the Python program can do one of the following:

  • recover safely
  • translate the exception into a stable domain error (especially as a value passed through the call stack via return values)
  • add essential context while preserving the cause
  • perform required cleanup not otherwise managed by a context manager

Err.1. Catch exceptions as specifically as possible

Reason: Robust programs consider at least one to three failure states/situations of each expression/procedure/function that is invoked. most-common failure cases. For each of the states and account for them

Bad Example

import json
from logging import Logger


def _read_val_from_config_file(file_name: str, logger: Logger) -> int | None:
    try:
        with open(file_name) as fob:
            data = json.load(fob)
        return int(data["some"]["key"]["that"]["may"]["not"]["exist"])
    except Exception:
        logger.exception("Something went wrong reading a value from %s", file_name)
        return None

Most exceptions are not very exceptional. In the above cases, most file names are not real files that exist on the current machine; a given user cannot open most files on a system; most files are not JSON, most JSON files do not have the used keys. The “exceptional” cases are actually common and should have explicit error messages so a user can easily address issues. Production-grade, reliable scripts and especially long-running services do not have the luxury of only considering the “happy-path” of a program.

Example

import json
from logging import Logger


def _read_val_from_config_file(file_name: str, logger: Logger) -> int | None:
    try:
        with open(file_name) as fob:
            data = json.load(fob)
    except json.JSONDecodeError:
        logger.exception("File %s is not valid JSON", file_name)
        return None
    except FileNotFoundError:
        logger.error("File %s not found", file_name)
        return None

    try:
        raw_extracted = data["some"]["key"]["that"]["may"]["not"]["exist"]
    except (KeyError, AttributeError):
        logger.exception("Failed to access data in JSON file %s", file_name)
        return None
    try:
        parsed = int(raw_extracted)
    except ValueError:
        logger.error(
            "Failed to parse value %s to integer in file %s", raw_extracted, file_name
        )
        return None
    return parsed

Exceptions: Bare except Exception can be reasonable very close to the entrypoint of a program to guarantee some sort of external behavior even in the presence of unhandled exceptions. Web frameworks frequently implement this in middleware to return a 500 HTTP status code and keep the server alive.

Err.2. Keep try-except blocks tight and surrounding as few failure-modes as possible

Reason try-except blocks that cover many statements complicate error attribution and the creation of helpful error messages. In a try-except-(finally)-block covering 5+ statements, many operations could throw, for example, a ValueError, which complicates the creation of a targeted, maximally useful error message.

Bad Example

In this example, consider a script that runs, for example, once per day and typically exits within 5 seconds.

# /// script
# requires-python = ">=3.11"
# dependencies = ["polars", "requests"]
# ///
from pathlib import Path
from datetime import date
import logging
import requests
import polars as pl


def etl_airline_data_to_parquet(
    base_url: str,
    output_dir: Path,
    publish_date: date,
    logger: logging.Logger,
) -> Path | None:
    """Download data from airline API to hive-style partitioned parquet for
    a date that has not been processed yet, returning the newly written parquet
    file if it was written, otherwise None if data was already processed/written"""
    try:
        output_pq_path = output_dir.joinpath(
            f"publish_date={publish_date.isoformat()}", "data.parquet"
        )
        if output_pq_path.is_file():
            return None
        resp_json = requests.get(
            base_url, data={"publish_date": publish_date.isoformat()}
        ).json()
        output_pq_path.parent.mkdir(exist_ok=True, parents=True)
        _schema = {"flight_id": pl.Int64, "passengers": pl.Int64, "revenue": pl.Int64}
        pl.LazyFrame(resp_json, schema=_schema).sink_parquet(output_pq_path)
        return output_pq_path
    except Exception as exc:
        print(f"Failed to download airline data {str(exc)} ")
        raise

The above error handling is simply lazy and has no place in production. Here are a few problems.

  • Type checking helps us guarantee that output_dir is actually a Path, in which case the joinpath cannot fail and should not be in a try block. The consequence is that our error message cannot include the file name/path we are attempting to write because output_pq_path might be unbound.
  • We cannot distinguish easily between IO errors creating the directory and writing the file without looking for string patterns in the error message. This is not very reliable, as error messages are not typically part of the public API/contract of packages. Exception types typically are more stable.

Bloated try-except blocks hinder useful error messages because such blocks typically leave the except blocks with too many possibly unbounded local variables.

Example

# /// script
# requires-python = ">=3.11"
# dependencies = ["polars", "requests"]
# ///
from pathlib import Path
from datetime import date
import requests
import polars as pl


def etl_airline_data_to_parquet(
    base_url: str,
    output_dir: Path,
    publish_date: date,
) -> Path | None:
    """Download data from airline API to hive-style partitioned parquet for
    a date that has not been processed yet, returning the newly written parquet
    file if it was written, otherwise None if data was already processed/written"""
    output_pq_path = output_dir.joinpath(
        f"publish_date={publish_date.isoformat()}", "data.parquet"
    )
    if output_pq_path.is_file():
        return None
    try:
        resp = requests.get(
            base_url,
            params={"publish_date": publish_date.isoformat()},
            timeout=(5.0, 5.0),
        )
    except requests.ConnectionError as err:
        _msg = f"Failed to connect to {base_url}. Check DNS resolution"
        raise RuntimeError(_msg) from err
    resp.raise_for_status()
    resp_json = resp.json()
    try:
        output_pq_path.parent.mkdir(exist_ok=True, parents=True)
    except OSError as err:
        _msg = f"""Failed to create directory for {output_pq_path.parent}.
        Check directory permissions"""
        raise OSError(_msg) from err
    _schema = {"flight_id": pl.Int64, "passengers": pl.Int64, "revenue": pl.Int64}
    try:
        dt_lzdf = pl.LazyFrame(resp_json, schema=_schema)
    except (ValueError, TypeError) as err:
        _msg = f"""Failed to coerce data to schema when processing airline for
        url={base_url} and publish_date={publish_date.isoformat()}. Did the
        alternative data team change their API again?"""
        raise ValueError(_msg) from err
    try:
        dt_lzdf.sink_parquet(output_pq_path)
    except OSError as err:
        _msg = f"""Failed to write parquet file {output_pq_path} . The
        filesystem might be full"""
        raise OSError(_msg) from err
    return output_pq_path

The above example does a better job of making each except block target a very specific case, rather than just “something went wrong.” In certain long-running services, this kind of structure opens the door to giving much better error messages and passing that information up the call stack to recover, possibly by returning an error value/enum variant everywhere this function raises.

Short of an errors-as-values approach, the purpose of re-raiseing is to include information that would be lost in the call stack. For example, the error message on requests.get would not include the query parameters, and so even if we log the exception stack trace, we would never see exactly which request gave the exception. Other error messages contain contextual information fundamentally not found in the code and that can be invaluable for whomever is called in to respond to this error state at 2:00AM.

Exceptions: On the other hand, any error state where the only correct action is to crash the program and where the exception message already contains the necessary information for debugging doesn’t need an explicit try-except.

Err.3. Consider errors-as-values

Iteration

It.1. Separate reused data generation from data collection with generator functions

Reason: Separation of “data generating logic” and collection into a data structure enables easier testing paths of both components and enables easier swapping of the target data structure, for example, from a list/tuple into a set/frozenset. Moreover collecting an iterator is frequently both faster and more memory efficient than repeated appends/inserts into a data structure for many reasons, including the following.

  • Different “consumers” of the data generating process can short-circuit/stop early, or choose to process the data iteratively or in batches.
  • Iterator collection pushes the actual iteration (repeated calls to __next__) to C rather than Python
  • Even invoking the .append or .add is a dict lookup in Python, while the insertion path is optimized in the C layer of CPython.
  • Iterator collection involves better size hints that reduce memory allocations when compared to calling .insert or .add many times.

Only break out an iterative process to the generator when doing so improves readability (if a nested loop warrants a comment, maybe a reified 3-4-word symbol/function name would improve readability), testability, or laziness. Many iterate+collect/accumulate/accrete to data structure operations – especially over small/well-controlled amounts of data – do not benefit from generator functions.

Example

from collections.abc import Iterator
from datetime import date, timedelta


def _saturdays_between(start: date, end: date) -> Iterator[date]:
    """Yield all saturdays between `start` and `end`, inclusive of both
    endpoints"""

    # .isoweekday() returns 6 for Saturday and 7 for Sunday
    days_until_saturday = (6 - start.isoweekday()) % 7
    first_saturday = start + timedelta(days=days_until_saturday)
    cursor = first_saturday
    while start <= cursor <= end:
        yield cursor
        cursor = cursor + timedelta(days=7)


def _saturdays_with_zero_traffic(
    *,
    start: date,
    end: date,
    n_requests_by_date: dict[date, int],
) -> tuple[date, ...]:
    """Returns dates that are saturday with no requests, assuming
    `n_requests_by_date` only contains entries for days with requests"""
    return tuple(
        d for d in _saturdays_between(start, end) if n_requests_by_date.get(d, 0) == 0
    )

Bad Example

from datetime import date, timedelta


def _saturdays_with_zero_traffic(
    *,
    start: date,
    end: date,
    n_requests_by_date: dict[date, int],
) -> tuple[date, ...]:
    res: list[date] = []
    days_until_saturday = (6 - start.isoweekday()) % 7
    first_saturday = start + timedelta(days=days_until_saturday)
    cursor = first_saturday
    while start <= cursor <= end:
        if n_requests_by_date.get(cursor, 0) == 0:
            res.append(cursor)
        cursor = cursor + timedelta(days=7)
    return tuple(res)

A simple inline comprehension expression also serves this purpose.

Exception: Certain algorithms may fit more naturally in one functional or logical scope.

Type Annotations

Type annotations and effective names for symbols (functions, classes, and variable names) remove the need for the vast majority of prose and structured/numpydoc-style documentation. When viewed as “statically analyzable documentation,” type annotations are well worth the verbosity, especially for programs whose reliability and maintainability. Changing a 1000+ line code base without any type annotations is a difficult endeavor, riddled with subtle traps. Especially past 5000 source lines of code, the local reasoning enabled by type annotations greatly aids maintainability. Programs that are strictly/strongly type annotated have a better path toward migration to a faster language, such as Golang because such programs facilitate local reasoning and make most details of the program explicit.

Effective type-checking involves the following.

  • Sparing usage of typing.Any, mainly near some external boundary
  • Typed function definitions in owned production code, especially function parameters and return types
  • Invoke type checking in CI

Type.1. All production grade Python programs that value reproducibility, reliability, and maintainability should use the strictest possible type annotations and incorporate a type checker – such as mypy – in its strictest mode in CI and developer workflows.

Reason: Type checkers catch a wide variety of bugs before they show up in production. Type annotations enable contributors to understand a section of code/function using “local reasoning” rather than needing to trace through 50 function calls to determine “what is in this object/dictionary.”

Example

def worst(*, val1, val2, val3): ...
def still_bad(*, val1: int, val2: dict, val3: str) -> str: ...
def good(*, val1: int, val2: dict[str, str], val3: str) -> str: ...

Type.2. All services/applications/libraries that use boto3 for Amazon Web Services (AWS) interactions should use the types-boto3 library or its async sister types-aioboto3 for type checking and analysis.

The ubiquity of AWS in many tech stacks and the complexity of the SDK warrant specific call outs. The AWS Python SDK parses schema definition files shared by all AWS language-specific SDKs and generates classes and functions at runtime. This invalidates most kinds of type inference that type-checkers such as mypy can perform. One of the few answers to such an approach is code-generating the type annotations from the same schema definition files, which is exactly the approach of these two packages.

Keep the stub version aligned with the deployed aioboto3 version. Prefer narrowly scoped extras such as types-aioboto3[s3,sqs] unless broad coverage is justified. Treat type stubs as development dependencies unless they are explicitly imported at runtime.

Example

# /// script
# requires-python = ">=3.11"
# dependencies = ["aioboto3", "types-aiobotocore[s3]", "types-aioboto3"]
# ///
from __future__ import annotations
from enum import unique, Enum
from typing import TYPE_CHECKING
import aioboto3

if TYPE_CHECKING:
    from types_aiobotocore_s3.client import S3Client


@unique
class _GetOldestVersionIDErrorReason(Enum): ...


async def _s3_get_oldest_version_id(
    bucket: str, key: str, s3_client: S3Client
) -> str | _GetOldestVersionIDErrorReason: ...


async def _near_script_entrypoint() -> None:
    sess = aioboto3.Session()
    async with sess.client("s3") as s3_client:
        version_id_or_err = await _s3_get_oldest_version_id(
            bucket="fake-bkt", key="fake-key", s3_client=s3_client
        )
        # Can reuse the client and session across async boundaries
        ...

Testing

TE.1. Use Pytest and write function-based pytest-style tests – avoid class-based unittest-style tests

Reason: Functions are flatter, simpler, and easier to compose. Pytest has a better story for test setup and teardown. In Python, we do not use classes as namespaces – we have modules for this purpose. unittest is useful within the standard library for testing Python implementations themselves without needing to make design decisions beyond “copy JUnit.” The decisions made by JUnit and the Python standard library’s unittest do not reflect the constraints of most properly packaged Python programs. pytest, is designed to allow for concise, to-the-point, maintainable tests that use the best of what Python has to offer, especially test context/fixtures via context managers. Moreover the pytest CLI offers extensive functionality, from surgical test selection, verbosity configuration, reporting configuration, and debugger integration. The pytest extension ecosystem is vast and formidable.

Example

from collections.abc import Iterator
from dataclasses import dataclass
import pytest


@dataclass(frozen=True)
class _TestDBContext: ...


@pytest.fixture
def seeded_empty_unauthenticated_database() -> Iterator[_TestDBContext]:
    with _seeded_empty_unauthenticated_database_impl() as test_db_context:
        yield test_db_context


def test_my_application_behavior(
    seeded_empty_unauthenticated_database: _TestDBContext,
) -> None: ...


def test_another_application_behavior(
    seeded_empty_unauthenticated_database: _TestDBContext,
) -> None: ...

Bad Example

class TestGrouping:
    def setUp(self) -> None: ...
    def tearDown(self) -> None: ...
    def test_my_application_behavior(self) -> None: ...
    def test_another_application_behavior(self) -> None: ...

Exceptions Only write class-based unittest-style tests if the existing code base is using unittest-style tests.

TE.2. Prefer dependency injection and ‘interface-compatible’ implementations over mocking.

“Dependency injection” most commonly involves giving functions more parameters.

Reason: Mocking stands in the way of fearless refactoring and changing a program – for both people and LLMs. Every mock represents untested application behavior. That can be fine for situations/scenarios that are difficult to set up/hermetically reproduce, especially those related to an external service or flakiness, but dependency injection at least tests the “happy path” of a program more thoroughly than mocking does.

Passing constants as function parameters almost always invalidates the need to mock them. Prefer to hoist the base URL of the HTTP service over mocking http clients. The tmp_path pytest fixture should remove the need to mock any filesystem operations.

Bad Example

# /// script
# requires-python = ">=3.11"
# dependencies = ["aiohttp", "pytest", "pytest-asyncio"]
# ///
"""Download information of top news stories from an API to the filesystem as
structured parquet"""

from datetime import date
from typing import NamedTuple
from pathlib import Path

import aiohttp

# app/library/cli/service module


class _NewsApiResponse(NamedTuple):
    article_ids: tuple[int, ...]


_OUTPUT_PARQUET_DIR = "/mnt/big_data_prod/data/reputable_news/top_stories/parquet"


async def _fetch_top_news_events_of_day(
    publish_date: date, top_k: int = 10
) -> _NewsApiResponse:
    params = {"publish_date": publish_date.isoformat(), "top_k": top_k}
    async with aiohttp.ClientSession() as session:
        async with session.get(
            "/api.reputable_news.xxyyzz/articles/top",
            params=params,
        ) as response:
            return _NewsApiResponse(tuple((await response.json())["article_ids"]))


async def etl_top_k_news_to_parquet(
    publish_date: date,
) -> None:
    records = await _fetch_top_news_events_of_day(publish_date)
    # write parquet to `_OUTPUT_PARQUET_DIR` with pyarrow
    ...


# Test module

from unittest.mock import patch
import pytest


@pytest.mark.asyncio
async def test_etl_top_k_news_to_parquet(tmp_path: Path) -> None:
    output_parquet_base_dir = tmp_path.joinpath("outpq")
    output_parquet_base_dir.mkdir()
    with (
        patch(
            "some.package._OUTPUT_PARQUET_DIR",
            str(output_parquet_base_dir.absolute()),
        ),
        patch("some.package.aiohttp.ClientSession", ...),
    ):
        await etl_top_k_news_to_parquet(date.today())
        assert ...  # data `some_other_directory_in_ci` matches expectations

There are many issues with the above test.

  • If we switch HTTP clients, our mock also needs to change.
  • The test doesn’t actually verify that we invoke the HTTP client properly – the mock hides any issues passing headers, data, auth, etc. into the HTTP client.
  • The test needs to know more about the implementation details than merely the ‘public’ interface.

We can fix all of these by hoisting some function parameters.

Example

# /// script
# requires-python = ">=3.11"
# dependencies = ["aiohttp", "pytest", "pytest-httpserver", "pytest-asyncio"]
# ///
from dataclasses import dataclass
from datetime import date
from typing import Final, NamedTuple
from pathlib import Path

import aiohttp

# app/library/cli/service


class _NewsApiResponse(NamedTuple):
    article_ids: tuple[int, ...]


_PROD_OUTPUT_PARQUET_DIR: Final[Path] = Path(
    "/mnt/big_data_prod/data/reputable_news/top_stories/parquet"
)


@dataclass(frozen=True)
class _TopArticlesRequestCtx:
    timeout: float
    session: aiohttp.ClientSession
    ...


@dataclass(frozen=True)
class _TopArticlesRequestParams:
    top_k: int
    publish_date: date
    ...

    def as_params(self) -> dict[str, str]:
        return {
            "top_k": str(self.top_k),
            "publish_date": self.publish_date.isoformat(),
            # ...
        }


async def _fetch_top_news_events_of_day(
    req_params: _TopArticlesRequestParams, ctx: _TopArticlesRequestCtx
) -> _NewsApiResponse:
    async with ctx.session.get(
        "/articles/top",
        params=req_params.as_params(),
        timeout=ctx.timeout,
    ) as response:
        response.raise_for_status()
        resp_json = await response.json()
    return _NewsApiResponse(tuple(resp_json["article_ids"]))


async def etl_top_k_news_to_parquet(
    req_params: _TopArticlesRequestParams,
    ctx: _TopArticlesRequestCtx,
    output_parquet_base_dir: Path = _PROD_OUTPUT_PARQUET_DIR,
) -> None:
    records = await _fetch_top_news_events_of_day(req_params, ctx)
    # write as parquet to `output_parquet_base_dir` using pyarrow
    ...


# Test module

from pytest_httpserver import HTTPServer
import pytest
from pytest_httpserver import RequestHandler
# from my_etl_script import _TopArticlesRequestCtx, _TopArticlesRequestParams


@pytest.mark.asyncio
async def test_etl_top_k_news_to_parquet(
    tmp_path: Path, httpserver: HTTPServer
) -> None:
    fixed_date = date.fromisoformat("2026-01-01")
    response_data = {"article_ids": list(range(10))}
    httpserver.expect_request("/articles/top", method="GET").respond_with_json(
        response_data
    )
    output_parquet_base_dir = tmp_path.joinpath("outpq")
    output_parquet_base_dir.mkdir()
    news_base_url = httpserver.url_for("/")
    async with aiohttp.ClientSession(base_url=news_base_url) as session:
        ctx = _TopArticlesRequestCtx(...)
        req_params = _TopArticlesRequestParams(...)
        await etl_top_k_news_to_parquet(req_params, ctx, output_parquet_base_dir)
    assert ...  # parquet data in `output_parquet_base_dir` matches expectations
    assert len(httpserver.log) == 1, "We can't flood the service with requests"

The key here is to design function/API boundaries that enable dependency injection, in this case for the API URL and the target output directory.

While spinning up and programming an HTTP server that “minimally” reproduces the target external service behavior may not enable testing all of the idiosyncracies of the service and arguably may not be too structurally different from mocking the HTTP Client, the “hitting an HTTP service that is compatible enough with the ‘real’ external service” strategy at least allows verification of the proper invocation of the HTTP client. Standing up an HTTP server just for tests also enables switching the HTTP client implementation without changing tests. In that sense, this strategy “mocks” a more fundamental/intrinsic boundary of the program: it mocks “the actual external service” rather than the HTTP client/response. Of course, we could also skip mocking entirely and just hit the live HTTP external service. This is better left to an integration test, however. If this HTTP service is internal to an enterprise, not all CI environments may be set up to route to the live service. Moreover, running mutating requests/RPC calls can have side effects on live services which may be undesirable, especially given that we expect that the program will have bugs until it passes CI with the rigorous test suite.

Exceptions: Certain scenarios encountered in production are difficult to reliably reproduce, especially in CI environments that do not permit containers. Examples include idiosyncratic SFTP configurations, or flakiness (e.g. ’the first 2 service calls fail due to server load or gateway issues, but the 3rd call succeeds with some backoff).

TE.3. Prefer to test edge-case behavior via the package’s public interface over private functions/implementation details.

“Public” interface can mean different things for different programs. For libraries, it means “the functions and invocations that callers will use.” For services (e.g. HTTP or GRPC), the exposed endpoints constitute the “public” interface. For an “offline” application such as program that drains a queue to a database, the guideline is less precise. While there is no “public interface,” there are core function boundaries as opposed to functions that are more implementation details.

Reason Tests of private implementation details stand in the way of fearless refactors. Good tests can guide the changing of implementation details.

Example

# /// script
# requires-python = ">=3.11"
# dependencies = ["fastapi", "httpx2", "pytest"]
# ///

# Pretend our app code is an another module
# from myapp import app
import pytest
from collections.abc import Iterator
import uuid

from fastapi.testclient import TestClient


@pytest.fixture
def app_client() -> Iterator[TestClient]:
    with TestClient(app) as c:
        yield c


@pytest.fixture
def admin_api_key() -> str: ...


@pytest.fixture
def username_password_of_registered_user_not_in_admin_group(
    app_client: TestClient,
    admin_api_key: str,
) -> tuple[str, str]:
    username = uuid.uuid4().hex[0:8]
    password = "fake_password"

    resp = app_client.post(
        "/api/v1/user",
        data={"groups": [], "username": username, "password": password},
        headers={"x-app-api-key": admin_api_key},
    )
    assert resp.status_code == 201
    return username, password


@pytest.fixture
def registered_job_id(
    app_client: TestClient,
    admin_api_key: str,
) -> int:
    username = uuid.uuid4().hex[0:8]
    password = "fake_password"

    resp = app_client.post(
        "/api/v1/job",
        json={"spec": {...}},
        headers={"x-app-api-key": admin_api_key},
    )
    assert resp.status_code == 201
    job_id = int(resp.json()["job_id"])
    return job_id


def test_user_not_in_admin_group_cant_stop_job(
    app_client: TestClient,
    username_password_of_registered_user_not_in_admin_group: tuple[str, str],
    registered_job_id: int,
) -> None:
    username, password = username_password_of_registered_user_not_in_admin_group
    resp = app_client.post(
        "/api/v1/login", data={"username": username, "password": password}
    )
    assert resp.status_code == 200
    cookies = resp.cookies
    resp = app_client.delete(f"/api/v1/job/{registered_job_id}", cookies=cookies)
    assert resp.status_code == 403

Exceptions: Functions whose execution context is difficult to reproduce in a testing/CI environment may be better served by targeted tests. Though they should have clearly outlined interface boundaries.

TE.4. Avoid conditional branching in test cases by splitting up each branch into different test cases

Reason: if-statements within test bodies can impede a reviewer’s efforts to determine exactly what a given test verifies. There are many situations where a conditionally-branched test may never take one of the branches. One can track this by computing code coverage specifically for tests. But avoiding branches guarantees that “every assert in the body of a properly constructed test function is executed and guarantees something about a program. Fundamentally, branchless test bodies increase the signal-to-noise ratio of any given test.

Bad Example

# /// script
# requires-python = ">=3.11"
# dependencies = ["pytest"]
# ///
import re
from typing import Final
# Application/service/library module

FILE_NAME_PATTERN: Final[re.Pattern[str]] = re.compile(
    r"(?P<iso3_country_code>[A-Z]{3})\-(?P<publish_year>\d{4})\.csv"
)

# Testing module
import pytest


@pytest.mark.parametrize(
    ("instr", "expected_captures"),
    [
        pytest.param(
            "CHN-2025.csv",
            {"iso3_country_code": "CHN", "publish_year": "2025"},
            id="CHN-2025.csv",
        ),
        pytest.param(
            "IND-1980.csv",
            {"iso3_country_code": "IND", "publish_year": "1980"},
            id="IND-1980.csv",
        ),
        pytest.param(
            "INDIA-1980.csv",
            None,
            id="bad_country_code",
        ),
        pytest.param(
            "FRA-123.csv",
            None,
            id="bad_publish_year",
        ),
    ],
)
def test_file_name_pattern_regex(
    instr: str, expected_captures: dict[str, str] | None
) -> None:
    got_mtch = FILE_NAME_PATTERN.fullmatch(instr)
    if got_mtch:
        assert expected_captures is not None
        assert got_mtch.groupdict() == expected_captures
    else:
        assert expected_captures is None

In the above example, the urge to add some verification of the regex is reasonable, as regexes can be surprisingly complicated. However, the test body is more complex than it should be. We can fix this in one of two ways. If we are not willing to change the app/service/library source code/function signature, we can split up the tests into “good cases” and “bad cases.”

Example:

# /// script
# requires-python = ">=3.11"
# dependencies = ["pytest"]
# ///
import pytest


@pytest.mark.parametrize(
    ("instr", "expected_captures"),
    [
        pytest.param(
            "CHN-2025.csv",
            {"iso3_country_code": "CHN", "publish_year": "2025"},
            id="CHN-2025.csv",
        ),
        pytest.param(
            "IND-1980.csv",
            {"iso3_country_code": "IND", "publish_year": "1980"},
            id="IND-1980.csv",
        ),
    ],
)
def test_file_name_pattern_regex(
    instr: str,
    expected_captures: dict[str, str],
) -> None:
    got_mtch = FILE_NAME_PATTERN.fullmatch(instr)
    assert got_mtch is not None
    assert got_mtch.groupdict() == expected_captures


@pytest.mark.parametrize(
    ("instr",),
    [
        pytest.param("FRA-123.csv", id="bad_publish_year"),
        pytest.param("INDIA-1980.csv", id="bad_country_code"),
    ],
)
def test_file_name_pattern_does_not_match(instr: str) -> None:
    assert FILE_NAME_PATTERN.fullmatch(instr) is None

One can also change the application/service/library code to return more information that collapses multiple conditional branches as follows.

Example

from dataclasses import dataclass
import re
from typing import Final
# Application/service/library module


@dataclass(frozen=True)
class FilenameComponents:
    iso3_country_code: str
    publish_year: int


_FILE_NAME_PATTERN: Final[re.Pattern[str]] = re.compile(
    r"(?P<iso3_country_code>[A-Z]{3})\-(?P<publish_year>\d{4})\.csv"
)


def parse_file_name(filename: str) -> FilenameComponents | None:
    mtch = _FILE_NAME_PATTERN.fullmatch(filename)
    if mtch is None:
        return None
    captured = mtch.groupdict()
    # These hard accesses and parses should be safe when the regex matches
    return FilenameComponents(
        iso3_country_code=captured["iso3_country_code"],
        publish_year=int(captured["publish_year"]),
    )


# Test module

import pytest


@pytest.mark.parametrize(
    ("instr", "expected_parsed"),
    [
        pytest.param(
            "CHN-2025.csv",
            FilenameComponents(iso3_country_code="CHN", publish_year=2025),
            id="CHN-2025.csv",
        ),
        pytest.param(
            "IND-1980.csv",
            FilenameComponents(iso3_country_code="IND", publish_year=1980),
            id="IND-1980.csv",
        ),
        pytest.param(
            "INDIA-1980.csv",
            None,
            id="bad_country_code",
        ),
        pytest.param(
            "FRA-123.csv",
            None,
            id="bad_publish_year",
        ),
    ],
)
def test_file_name_pattern_regex(
    instr: str, expected_parsed: FilenameComponents | None
) -> None:
    got = parse_file_name(instr)
    assert got == expected_parsed

Logging

Log.1. Log One Canonical, Structured, Wide Completion/Outcome Event Per Unit of Work

Accumulate context over a meaningful unit of work and emit a comprehensive structured event, possibly with isolated logs for important events, such as security-relevant events, important state transitions, heartbeats/progress updates for especially long-running operations, or circuit-breakers. This applies for long-running HTTP/gRPC services, scripts, background jobs, queue workers, and workflows.

This is especially useful in programs that will run for longer than 10 seconds or programs that will be deployed to production because we typically ask “what is x program doing now/what is taking so long” of those such programs.

Reason: In the context of a long-running backend service that provides real value in a business setting and that requires debugging during outages, structured canonical events fundamentally transform debugging from “trudging through swamps of text” to “answering directed questions with queries over structured data.” The 8-32 fields emitted in a wide structured log event, especially high-cardinality non-sensitive fields (e.g. n_rows_processed, total_postgres_users_query_time_ms, n_s3_objects_put, job_status, bytes_written_to_s3, subscription_tier, job_id, route, status_code_detailed, environment) at the end of each unit of work overlap tremendously with the (typically lower cardinality) fields that are the bread-and-butter of dedicated metrics systems such as Datadog, the Grafana LGTM Stack (Loki, Grafana, Tempo, Mimir) , VictoriaLogs, and Splunk. Machine-parsable logs feed nicely into metrics systems and provide a quick path to more dedicated metrics-based observability systems (e.g. Datadog logs vs metrics). In production environments, logs have costs. Log storage costs money. Indexing for searchability costs money. Noisy logs cost time when debugging. Extra downtime from trudging through log noise costs money. Each log line must pull its weight and actively help in some reasonably probable production outage scenario.

Example

# /// script
# requires-python = ">=3.11"
# dependencies = ["fastapi"]
# ///

import asyncio
from dataclasses import dataclass
from http import HTTPStatus
import logging
from enum import unique, StrEnum
from typing import Literal

from fastapi import APIRouter, FastAPI, Request, Response
from pydantic import BaseModel

_LOGGER = logging.getLogger(__name__)

_RawWideEventT = dict[str, str | int]  # passed into _LOGGER.info(..., extra=...)


@dataclass(frozen=True)
class _GetJobSpecLogInfo:
    """Log-worthy information accreted while processing a GET job spec request"""

    ...


@unique
class _GetJobSpecificationErrorReason(StrEnum):
    JOB_NOT_FOUND = "JOB_NOT_FOUND"
    TIMEOUT = "TIMEOUT"
    ...


@dataclass(frozen=True)
class _GetJobSpecificationWideEventSuccess:
    request_duration_ms: int
    status: Literal["success"] = "success"
    ...

    def as_raw(self) -> _RawWideEventT: ...


@dataclass(frozen=True)
class _GetJobSpecificationWideEventError:
    request_duration_ms: int
    status_code_detailed: _GetJobSpecificationErrorReason
    error_message: str
    status: Literal["error"] = "error"
    ...

    def as_raw(self) -> _RawWideEventT: ...


_GetJobSpecificationWideEvent = (
    _GetJobSpecificationWideEventSuccess | _GetJobSpecificationWideEventError
)


class JobSpecificationResponse(BaseModel): ...


api_v1_router = APIRouter(prefix="/api/v1")
app = FastAPI()


async def _get_job_spec_by_id(
    job_id: int,
) -> tuple[JobSpecificationResponse, _GetJobSpecLogInfo]:
    """Query the postgresql database for a job spec"""


def _basic_wide_event(
    ctx: _GetJobSpecLogInfo, request: Request
) -> _GetJobSpecificationWideEventSuccess:
    """Canonical event metadata for a success getting a job spec. Just moving
    data around, and should never raise an error"""
    ...


def _get_job_uncaught_exception_wide_event(
    job_id: int, request: Request
) -> _GetJobSpecificationWideEventError:
    """Canonical wide event metadata for an unexpected error getting a job spec"""
    ...


def _get_job_timeout_wide_event(
    job_id: int, request: Request
) -> _GetJobSpecificationWideEventError:
    """Canonical wide event metadata for a timeout error getting a job spec"""
    ...


def _log_and_publish_metric(event: _GetJobSpecificationWideEvent) -> None: ...


@api_v1_router.get(
    "/job/{job_id}", response_model=JobSpecificationResponse
)  # Add more params for documentation as needed
async def _get_job_id(job_id: int, request: Request):
    """Get a job specification by its ID"""
    try:
        resp, ctx = await asyncio.wait_for(_get_job_spec_by_id(job_id), timeout=5.0)
    except asyncio.TimeoutError:
        _event = _get_job_timeout_wide_event(job_id, request)
        _log_and_publish_metric(_event)
        return Response(status_code=HTTPStatus.REQUEST_TIMEOUT)
    except Exception:
        _event = _get_job_uncaught_exception_wide_event(job_id, request)
        _log_and_publish_metric(_event)
        return Response(status_code=HTTPStatus.INTERNAL_SERVER_ERROR)

    _event = _basic_wide_event(ctx, request)
    _log_and_publish_metric(_event)
    return resp


app.include_router(api_v1_router)

Any service worthy of existence and whose absence will be missed should be able to give some indication of its state per some unit of work that can be exposed at least via logs.

Command Line Interfaces

CLI.1. Use the standard library’s argparse module for command line interfaces.

Reason: The argparse module conforms with most unix-style CLI conventions so that CLIs “feel” at home in unix-like environments, has very effective error handling, gets regular improvements – including more humane error handling and color output introduced in Python 3.14 –, and is featureful enough for the vast majority of CLIs. Most CLIs should be simple and argparse tastefully implements all of the core components of simple CLIs.

Example

This example is slightly more than minimal precisely to demonstrate how to integrate a CLI into, for example, a data processing operation.

#!/usr/bin/env python3
"""An example showing common features of the _many_ Python CLIs that I
write."""

from __future__ import annotations
import logging
from collections.abc import Sequence
from argparse import ArgumentParser, Namespace, ArgumentDefaultsHelpFormatter
from pathlib import Path
import sys
from datetime import date, timedelta

LOGGER = logging.getLogger(__name__)

_VERBOSITY_TO_LOG_LEVEL: dict[int, int] = {
    0: logging.WARNING,
    1: logging.INFO,
    2: logging.DEBUG,
}


class _MyProgOpts(Namespace):
    verbosity: int
    start_date: date
    end_date: date
    input_data_base_dir: Path
    output_parquet_base_dir: Path


def _get_parser() -> ArgumentParser:
    parser = ArgumentParser(
        description=__doc__, formatter_class=ArgumentDefaultsHelpFormatter
    )
    parser.add_argument(
        "-v",
        "--verbose",
        action="count",
        default=0,
        help="Logging verbosity. Provide 0 to 2 times.",
        dest="verbosity",
    )
    parser.add_argument(
        "--start-date",
        type=date.fromisoformat,
        default=date.today() - timedelta(days=3),
        help="The first date (inclusive) of data to process. ISO-8601 date format",
    )
    parser.add_argument(
        "--end-date",
        type=date.fromisoformat,
        default=date.today(),
        help="The last date (inclusive) of data to process. ISO-8601 date format",
    )
    parser.add_argument(
        "input_data_base_dir",
        type=Path,
        help="The directory of the raw x dataset. zip archives named like 'some_data_2026-01-01.zip'",
    )
    parser.add_argument(
        "output_parquet_base_dir",
        type=Path,
        help="The base directory of the hive-style partitioned parquet dataset to write/insert into",
    )
    return parser


def _run(opts: _MyProgOpts) -> int:
    logging.basicConfig(
        level=_VERBOSITY_TO_LOG_LEVEL.get(opts.verbosity, logging.DEBUG)
    )
    LOGGER.info("Beginning sample program with opts %s", opts)
    from mypackage._workhorse import etl_data

    written_paths = etl_data(
        start=opts.start_date,
        end=opts.end_date,
        input_data_base_dir=opts.input_data_base_dir,
        output_parquet_base_dir=opts.output_parquet_base_dir,
    )
    LOGGER.info(
        "Wrote %s files to %s", len(written_paths), opts.output_parquet_base_dir
    )
    return 0


def main(args: Sequence[str] | None = None) -> int:
    parser = _get_parser()
    opts: _MyProgOpts = parser.parse_args(args, _MyProgOpts())
    return _run(opts)


if __name__ == "__main__":
    sys.exit(main())

There is no need to comment the main, _run, _get_parser or the Namespace implementing class here. These are idioms and well understood among professionals who write reliable Python. The great deal of functionality in the CLI justifies the length.

Each function has a very specific purpose. main is the core interface between the (C)Python runtime and the OS, insofar as parse_args typically receives None at runtime and pulls from sys.argv. main bridges the world of a string list of arguments with the world of static analyzability. We pass args as a parameter to main to enable easy testing of main – just have the test import main and call main with the same CLI flags one would pass to the program. The run function is the bridge between argparse and the core behavior of our program. In larger multi-file applications, prefer to not make the CLI entrypoint module much larger than this. There are likely more useful interfaces/invocation methods such as a HTTP service, Apache Airflow DAG, GRPC service, or even as a library.

Even though type checkers do not entirely validate the information-flow from argparse.ArgumentParser to the subclassed instance of argparse.Namespace, the proximity of those two within the same module typically enables simple debugging.

Bad Example Using a third-party CLI package such as typer for a simple CLI with fewer than 16 parameters/options is a bad idea. While typer may cut down on a bit of code and titillate the senses of those who like the idea of exposing a decorated function and driving behavior from type annotations, the supply chain risks from extra dependencies that add no runtime or “user”-facing value (many CLIs are predominantly used by computers rather than by humans) outweigh the burden of maybe 5-10 more lines of code.

Exception: More involved terminal-based user interfaces (e.g. TUIs, “terminal-widget”-style libraries) may benefit from a CLI that is not argparse. “Flashy” UI is not the distinguishing factor of most CLIs, however. Of course, many “single-use”/“throwaway” Python scripts do not need to abide by this example. Maybe a good rule of thumb is “if you ever expect a human or LLM to use/understand the CLI more than once, use argparse in the style exemplified above.

Packaging Python Programs

Pkg.1. Manually setting the PYTHONPATH environment variable is ALWAYS suspect.

This entire section largely consists of better alternatives.

Reason: There is almost never a need to set PYTHONPATH for a properly packaged Python program. Very few who write Python programs in professional settings understand the module resolution rules of PYTHONPATH. Indeed there are many “gotchas.”

Pkg.2. Use PEP-723-style inline metadata for single-file standalone scripts that only require a couple of dependencies

Reason: A dedicated task-runner such as pipx or uv run can reliably execute a script with dependencies without necessarily needing a full package structure. PEP-723 is a nice middle ground between a fully reproducible package structure and non-reproducible programs that use third-party dependencies without declaring them.

Pkg.3. Prefer to put packaging metadata and tool configuration in pyproject.toml over setup.cfg and setup.py.

Reason: pyproject.toml, initially specified in PEP-518, refined in PEP-621, and continually specified by the Python Packaging Authority, works with many tools such as hatch, uv, and pdm. By contrast, setup.py and setup.cfg are only a part of setuptools. There are situations where setuptools’ tendency to put egg files and other metadata in the project and source directories is undesirable. Moreover the ability to easily switch between build backends by only changing the build-system section of pyproject.toml setup.cfg format have formalized specifications.

Exceptions: Certain teams center deployment tooling around setup.py. In such cases, keeping with team conventions allows the team to operate coherently and uniformly, at least until those tools and standards have deliberately considered the potential benefits of using pyproject.toml.

Pkg.4. Minimize runtime/deployed dependencies

Reason: Every dependency at runtime is a potential attack vector. 2026 saw many high-profile supply chain attacks, many of which operated on transitive dependencies and hijacked deployment pipelines without the knowledge of library maintainers. Dependencies that only operate in CI (e.g. linters), may be less likely to cause problems for deployed applications. The days of PyPI mainly consisting of good faith actors are long gone.

Examples

  • Prefer the zoneinfo standard library module to the older pytz third party package. Added in Python 3.9, zoneinfo is now available in every currently supported Python version. Prefer to also depend on the tzdata package maintained under the Python umbrella/organization when deploying to minimal environments lacking the olson/IANA timezone database (e.g. distroless containers).
  • Prefer the standard library’s datetime functionalities over dateutil. The standard library properly and reliably handles almost all use cases of this once venerable package.
  • Prefer the dataclasses standard library module over the attrs third party package that inspired it for the vast majority of “plain-old-data” containers especially for data that is generated within the Python program, unless there is a very good reason.
  • Prefer the json standard library module over the orjson third party package unless the application has specific performance requirements and JSON (de)serialization is a measured bottleneck.

The previous examples, though they involve fairly stable and reliable third-party packages, exist to highlight the fact that smaller dependency graphs translate to faster and cheaper CI runs, smaller supply chain attack surfaces for deployed services, quicker ramp-up for newer contributors (Python developers are expected to understand the standard library), and a better maintenance story insofar as more eyeballs around the world will notice security issues in CPython than in some third party package.

Exceptions: At the same time, avoid contorting the standard library into something outside of its intended scope. For example, the standard library’s urllib.request is sufficient for simple scripts that make one or two HTTP requests to known ‘well-behaved’ services, however long-running services that will make many requests should consider a battle-tested third party HTTP client such as aiohttp for a better path to handling real-world considerations like connection pooling, redirects, authentication, cookies, stream decompression, and multipart form uploads. The lack of a production-grade HTTP client is a known limitation of the Python standard library as of 2026. As another example, database clients typically implement involved wire protocols (sometimes wrapping high-performance C, C++, or Rust implementations) that no HTTP or gRPC service should reimplement.

Pkg.5. Pin dependencies for long-running services

While libraries must permit a range of dependencies to prevent conflicts of consumer applications, the ’leaves’ in a dependency graph can and should pin dependencies, while making dependency upgrades dedicated commits.

Reason: This helps minimize supply chain attacks. A minor bug fix must not introduce vulnerabilities far beyond the scope of the affected code. Pinning dependencies cuts off many attack vectors and makes the risk explicit with dedicated dependency bump commits.

This can be accomplished in a few ways. For example, one can have a requirements-minimal.txt, and generate frozen/pinned requirements.txt with uv pip compile/pip-compile that is the actual list of dependencies used by the service. uv.lock can also accomplish this, while also enforcing package checksums and keep track of package source registries.

Polymorphism in Python

Poly.1. Prefer functools.singledispatch over inheritance for “closed” polymorphism when the input and output types are well-specified, when the different strategies are determined by externally defined specifications, and when one module/file can control all strategies/implementations.

Reason: This cleanly separates “behavior” from “data” without all of the extra baggage of inheritance, especially when disallowing subclassing of strategies/variants. Placing all implementations/registered functions in one module minimizes the complexity of function registry via import side effect. singledispatch tends to outlive its usefulness when there are enough implementations to render organizing them into 1 module impractical.

functools.singledispatch implements a multi-method by dispatching on the type of the first function parameter.

Example

from dataclasses import dataclass


@dataclass(frozen=True)
class _PreprocessContext: ...


# spec.py, possibly in another package/part of the monorepo
@dataclass(frozen=True)
class PDFAttachmentPreprocess:
    compression_level: int


@dataclass(frozen=True)
class JpegAttachmentPreprocess:
    max_target_size_bytes: int


@dataclass(frozen=True)
class ZipArchivePreprocess: ...


TAttachment = bytes
Preprocess = PDFAttachmentPreprocess | JpegAttachmentPreprocess | ZipArchivePreprocess


def infer_preprocess_procedure(attachment: TAttachment) -> Preprocess | None:
    """Use libmagic to determine what kind of file we are dealing with"""


# Dedicated registry module with all implementations
from functools import singledispatch


@singledispatch
def preprocess_attachment_before_upload(
    preprocess: Preprocess,
    attachment: TAttachment,
    ctx: _PreprocessContext,
) -> bytes:
    raise NotImplementedError()


@preprocess_attachment_before_upload.register
def _(
    preprocess: JpegAttachmentPreprocess,
    attachment: TAttachment,
    ctx: _PreprocessContext,
) -> bytes:
    ...
    # Maybe each implementation keeps the core body in separate modules and this
    # registry forwards parameters to the variant-specific modules, leaving this
    # registry module as a unified entrypoint to multiple fully known and
    # controlled backends


@preprocess_attachment_before_upload.register
def _(
    preprocess: PDFAttachmentPreprocess,
    attachment: TAttachment,
    ctx: _PreprocessContext,
) -> bytes: ...

Now we can either test just the preprocess_attachment_before_upload function or we can give the registered function an actual name and call those in tests, though we should prefer to just use preprocess_attachment_before_upload because the rest of the library/service/application code will call that.

For single-file scripts where all of the strategies are specified and implemented in one module, an if-elif-else chain encapsulated in a function is simpler and faster.

Performance Note: singledispatch has some performance overhead when compared to inheritance, measured in single-digit microseconds. This type of dispatch typically happens once or a couple of times per work unit, rather than millions of times in a tight loop. Profile before determining that this matters. And if it does matter, maybe Python is not the best language for the service/application.

Poly.2. Prefer a plugin system via packaged entrypoints for “open” polymorphism where the library does not control all implementations.

TODO.

Style

Style.1. Drive the core behavior of an application or service via functions rather than by classes/methods

Reason: Classes are for storing data or state, and for enforcing invariants of encapsulated/hidden attributes. Functions are fundamentally simpler and “flatter” than classes insofar as one rarely considers the distinctions between state, value, identity when using functions.

Bad Example

class APIDownloader:
    def __init__(self, api_key: str, url: str) -> None: ...
    async def download_to_file(self, output_file_name: str) -> None: ...

There are two dead giveaways that this class is unnecessary. First, a class with 2 public/invoked methods – one of them being __init__ – should almost always be a function. Second, a class whose name is a noun derived from a verb should likely be a function.

Example

async def download_from_api_to_file(
    *,
    url: str,
    api_key: str,
    output_file_name: str,
) -> None: ...