Jupyter Notebooks
Jupy.0. Good notebooks use data to tell a story and comments need to
concisely and tastefully facilitate the storytelling.
Jupy.1. Persist notebooks in a format that is amenable to version-control,
executable, and well-documented. Export to HTML or PDF as desired. As of 2026,
jupytext works well for this.
Reason: Version control and well-documented formats enable reproducibility and repeatability of analysis. With jupytext, some guidelines enable all analytical notebooks to be executable and immediately reproducible.
Jupy.2. Hoist environment-specific values that are constant for the
notebook runtime to the top of the notebook.
Jupy.3. Structure the majority of the analysis, data-fetching,
visualization, and other behaviors around functions that have individual/dedicated cells/blocks for performing side-effects such as drawing a graph to the screen, writing a file, or displaying a small table.
Reproducibility
Repro.1. Insight-generating scripts/notebooks should be executable in under
2 minutes.
Reason: Analytical notebooks (e.x. Jupyter/Jupytext), at their best, combine simple code – flat, human-readable, familiar – with formatted text and some reproducible images/tables for telling stories with data. They are terrible for substantial data processing/ingestion, ETL, and long-running jobs. In an ideal world, heavy data processing is handled in the deliberately-tuned data warehouse, such as Clickhouse, Postgresql, or BigQuery. Long running jobs have more failure modes than quick jobs.
Performance
Perf.1. Move as much computation as possible to dedicated and finely-tuned services,
such as data warehouses, or services in front of data warehouses.
Perf.2. Use dedicated computation tools such as duckdb, polars, and
chdb.
Perf.3. Always run queries on dedicated databases or atleast parquet files
and never start performing analysis on CSVs, JSON, XML, HTML until converting/ingesting it to a statically-typed, optimized, compressed, possibly indexed, column-oriented representation.
Perf.3. Put SQL at the forefront of the “data preparation/transformation”
that will feed into dataframes input into regression/predictive/ models and plotting tools.