Jupyter notebook
Skill furkangonel/cowrangler/bundled_skills/data-science/jupyter-notebook
Autonomous terminal AI agent for workflows and feasible project procedures. Co-Worker Co-Wrangler π
npx -y skills add furkangonel/cowrangler --skill jupyter-notebookAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 2 stars2 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Jupyter notebook structure, reproducibility, and best practices for data science.
SKILL.md
9.3 KB, as published. Nobody here has run it
Jupyter Notebook SOP
When to Use
- User is writing or improving a Jupyter notebook
- User wants to share or publish a notebook as a report
- User has reproducibility issues (notebook works for them but not others)
- User wants to structure an analysis or ML experiment cleanly
Part 1 β Recommended Notebook Structure
Use these sections in order. Each section is a Markdown cell followed by code cells.
1. Title & Metadata
2. Imports & Configuration
3. Data Loading
4. Exploratory Data Analysis (EDA)
5. Feature Engineering / Preprocessing
6. Modeling (if applicable)
7. Results & Conclusions
8. Appendix (optional)
Section 1 β Title & Metadata
# Analysis Title
**Author:** Your Name
**Date:** 2025-05-18
**Dataset:** dataset_name.csv (source / version)
**Purpose:** One sentence on what question this notebook answers.
## Summary
Key findings in 3-5 bullet points β fill in after completing the notebook.
Section 2 β Imports & Configuration
# ββ Standard library ββββββββββββββββββββββββββββββββββββββββββββββ
import os
import json
from pathlib import Path
from datetime import datetime
# ββ Data manipulation ββββββββββββββββββββββββββββββββββββββββββββββ
import numpy as np
import pandas as pd
# ββ Visualization βββββββββββββββββββββββββββββββββββββββββββββββββ
import matplotlib.pyplot as plt
import seaborn as sns
import plotly.express as px
# ββ ML (if needed) ββββββββββββββββββββββββββββββββββββββββββββββββ
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
# ββ Config ββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
RANDOM_SEED = 42
DATA_DIR = Path("../data")
OUTPUT_DIR = Path("../outputs")
OUTPUT_DIR.mkdir(exist_ok=True)
# Display settings
pd.set_option("display.max_columns", 50)
pd.set_option("display.max_rows", 100)
pd.set_option("display.float_format", "{:.4f}".format)
plt.rcParams["figure.figsize"] = (12, 6)
plt.rcParams["figure.dpi"] = 100
sns.set_theme(style="whitegrid", palette="muted")
np.random.seed(RANDOM_SEED)
print(f"NumPy: {np.__version__}, Pandas: {pd.__version__}")
Section 3 β Data Loading
# Load data
df = pd.read_csv(DATA_DIR / "raw" / "dataset.csv")
# Confirm load
print(f"Shape: {df.shape}")
print(f"Memory: {df.memory_usage(deep=True).sum() / 1e6:.1f} MB")
df.head()
Section 4 β EDA (see Part 2)
Section 5 β Feature Engineering
# Create new features
df["feature_ratio"] = df["col_a"] / df["col_b"].replace(0, np.nan)
df["log_amount"] = np.log1p(df["amount"])
# Encode categoricals
df = pd.get_dummies(df, columns=["category"], drop_first=True)
# Save processed data
df.to_parquet(DATA_DIR / "processed" / "clean_dataset.parquet", index=False)
print("Saved processed dataset")
Part 2 β EDA Checklist
Run these in sequence for any new dataset:
# ββ 1. Shape and types ββββββββββββββββββββββββββββββββββββββββββββ
print(df.shape)
df.dtypes.value_counts()
# ββ 2. Missing values βββββββββββββββββββββββββββββββββββββββββββββ
missing = df.isnull().sum()
missing_pct = (missing / len(df) * 100).round(2)
pd.DataFrame({"count": missing, "pct": missing_pct}).query("count > 0").sort_values("pct", ascending=False)
# ββ 3. Duplicates βββββββββββββββββββββββββββββββββββββββββββββββββ
print(f"Duplicates: {df.duplicated().sum()} ({df.duplicated().sum()/len(df)*100:.1f}%)")
# ββ 4. Basic statistics ββββββββββββββββββββββββββββββββββββββββββββ
df.describe(include="all").T
# ββ 5. Distributions β numeric ββββββββββββββββββββββββββββββββββββ
numeric_cols = df.select_dtypes(include=np.number).columns
df[numeric_cols].hist(bins=30, figsize=(16, 10))
plt.tight_layout()
plt.show()
# ββ 6. Distributions β categorical βββββββββββββββββββββββββββββββ
cat_cols = df.select_dtypes(include="object").columns
for col in cat_cols[:5]: # limit to first 5
print(f"\n{col}: {df[col].nunique()} unique values")
print(df[col].value_counts().head(10))
# ββ 7. Correlations βββββββββββββββββββββββββββββββββββββββββββββββ
corr = df[numeric_cols].corr()
mask = np.triu(np.ones_like(corr, dtype=bool))
fig, ax = plt.subplots(figsize=(12, 10))
sns.heatmap(corr, mask=mask, annot=True, fmt=".2f", cmap="coolwarm",
center=0, vmin=-1, vmax=1, ax=ax)
plt.title("Correlation Matrix")
plt.tight_layout()
plt.show()
# ββ 8. Target distribution (if supervised learning) βββββββββββββββ
# df["target"].value_counts(normalize=True).plot(kind="bar")
Part 3 β Reproducibility Checklist
Before sharing or publishing:
- Seeds set at the top of the notebook (
RANDOM_SEED = 42, applied to numpy, sklearn, torch) - Relative paths used (
Path("../data/raw/file.csv"), not/home/username/...) - Kernel restart + run all executed β notebook runs clean from top to bottom
- No hidden state β all variables defined in the cell they're first used
- Requirements pinned:
pip freeze > requirements.txt # or with conda: conda env export > environment.yml - Large files excluded from git (use
.gitignoreor git-lfs) - Outputs cleared before committing (
Kernel β Restart & Clear Output) β or use nbstripout - nbconvert tested if sharing as HTML/PDF:
jupyter nbconvert --to html notebook.ipynb jupyter nbconvert --to pdf notebook.ipynb # requires LaTeX
Part 4 β Common Pitfalls
Hidden State (most common bug)
# Cell A
x = 10
# Cell B
y = x + 5
# Cell A (run again with x = 20)
x = 20
# Cell B (y is still 15 from before β hidden state!)
# Restart kernel to avoid this
Fix: Always use Kernel β Restart & Run All before sharing.
In-place DataFrame Mutations
# WRONG β silently modifies df, affects all subsequent cells
df.drop(columns=["col_to_remove"], inplace=True)
# CORRECT β create a new variable
df_clean = df.drop(columns=["col_to_remove"])
Mutable Defaults / Loop Variables
# WRONG β all lambdas capture the same `i` (Python closure gotcha)
fns = [lambda: i for i in range(3)]
[f() for f in fns] # [2, 2, 2]
# CORRECT
fns = [lambda i=i: i for i in range(3)]
Missing Data Silently Propagates
# NaN in one column can silently make aggregations wrong
df["amount"].mean() # NaN if any value is NaN!
df["amount"].mean(skipna=True) # explicit β recommended
Part 5 β Useful Magic Commands
# Timing
%time df.groupby("category").agg({"value": "sum"})
%timeit -n 3 df["col"].apply(lambda x: x**2)
# Memory profiling (requires memory_profiler)
%load_ext memory_profiler
%memit df.merge(df2, on="id")
# Run shell commands
!ls ../data/raw/
!pip install polars --quiet
# Show all variables
%whos DataFrame
# Reload a module after editing it externally
%load_ext autoreload
%autoreload 2
Part 6 β nbconvert Export
# HTML report (self-contained, shareable)
jupyter nbconvert --to html --no-input notebook.ipynb
# --no-input hides code cells β good for stakeholder reports
# Slides (reveal.js)
jupyter nbconvert --to slides notebook.ipynb --post serve
# Python script (for productionizing)
jupyter nbconvert --to script notebook.ipynb
# Strip outputs before committing
pip install nbstripout
nbstripout notebook.ipynb
# Or configure globally:
nbstripout --install
Agent Instructions
- When writing cells, group logically related operations in one cell β not one operation per cell
- Add a Markdown header before each major section β makes the notebook navigable
- Print shape and sample (
df.head()) after every data transformation β reduces debugging time - Always set and use
RANDOM_SEEDin the config section - Suggest
Kernel β Restart & Run Allbefore the user shares or declares the notebook "done" - If the user has a reproducibility bug, first ask: "Did you restart the kernel and run all cells fresh?"
- For large datasets (>1M rows), suggest
polarsor chunkedpandasreading