Python performance
Skill sairam0424/MindForge/.mindforge/skills/python-performance
MindForge: The Enterprise Agentic Framework for Claude Code & Antigravity. High-performance autonomous execution, wave-parallelism, and multi-tier governance for production-grade AI engineering.From the repository description
npx -y skills add sairam0424/MindForge --skill python-performanceAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
7.8 KB, ~1.8k tokens by cl100k_base, as published. Nobody here has run it
Skill — Python Performance
When this skill activates
Any task involving Python runtime performance: identifying bottlenecks, reducing execution time, fixing memory leaks, optimizing async code, parallelizing workloads, or replacing slow Python loops with vectorized operations.
Mandatory actions when this skill is active
Before writing any code
- Establish a baseline measurement. Never optimize without numbers. Identify:
- What is slow? (wall time, CPU time, memory growth, I/O wait)
- How slow? (current measurement with units)
- What is the target? (acceptable latency/throughput/memory)
- Select the correct profiling tool for the problem:
Symptom Tool Command "It's slow" (general) cProfile + snakeviz python -m cProfile -o prof.out script.py && snakeviz prof.out"This function is slow" (line-level) line_profiler kernprof -l -v script.py(decorate with@profile)"Memory keeps growing" memory_profiler python -m memory_profiler script.py(decorate with@profile)"Memory leaks in long-running process" tracemalloc tracemalloc.start(); snapshot = tracemalloc.take_snapshot()"I/O bound" py-spy (sampling) py-spy top --pid <PID>"GIL contention" py-spy py-spy record --native -o profile.svg --pid <PID> - Do NOT guess. Profile first, then optimize the hottest path.
During implementation
Profiling Workflow
# Quick cProfile workflow
import cProfile
import pstats
profiler = cProfile.Profile()
profiler.enable()
# ... code under test ...
profiler.disable()
stats = pstats.Stats(profiler).sort_stats("cumulative")
stats.print_stats(20) # Top 20 by cumulative time
Asyncio Optimization
- Use
asyncio.TaskGroup(Python 3.11+) overasyncio.gatherfor structured concurrency:async with asyncio.TaskGroup() as tg: task1 = tg.create_task(fetch_user(user_id)) task2 = tg.create_task(fetch_orders(user_id)) # Both complete or both cancel on error - Never block the event loop with synchronous I/O. Offload with
loop.run_in_executor:result = await loop.run_in_executor(None, blocking_io_function, arg) - Avoid
awaitin tight loops. Batch operations:# Bad: sequential awaits for item in items: await process(item) # Good: concurrent with bounded concurrency sem = asyncio.Semaphore(10) async def bounded(item): async with sem: return await process(item) results = await asyncio.gather(*(bounded(i) for i in items)) - Use
asyncio.Queuefor producer-consumer patterns instead of polling. - Set appropriate timeouts on all network calls:
async with asyncio.timeout(5.0).
Multiprocessing
- Use
ProcessPoolExecutorfor CPU-bound work (bypasses GIL):from concurrent.futures import ProcessPoolExecutor with ProcessPoolExecutor(max_workers=os.cpu_count()) as pool: results = list(pool.map(cpu_heavy_function, data_chunks)) - Use
multiprocessing.shared_memoryfor large data to avoid serialization overhead:from multiprocessing import shared_memory shm = shared_memory.SharedMemory(create=True, size=array.nbytes) shared_array = np.ndarray(array.shape, dtype=array.dtype, buffer=shm.buf) - Prefer
Pool.map/Pool.imap_unorderedover manual Process management. - Watch for pickling overhead: large objects passed to workers are serialized. Keep payloads small.
- Use
threading(not multiprocessing) for I/O-bound parallelism (GIL is released during I/O).
NumPy Vectorization
- Replace Python loops over numeric data with vectorized operations:
# Bad: 100x slower result = [x * 2 + 1 for x in data] # Good: vectorized result = data * 2 + 1 - Use boolean indexing instead of conditional loops:
# Bad filtered = [x for x in data if x > threshold] # Good filtered = data[data > threshold] - Use
np.wherefor conditional assignment,np.einsumfor tensor operations. - Avoid creating intermediate arrays in chains. Use
out=parameter or in-place operations. - For operations NumPy cannot vectorize: try
numba.jit(nopython=True).
Generator Pipelines (Memory Efficiency)
- Use generators for processing large datasets that do not fit in memory:
def read_chunks(path, chunk_size=8192): with open(path, "rb") as f: while chunk := f.read(chunk_size): yield chunk def process_pipeline(path): chunks = read_chunks(path) decoded = (chunk.decode("utf-8") for chunk in chunks) lines = (line for chunk in decoded for line in chunk.splitlines()) return (parse(line) for line in lines if line.strip()) - Use
itertools(chain, islice, groupby) for composing lazy pipelines. - For pandas: use
chunksizeparameter inread_csvfor row-by-row processing.
Memory Optimization
- Use
__slots__on data classes with many instances:class Point: __slots__ = ("x", "y", "z") def __init__(self, x: float, y: float, z: float): self.x, self.y, self.z = x, y, z # Saves ~40% memory per instance vs regular class - Use
sys.getsizeofandpympler.asizeofto measure object memory. - Prefer
array.arrayover lists for homogeneous numeric data. - Use
weakreffor caches that should not prevent garbage collection. - Intern frequently repeated strings:
sys.intern(string).
C Extension Paths (When Python Is Not Fast Enough)
- Cython: annotate hot functions with
cdeftypes for 10-100x speedup. - pybind11: wrap existing C++ code with minimal boilerplate.
- ctypes/cffi: call existing shared libraries without compilation.
- Decision rule: Only reach for C extensions after profiling proves Python is the bottleneck AND vectorization/multiprocessing cannot solve it.
After implementation
- Re-run the same profiling tool used for the baseline.
- Compare before/after measurements with the same input data.
- Document the optimization in a comment or docstring explaining WHY and the measured improvement:
# Optimized from 4.2s to 0.3s by vectorizing the distance calculation. # See: profiling results in docs/perf/distance-calc-2024-03.md - Verify correctness: optimized code must produce identical results to the original.
- Run the test suite to ensure no regressions.
Performance anti-patterns to flag
- Premature
@lru_cacheon functions with large/unhashable arguments (memory leak). globalinterpreter lock assumptions (using threads for CPU-bound work).- String concatenation in loops (
+=creates new string each time; use"".join(parts)). - Repeated dictionary/attribute lookups in tight loops (hoist outside).
- Using
pandasfor row-by-row iteration (iterrowsis 100x slower than vectorized). - Creating DataFrames inside loops (pre-allocate or use list-of-dicts then single concat).
Self-check before task completion
Before marking a task done when this skill was active:
- Baseline measurement exists (before optimization).
- After measurement shows quantified improvement.
- Profiling tool output confirms the hotspot was addressed.
- Correctness verified (same output as before optimization).
- No new memory leaks introduced (check with tracemalloc for long-running processes).
- Test suite passes without regressions.
- Optimization is documented with measured numbers in code or docs.
- No premature optimization: only the measured bottleneck was addressed.
What ships with it
Read from the repository
Just SKILL.md. No reference files, no scripts.