agentsclimarketplace

Concurrency tuning

Skill Amey-Thakur/AI-SKILLS/skills/performance/concurrency-tuning

Plug-and-play skills and prompts for every AI coding agent

Install
npx -y skills add Amey-Thakur/AI-SKILLS --skill concurrency-tuning

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • 21 days oldThe repository was created 21 days ago. New is not bad, but a brand new repository carrying a familiar-sounding name is the shape a typosquat arrives in, and there has been no time for anyone else to find a problem with it.
  • 4 stars4 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Size worker pools, queue depths, and batch widths against the real bottleneck so parallelism adds throughput instead of contention. Use when a parallel job stops scaling, saturates a resource, or spends more time coordinating than working.

SKILL.md

2.8 KB, 610 tokens by cl100k_base, as published. Nobody here has run it

Concurrency tuning

Adding workers helps only until some shared resource saturates: a CPU core, a connection pool, a disk head, a lock. Past that point every extra worker adds context switches and contention while throughput flattens or drops. Tuning means finding that knee and parking just short of it.

Method

  1. Classify the work first. CPU-bound work scales to roughly the core count: set the pool near nproc (or os.cpu_count()), not higher. IO-bound work waits on the network or disk, so it wants far more concurrency than cores, bounded by the downstream limit instead.
  2. Find the binding resource before touching worker count. Run the job and watch htop, iostat -x 1, and the database's active-connection count at once. The resource pegged at 100 percent is your ceiling; raising workers past it only lengthens queues.
  3. Bound queue depth on purpose. An unbounded queue turns a slow consumer into an out-of-memory crash. Cap it (a Semaphore, a fixed maxsize, a channel buffer) so producers block and apply backpressure rather than buffering gigabytes of pending work.
  4. Respect Amdahl's law. If 20 percent of the wall time is serial (setup, a global lock, a final merge), maximum speedup is 5x no matter how many workers you add. Measure the serial fraction and attack it before buying more parallelism that cannot pay off.
  5. Match pool size to the scarcest downstream limit. Forty workers hitting a database capped at 20 connections means 20 threads block on the pool. Set worker count at or below the connection ceiling, or the extra threads are pure overhead.
  6. Sweep, do not guess. Run the job at 2, 4, 8, 16, 32 workers and plot throughput. Pick the point where the curve flattens; the setting past the knee costs memory and tail latency for no gain.

Signals

  • Does throughput actually rise between your current setting and the next step up, or has the curve already gone flat?
  • Is exactly one resource pegged at 100 percent, confirming the real ceiling?
  • Is every queue in the pipeline bounded, so a stall blocks rather than balloons memory?
  • Does the measured speedup track the serial fraction Amdahl predicts?

Boundaries

This skill sizes pools and queues for throughput. Correctly sharing mutable state between those workers is a separate concern: race conditions, lock ordering, and atomicity belong to concurrency-safety review. Distributing work across machines rather than threads is a scheduling and partitioning problem, not a pool-tuning one.

What ships with it

Read from the repository

Just SKILL.md. No reference files, no scripts.

Keep looking

Skills are one crate of 327,069. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.