agentsclimarketplace

Caching

Skill MARUCIE/openclaw-foundry/web/public/packs/spellbook-code-reviewer/skills/caching

The curated AI Agent skill marketplace — 37K+ vetted skills, S/A/B/C ratings, deploy anywhere

Install
npx -y skills add MARUCIE/openclaw-foundry --skill caching

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

One thing to look at

  • 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

What its author says it does

Copied from the file, not written here

Use when adding or debugging caching in a service — choosing a cache strategy, designing TTLs, preventing stampedes, reasoning about invalidation, or configuring HTTP Cache-Control headers.

SKILL.md

14.3 KB, as published. Nobody here has run it

是什么

这是一份缓存设计规范,帮团队判断什么场景该加缓存、用哪种缓存策略、怎么避免脏数据,让接口响应时间从几百毫秒降到几十毫秒,同时不会因为缓存失效带来线上故障。

怎么用

  1. 接口变慢或数据库压力大时,先按本规范判断属于读多写少还是热点数据,再决定加哪一层缓存。
  2. 设计缓存键命名和 TTL(过期时间)时,遵循本文档约定,避免 key 冲突和无限堆积导致内存爆炸。
  3. 选用 Cache-Aside(旁路缓存)、Read-Through(读穿透)、Write-Through(写穿透)等模式前,对照失效场景小节评估数据一致性风险。
  4. Code Review 时检查同事的缓存改动是否覆盖了缓存击穿、雪崩、穿透三大经典问题。
  5. 上线后通过命中率指标验证缓存效果,命中率低于 80% 说明策略需要回炉。

架构图

flowchart LR
    A[请求到达] --> B{命中缓存?}
    B -->|是| C[直接返回]
    B -->|否| D[回源数据库]
    D --> E[写入缓存]
    E --> C

Caching Patterns

Strategies and implementation patterns for application-level, distributed, and HTTP caching.

When to Activate

  • Adding Redis or Memcached to reduce database load or API latency
  • Designing TTL values and cache invalidation strategies
  • Preventing cache stampede on high-traffic keys
  • Configuring HTTP Cache-Control and CDN caching rules
  • Choosing between cache-aside, write-through, or write-behind
  • Debugging stale data, cache poisoning, or thundering herd problems
  • Sizing a cache or deciding what to cache vs. not cache

Strategy Selection

StrategyHowBest For
Cache-aside (lazy)App checks cache first; on miss, loads from DB, populates cacheGeneral-purpose read caching
Write-throughWrite to cache and DB simultaneouslyData that's read immediately after write
Write-behind (write-back)Write to cache; async flush to DBHigh write throughput, tolerance for small loss window
Read-throughCache fetches from DB on miss (cache manages itself)Managed caches (ElastiCache DAX, Momento)
Refresh-aheadProactively refresh before expiryPredictable access patterns, zero-miss latency required

Cache-Aside (Most Common)

# Python — cache-aside with Redis
import redis, json, hashlib
from typing import Callable, TypeVar

T = TypeVar("T")
r = redis.Redis(host="redis", port=6379, decode_responses=True)

def get_or_set(key: str, loader: Callable[[], T], ttl: int = 300) -> T:
    cached = r.get(key)
    if cached is not None:
        return json.loads(cached)

    value = loader()
    r.setex(key, ttl, json.dumps(value, default=str))
    return value

# Usage
user = get_or_set(f"user:{user_id}", lambda: db.query(User).get(user_id), ttl=600)
// TypeScript — cache-aside
import { createClient } from "redis";

const redis = createClient({ url: "redis://redis:6379" });

async function getOrSet<T>(
  key: string,
  loader: () => Promise<T>,
  ttlSeconds = 300,
): Promise<T> {
  const cached = await redis.get(key);
  if (cached) return JSON.parse(cached) as T;

  const value = await loader();
  await redis.setEx(key, ttlSeconds, JSON.stringify(value));
  return value;
}
// Go — cache-aside
func (c *Cache) GetOrSet(ctx context.Context, key string, loader func() (any, error), ttl time.Duration) (any, error) {
    val, err := c.redis.Get(ctx, key).Result()
    if err == nil {
        var result any
        json.Unmarshal([]byte(val), &result)
        return result, nil
    }
    if !errors.Is(err, redis.Nil) {
        return nil, err
    }

    data, err := loader()
    if err != nil {
        return nil, err
    }
    b, _ := json.Marshal(data)
    c.redis.SetEx(ctx, key, string(b), ttl)
    return data, nil
}

Cache Key Design

# Pattern: <service>:<entity>:<id>[:<variant>]
user:profile:123
user:orders:123:active
product:detail:sku-456
search:results:<md5(query+filters)>

# BAD: too broad — invalidation nukes unrelated data
cache_key = "users"

# BAD: too granular — misses sharing opportunity
cache_key = f"user_orders_by_{user_id}_status_{status}_page_{page}"

# GOOD: namespace + entity + discriminator
cache_key = f"user:{user_id}:orders:{status}"   # paginate in app, not in key

TTL Design

Data TypeTTL RangeReasoning
User session15–60 min (sliding)Balance UX vs. stale auth
User profile5–15 minInfrequent changes, high read volume
Product catalog1–24 hrChanges only on explicit update
Search results1–5 minAcceptable staleness for non-personalized
Rate limit countersMatch the window (60s, 3600s)Must expire with the window
One-time tokensExact validity periodNo grace period
Computed aggregates1–10 minTrade accuracy for throughput
# Sliding TTL for sessions — reset on every access
def get_session(session_id: str) -> dict | None:
    key = f"session:{session_id}"
    data = r.get(key)
    if data:
        r.expire(key, 1800)  # extend on access
        return json.loads(data)
    return None

Cache Stampede Prevention

When a popular key expires, many requests hit the DB simultaneously.

Probabilistic Early Recomputation (XFetch)

import math, random, time

def fetch_with_xfetch(key: str, loader: Callable[[], T], ttl: int, beta: float = 1.0) -> T:
    cached_raw = r.get(key)
    if cached_raw:
        entry = json.loads(cached_raw)
        delta = entry["compute_time"]
        remaining_ttl = r.ttl(key)
        # probabilistically recompute before expiry
        if remaining_ttl - beta * delta * math.log(random.random()) < 0:
            cached_raw = None  # trigger recompute
        else:
            return entry["value"]

    start = time.monotonic()
    value = loader()
    compute_time = time.monotonic() - start
    r.setex(key, ttl, json.dumps({"value": value, "compute_time": compute_time}, default=str))
    return value

Mutex Lock (Simpler)

import time

def get_with_lock(key: str, loader: Callable[[], T], ttl: int) -> T:
    cached = r.get(key)
    if cached:
        return json.loads(cached)

    lock_key = f"{key}:lock"
    acquired = r.set(lock_key, "1", nx=True, ex=10)  # 10s lock timeout

    if acquired:
        try:
            value = loader()
            r.setex(key, ttl, json.dumps(value, default=str))
            return value
        finally:
            r.delete(lock_key)
    else:
        # Wait and retry — another worker is computing
        time.sleep(0.1)
        return get_with_lock(key, loader, ttl)

Redis Data Structures

# String — simple values, counters
r.set("config:feature_x", "enabled")
r.incr("counter:api_calls:2025-06-01")

# Hash — object fields (avoids full serialization for partial updates)
r.hset("user:123", mapping={"name": "Alice", "plan": "pro"})
r.hget("user:123", "plan")
r.hgetall("user:123")

# Set — membership, deduplication
r.sadd("online_users", "user:123", "user:456")
r.sismember("online_users", "user:123")

# Sorted Set — leaderboards, rate limiting with sliding window
r.zadd("leaderboard", {"user:123": 1500, "user:456": 2000})
r.zrevrange("leaderboard", 0, 9, withscores=True)  # top 10

# List — queues, recent activity
r.lpush("recent:user:123", "order:789")
r.ltrim("recent:user:123", 0, 49)  # keep last 50

# Stream — event log with consumer groups (lightweight Kafka alternative)
r.xadd("events:orders", {"event_type": "placed", "order_id": "abc"})

Invalidation Strategies

# 1. TTL expiry — simplest, eventual consistency
r.setex(key, 300, value)

# 2. Explicit delete on write — strong consistency
def update_user(user_id: str, data: dict):
    db.update(User, user_id, data)
    r.delete(f"user:profile:{user_id}")  # invalidate immediately
    r.delete(f"user:orders:{user_id}:*")  # careful: KEYS is O(N), use SCAN

# 3. Tag-based invalidation — invalidate groups of keys
def set_with_tag(key: str, value: any, tag: str, ttl: int):
    r.setex(key, ttl, json.dumps(value))
    r.sadd(f"tag:{tag}", key)
    r.expire(f"tag:{tag}", ttl + 60)

def invalidate_tag(tag: str):
    keys = r.smembers(f"tag:{tag}")
    if keys:
        r.delete(*keys)
    r.delete(f"tag:{tag}")

# 4. Cache-aside with versioning — no explicit invalidation needed
def versioned_key(entity: str, entity_id: str) -> str:
    version = r.get(f"version:{entity}:{entity_id}") or "0"
    return f"{entity}:{entity_id}:v{version}"

def invalidate(entity: str, entity_id: str):
    r.incr(f"version:{entity}:{entity_id}")  # old keys naturally expire

Rate Limiting with Redis

# Sliding window counter
def is_rate_limited(user_id: str, limit: int = 100, window: int = 60) -> bool:
    key = f"ratelimit:{user_id}"
    now = time.time()
    window_start = now - window

    pipe = r.pipeline()
    pipe.zremrangebyscore(key, 0, window_start)  # remove old entries
    pipe.zadd(key, {str(now): now})              # add current request
    pipe.zcard(key)                              # count in window
    pipe.expire(key, window)
    results = pipe.execute()

    return results[2] > limit

# Token bucket (alternative — smoother bursting)
def consume_token(key: str, rate: float, capacity: int) -> bool:
    lua = """
    local tokens = tonumber(redis.call('GET', KEYS[1])) or tonumber(ARGV[2])
    local last = tonumber(redis.call('GET', KEYS[2])) or tonumber(ARGV[3])
    local now = tonumber(ARGV[3])
    local rate = tonumber(ARGV[1])
    local capacity = tonumber(ARGV[2])
    tokens = math.min(capacity, tokens + (now - last) * rate)
    if tokens >= 1 then
        redis.call('SET', KEYS[1], tokens - 1)
        redis.call('SET', KEYS[2], now)
        return 1
    end
    return 0
    """
    # Use redis.eval() with Lua for atomic token bucket

HTTP Caching

Cache-Control Headers

# Static assets — long cache, versioned URLs
Cache-Control: public, max-age=31536000, immutable   # 1 year; URL changes on update

# API responses — CDN-cacheable, short TTL
Cache-Control: public, max-age=60, s-maxage=300      # browser 1min, CDN 5min

# Authenticated API responses — never CDN-cache
Cache-Control: private, max-age=0, must-revalidate

# Never cache
Cache-Control: no-store

# Revalidate with ETag
Cache-Control: no-cache                               # always revalidate; use ETag
ETag: "abc123"

# Vary header — CDN stores separate copies per value
Vary: Accept-Encoding, Accept-Language

ETags and Conditional Requests

from hashlib import md5
from flask import request, jsonify, make_response

@app.get("/api/products/<product_id>")
def get_product(product_id: str):
    product = get_product_from_db(product_id)
    etag = md5(json.dumps(product, sort_keys=True).encode()).hexdigest()

    if request.headers.get("If-None-Match") == etag:
        return "", 304  # Not Modified — no body, saves bandwidth

    response = make_response(jsonify(product))
    response.headers["ETag"] = etag
    response.headers["Cache-Control"] = "public, max-age=60"
    return response

Distributed Cache Pitfalls

# 1. Cache penetration — repeated misses for non-existent keys
Solution: cache null/"not found" with short TTL (30–60s)
r.setex(key, 60, json.dumps(None))

# 2. Cache avalanche — many keys expire simultaneously
Solution: add jitter to TTL
ttl = base_ttl + random.randint(0, base_ttl // 10)

# 3. Hot key — single key receiving disproportionate traffic
Solution: local in-process cache as L1, Redis as L2
from functools import lru_cache
@lru_cache(maxsize=1000)
def get_config(key: str): ...  # millisecond in-process cache

# 4. Large values — serializing/deserializing huge objects
Solution: store field-level with Redis Hash; never cache full result sets > 1MB

# 5. Stale reads after failover
Solution: use Redis Sentinel or Cluster; never rely on single-node without replication

See also: performance, database-design, api-design

Red Flags

  • Cache stampede on simultaneous key expiry — all requests hit the DB at once when a hot key expires; use probabilistic early expiry, a distributed lock, or staggered TTLs to prevent the pile-on
  • No TTL on cached values — keys accumulate indefinitely and consume memory; every cached value must have an expiry unless explicitly justified as permanent
  • Missing ownership context in cache keys — a key without tenant or user ID can serve one user's data to another; always include the ownership scope in every cache key
  • Write-through without invalidating on write failure — a failed DB write while the cache shows success creates a stale-read window; invalidate the cache key on any write failure
  • In-process LRU cache in a multi-worker service — forked workers maintain separate memory; a cache write in one worker is invisible to others; use Redis for cross-process sharing
  • Cache-Control: no-store on versioned static assets — disabling caching on content-hashed JS/CSS/images forces a full download on every page load; use max-age=31536000, immutable for versioned assets
  • Caching at the wrong layer — caching computed aggregates that are rarely requested wastes memory; cache at the layer closest to the hot query, and measure hit rates before adding any new cache

Checklist

  • Cache keys follow <service>:<entity>:<id> namespace convention
  • TTL values justified per data type — not a single global default
  • Cache-aside pattern implemented; null results cached with short TTL (prevents cache penetration)
  • TTL jitter applied to prevent cache avalanche on mass expiry
  • High-traffic keys protected against stampede (mutex lock or XFetch)
  • Invalidation strategy defined: TTL only, explicit delete on write, or versioned keys
  • Sensitive data (auth tokens, PII) uses private Cache-Control or not cached at all
  • Static assets served with long max-age + immutable + content-hashed URLs
  • Rate limiters use atomic Redis operations (Lua scripts or pipeline)
  • Redis connection pooling configured; not creating new connection per request
  • Cache hit rate monitored; eviction policy set (allkeys-lru or volatile-lru)
  • No KEYS * in production — use SCAN for bulk operations

Keep looking

Skills are one crate of 328,083. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.