Mobile backend infra
Skill almasumdev/awesome-mobile-backend-agent-skills/.github/skills/infra/mobile-backend-infra
Agent skills for the backend-for-mobile layer: APIs, auth, push, sync, and BaaS integrations.
npx -y skills add almasumdev/awesome-mobile-backend-agent-skills --skill mobile-backend-infraAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
2 things to look at
- no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
- 1 stars1 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
What its author says it does
Copied from the file, not written here
Baseline infrastructure for a mobile backend -- edge caching, feature flags, and multi-region strategy. Use when planning or auditing the deployment shape of a mobile backend.
SKILL.md
6.0 KB, as published. Nobody here has run it
Mobile Backend Infra
Instructions
Mobile traffic is globally distributed, latency-sensitive, and correlated with user behavior (morning/evening peaks). The infra needs a CDN in front, feature flags everywhere, and a sober region strategy.
1. Edge Caching
Put a CDN in front of every public endpoint, even JSON APIs.
- Static/media: long
Cache-Control: public, max-age=31536000, immutablewith content-hashed URLs. - Cacheable reads (feeds, listings): short TTLs (10–60 s) with
Cache-Control: public, max-age=30, stale-while-revalidate=60-- reduces thundering herd when a release hits. - Authenticated but cacheable per-user:
Cache-Control: private, max-age=30; keep out of the shared CDN tier.
Key the CDN cache on Accept, Accept-Language, and the major version (/v1). Avoid caching by Authorization (cardinality explosion).
Purge by URL or by tag (Fastly surrogate keys, Cloudflare cache tags) on mutation events:
await cdn.purgeByTag(`article:${id}`);
2. Feature Flags
Every user-visible feature ships behind a flag. Flags are evaluated at two layers:
- Server-side (authoritative): mobile clients receive the computed flag value in
GET /v1/configor in session tokens. - Client-side (optional): the flag SDK re-evaluates locally for offline usage; the server value wins on reconnect.
GET /v1/config
{
"flags": {
"home_v2": true,
"checkout_stripe_intent": true,
"ai_suggestions": { "enabled": true, "model": "v3" }
},
"refresh_after": 1713522660
}
Principles:
- Kill switch per critical path. A bad flag must roll back in < 60 s globally.
- Ramp gradually: 1% -> 10% -> 50% -> 100%, with automatic guardrails (error rate, p95).
- Flag values are cached with short TTL on device (e.g., 5 min) so a mid-session switch does not break UX.
Vendors: LaunchDarkly, Statsig, ConfigCat; or self-host with a simple config service + pubsub.
3. Region Strategy
Three common postures:
A. Single region. Cheapest and simplest. Acceptable for early-stage apps with latency-tolerant use cases (< 300 ms global TTFB is achievable with edge caching + HTTP/3).
B. Multi-region read, single-region write. Reads served from the nearest region, writes forwarded to the primary. Fits most apps:
[us-east-1] (primary writes) <---replication--- [eu-west-1] [ap-south-1]
^ ^ ^
+----- mobile write ----+ | |
| v v
BFF/edge -----------> nearest read region
Replication: Postgres logical replication, Aurora Global, Cloud Spanner, DynamoDB global tables -- pick the one that fits your datastore.
C. Active-active multi-region. Complex. Needed for regulatory locality (EU data stays in EU) or low-latency writes globally. Budget for conflict resolution, global id generation (ULID / Snowflake), and a year of operational learning.
4. Pod and Cluster Sizing
- API pods: CPU-bound usually. Start at 2 vCPU / 2 GiB; scale out horizontally. HPA on CPU + request-per-second.
- Realtime nodes: memory-bound. See
realtime-scaling. - Background workers: separate cluster / nodepool; noisy neighbors are common.
- Use a PodDisruptionBudget to keep at least (N-1) pods up during rollouts.
5. Networking
- TLS 1.2+ everywhere; TLS 1.3 where possible.
- HTTP/2 at the LB; HTTP/3 where the CDN supports it -- big gains on lossy cellular networks.
- Timeouts set at every layer: LB > ingress > pod > downstream; each smaller than the one above so upstream can retry intelligently.
6. Observability Stack
- Tracing: OpenTelemetry; propagate
traceparentfrom the mobile SDK through the stack. - Logs: structured JSON with
request_id,user_id(hashed),endpoint,latency_ms,status,error.code. - Metrics: RED (rate, errors, duration) per endpoint; saturation for pools and queues.
- SLOs: e.g., reads p95 < 200 ms from mobile edge, error rate < 0.5%. Page on SLO burn rate, not on every error.
7. Secrets and Config
- Secrets in a manager (AWS Secrets Manager, GCP Secret Manager, Vault). No secrets in env vars baked into images.
- Config separated from code; feature flags above; environment-specific values in a config service.
8. Cost Controls
Mobile backends easily spend > 50% of budget on egress and object storage. Watch:
- Egress by endpoint / region; cache miss storms on deploy day.
- Object storage lifecycle rules: transition originals to cheaper tiers after 90 days.
- Realtime fanout growth: linear in connections, quadratic if you misuse channels.
9. Incident Readiness
- Runbooks for the top 20 alerts.
- Feature-flag kill switches for every risky path.
- Regional failover drills quarterly.
- A rollback-first culture: the mobile app pins a minimum supported backend version; breaking a contract means roll forward a fix, not "force update".
10. Security Baseline
- WAF at the edge (managed rule set + your own for auth endpoints).
- Bot detection for public endpoints (login, signup, forgot-password).
- Attestation (App Attest, Play Integrity) bumps trust for critical flows.
- Regular dependency and image scans; patch cadence documented.
Checklist
- CDN in front with documented cache key and TTLs per endpoint class.
- Feature flags for every risky feature with kill-switch ramp-down paths.
- Region posture chosen and documented.
- HTTP/2+ and TLS 1.2+ enforced; HTTP/3 where supported.
- OpenTelemetry tracing propagated from client through downstream services.
- SLOs defined per endpoint class with burn-rate alerting.
- Secrets in a manager; config decoupled from images.
- Egress and object storage lifecycle under cost monitoring.
- WAF + bot detection on public endpoints.
- Runbooks and quarterly failover drills in place.