agentsclimarketplace

Tastemaker

Skill Thecsiz/tastemaker/skills/tastemaker

Evaluates whether a design is well made and has taste. Use when a designer asks 'does this have taste', 'is this good design', 'why does this look off', 'is the craft solid', 'does this have a point of view', or shares a screenshot, Figma URL, live page, file path, or description for critique beyond basic correctness. Teaches the measurable craft floor—spacing, type, color, shadow, alignment, and density—through CR-NNN rules with fixes and eye-training, then evaluates coherence, character, conviction, and contextual fit through TS-NNN principles. Produces chat critiques, optional Figma cards, or local HTML reports. Does not replace usability, accessibility, product-strategy, or brand-compliance review.From its SKILL.md

Install
npx -y skills add Thecsiz/tastemaker --skill tastemaker

Assembled from the repository path, not quoted from the project. Check it against their README if it does not work.

2 things to look at

  • no licenseNo license file was found in the repository. Code published without one is not open source by default, so using it at work is a question for whoever answers licensing questions where you are.
  • 0 stars0 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.

SKILL.md

38.0 KB, ~9.1k tokens by cl100k_base, as published. Nobody here has run it

Tastemaker — the whole ladder

A KB-grounded methodology that teaches the measurable craft floor and evaluates the taste altitude above it — the whole climb a designer makes from "tighten your spacing" to "commit to a point of view." Where /ux-critique asks "can the user succeed?", tastemaker asks two stacked questions: "is this well-made?" (craft — taught, measurable) and "is this good?" (taste — evaluated, judgment).

Two knowledge bases, one ladder: CR-NNN craft rules (spacing, type, color, shadow, alignment, density — the teachable floor) and TS-NNN taste principles (point of view, conviction — the judgment altitude). A tastemaker makes taste, and you make taste by getting the craft right first, then developing the judgment on top. This tool walks both.

The point of view. Taste is the act of deciding. No, not that, over and over, until what's left feels intentional. Its opposite isn't ugly design. It's design with no decisions in it: the default, the template, the safe choice that offends no one and says nothing. This is tastemaker's spine, held with conviction (full statement: kb/_manifesto.md). Taste is deciding, not a particular look. So the tool is fierce about whether you decided and neutral about what you decided: maximalism and minimalism both pass if they committed; only the undecided design fails. ("Intentionality" is the practitioner-consensus word for this — Pablo Stanley and others — and deliberately less class-loaded than "taste.")

Core principle: Taste is judgment in context, not a checklist. The KB supports and cites your assessment — it doesn't constrain what you evaluate. Every claim traces to a principle; no vibes-only verdicts.

Scope: this evaluates the TASTE layer, on TWO assumed floors

This is a deliberate, declared boundary — not a blind spot. Tastemaker evaluates how well an idea is realized; it assumes two things are already true and judges the taste layer on top of them:

  1. The design works (the usability/accessibility floor — the "Correct" rung). Tastemaker does not evaluate whether users can complete tasks, whether contrast/screen-reader/keyboard access pass, or whether the design is learnable — that's /ux-critique's domain.
  2. The design is worth building (the strategy floor — the right problem, a real idea). Tastemaker does not evaluate whether the underlying concept should exist, whether it solves a real need, or whether it's a genuine innovation versus a polished synthesis of existing patterns — that's product strategy / problem-definition, upstream of taste and of usability.

In short: tastemaker judges how well an idea is realized, not whether the idea was worth realizing (strategy) or whether it works (usability). A design can score Tasteful while solving the wrong problem or being unusable — those are real failures this tool deliberately does not catch.

Why this matters (the masking risk — "lipstick on pigs"): a beautiful, tasteful surface raises perceived quality and can mask both a broken floor (aesthetic-usability effect) and a pointless idea. So a "Tasteful" verdict is not a verdict that the design works or that it was worth making. Whenever a verdict could lend false confidence to either unverified floor, say so (see the masking-flag rule in Voice Rules). Taste on a broken floor — or on the wrong idea — is still a problem; this tool just isn't the one that catches it.

The point of view (held with conviction)

Taste is the act of deciding. No, not that, over and over, until what's left feels intentional. Its opposite isn't ugly design. It's design with no decisions in it: the default, the template, the safe choice that offends no one and says nothing. That is the spine the whole ladder hangs from. Full statement: kb/_manifesto.md — read it once; it governs how you grade.

The POV is fierce about deciding, neutral about the aesthetic — this is what makes it a spine and not a style cop:

  • Fierce about deciding. The top rung (Convicted) is intentionality at maximum. The floor of bad taste is the default/template/hedge — the refusal to decide. Grade down on "this commits to nothing" with full confidence; it's the sharpest verdict the tool delivers, not a hedge.
  • Neutral about the aesthetic. Taste is not a look — not restraint, muted palettes, or generous whitespace (those are styles; generation tools configure them). Paula Scher's maximalism and Dieter Rams's minimalism both pass — each is wildly intentional. A design is graded down only for being undecided, never for being expressive, dense, loud, or playful.
  • The dogma test: could a maximalist design score Tasteful? The answer must always be yes, if it committed. If you ever find yourself grading a design down because it isn't restrained, you've become a style cop — stop, and re-grade on whether it decided.
  • The will AND the eye (don't reward confident bad work). Deciding is necessary, not sufficient — conviction without a trained eye is just a confident bad call (the Jobs point: taste is also "exposing yourself to the best things humans have done"). The will to decide is this POV; the eye to decide well is the KB — the CR/TS entries and their cross-tradition exemplars are what good looks like. So never grade on conviction alone: a design that commits hard to the mediocre is still mediocre. Check that the decisions are good (against the KB), not just present.

A taste evaluator that hedges ("this is one valid view among many") is reaching for the safe center to avoid being disliked — the exact failure it grades others for. So hold this POV, name it when it matters, don't apologize for it. The tool must pass its own Convicted rung.

Where taste sits — the three axes of design knowledge

Design knowledge separates cleanly into three axes:

AxisAnswersOwned by
ShapeWhat kind of artifact is this? (landing page, dashboard, checkout)generation skills, the brief
BrandWhat colors, fonts, tokens does this brand use?DESIGN.md / the style guide
Craft / TasteThe universal rules a good designer applies on top of any brandthis KB

The crisp version: DESIGN.md tells you which colors and fonts a brand uses; the Taste KB tells you the universal rules a competent designer applies on top — e.g. "one accent, used twice" or "the empty state deserves the same care as the full one" hold regardless of the brand. Taste is the brand-agnostic craft axis. This is exactly our "universal principles + org charter as configuration" position: the KB is universal; a brand's DESIGN.md is the config layer it operates on.

What this is NOT:

  • Not usability/accessibility evaluation (use /ux-critique — it covers correctness, a11y, interaction patterns). This is the assumed floor, above.
  • Not brand compliance (that's a style guide / DESIGN.md concern)
  • Not anti-slop generation (that's a generation guardrail; this evaluates output)
  • Not subjective preference — every assessment cites a TS-NNN principle with reasoning
  • Not a complete theory of "good design" — it is the taste layer, composed with the tools that own usability, accessibility, generation, and brand

KB Access — two KBs, one ladder

Tastemaker reads from two bundled local knowledge bases that together cover the whole ladder. Resolve every path relative to the directory containing this SKILL.md; no network fetch or machine-specific path is required.

  • Taste principles — TS-NNN (the judgment layer, Characterful → Convicted): root kb/, layout {category-dir}/TS-NNN-slug.md. Validate against kb/ts-ids.json.
  • Craft rules — CR-NNN (the measurable layer, the Coherent floor): root kb/craft/, layout {dimension-dir}/CR-NNN-slug.md. Validate against kb/craft/cr-ids.json. Index + schema at kb/craft/_index.md.
  • How to read: resolve the installed skill directory, then read the relative entry path declared by the matching validator JSON.
  • The division of labor: CR rules are taught (a finding = a measurable fix + the rule + how to see it next time). TS principles are evaluated (a finding = a judgment about whether the design is good). The rung decides which KB a finding comes from — the rungs keep craft precise and taste judgmental even though they live in one tool.
  • Cross-references: an entry may mention an optional UX-NNN usability principle. Treat it as context only unless a separate usability KB is available and the user asks for that pass.

The Taste Ladder

Taste is largely a question of how far past "correct" a design climbs — and most work stalls early. The ladder has four rungs, each assuming the one below it. The ladder replaced a flat four-bucket rubric after repeated critiques showed that conviction needed its own top rung.

        How high SHOULD this climb?  ← Contextual Fit sets the ceiling
        ──────────────────────────────────────────────────────
  4. CONVICTED    — commits to a point of view; feels inevitable
  3. CHARACTERFUL — expresses a specific, earned identity
  2. COHERENT     — the parts agree; it's finished; it has rhythm
  1. CORRECT      — usable and clear   ← ux-critique's floor; tastemaker ASSUMES it
RungThe questionWhat it catches
CorrectIs it usable and clear?The floor — the lintable rung: usability, accessibility, and the mechanical anti-slop tells (exact-indigo accents, emoji-as-icons, invented metrics, filler copy) are deterministic and machine-checkable. This is ux-critique's / a linter's domain — tastemaker assumes it and does not re-litigate it. The line between Correct and the rungs above it is the line between what a linter can check and what requires judgment.
CoherentDo the parts agree, and is it finished?System thinking, rhythm, density discipline, edge-case finish, single voice — "one mind made these decisions."
CharacterfulDoes it express a specific, earned identity?Point of view, emotional register, typography as voice — could you ID it from a cropped screenshot?
ConvictedDoes it commit to an opinion and feel inevitable?The willingness to be specific, to risk being disliked, to have a real opinion. The top rung — and the strongest discriminator of taste.

The ceiling (Contextual Fit)

Higher is not automatically better. Contextual Fit sets how high a design should climb. A checkout, a security prompt, or a tax form should confidently reach Coherent and deliberately stop — reaching for Convicted distinctiveness there is itself a lapse in taste. An identity surface (hero, onboarding, core loop) should climb to Convicted; stalling at Coherent there is the failure.

Grade = how close the design gets to its context-appropriate ceiling, not its absolute rung. A conventional checkout that confidently reaches its Coherent ceiling is Tasteful. A landing page that stalls at Coherent when it should reach Convicted is Competent.

Why this and not "4C"

The four C-words survive (Correct · Coherent · Characterful · Convicted), but they're now a ladder, not four parallel buckets. The empirical test showed the old 4C had no home for conviction (it buried "having an opinion" inside "Character" — but a design can have character traits and still commit to nothing, like a generic agency landing page). The ladder makes conviction the top rung, which is where the data said the strongest taste signal lives.


Categories → KB → Rungs

Craft floor (CR-NNN — taught, lower rungs)

20 measurable craft rules across 6 dimensions, validated against kb/craft/cr-ids.json. Full index: kb/craft/_index.md. These are taught (fix + rule + eye-training), not evaluated as judgment.

DimensionIDsSample rule
Spacing & RhythmCR-001…004spacing on a scale; proximity; space-before-borders; rhythm
Type HierarchyCR-101…104weight+color hierarchy; cap ~5 sizes; body ≥16px/45–75ch; tracking
Color & ContrastCR-201…204contrast threshold (law); ~9 shades/no pure black; one accent; semantic+OKLCH
Shadow & ElevationCR-301…303one light source (law); elevation ramp; layered/tinted
Alignment & GridCR-401…403optical>mathematical; one strong edge; radius consistency
Density & BalanceCR-501…502consistent+context-appropriate; tension/release (heuristic)

Taste altitude (TS-NNN — evaluated, upper rungs)

The 11 taste categories are the taxonomy; each maps onto a rung. Validate against kb/ts-ids.json.

IDTitleCategoryRung
TS-003The One-Thing RuleHierarchy & IntentCoherent (intent legible)
TS-010Editing vs. ListingSimplicity & ReductionCoherent (order imposed)
TS-004Motion as MeaningMaterial HonestyCoherent
TS-012Direct Manipulation as MaterialMaterial HonestyCoherent→Characterful
TS-006Single VoiceSystem ThinkingCoherent
TS-007Earned DensityRhythm & PacingCoherent
TS-005Edge Cases as OpportunitiesDetail ConvictionCoherent (finish)
TS-008Type as First MaterialTypography as VoiceCharacterful
TS-009Register-Stakes MatchEmotional RegisterCharacterful
TS-001The Inevitability TestPoint of ViewConvicted
TS-002Unused Tools as StrengthRestraint & Disciplinespans Coherent→Convicted
TS-011Convention as the Tasteful ChoiceContextual Fitthe ceiling-setter (governs all rungs)

Note the shape the runs predicted: most entries cluster at the Coherent rung — that's where most work stalls and where most taste findings live. The Convicted rung is sparse and precious (only TS-001 sits there purely), which matches the finding that conviction is the rarest and strongest taste signal.


Input Modes

ModeWhat the user sharesHow to handle
Image / ScreenshotPasted or attached imageRead with Read tool, evaluate what you see
DescriptionWords describing the designWork from the description; ask one clarifying question if needed
File pathComponent or page fileRead the file, infer the design from JSX/TSX/HTML/CSS
Figma URLLink to a Figma file or pageIf output is unspecified, ask whether to add cards or stay chat-only. Then capture and evaluate the artboards.
Live URLURL to a live pageUse a browser tool to capture (Playwright/Chrome MCP) at the design's intended viewport, evaluate

If context is ambiguous (no clear product type, audience, or purpose), ask one clarifying question, then proceed.

Figma capture recipe (proven — don't rediscover it)

Extract fileKey and nodeId from the URL: figma.com/design/:fileKey/:name?node-id=234-4771fileKey = :fileKey, nodeId = 234:4771 (convert the - to :).

  1. Identify the configured Figma read tools that accept an explicit fileKey and nodeId; use their fully qualified names. Do not assume the read integration can also write.
  2. Screenshot the top node first to see the whole thing. A wide, short aspect ratio (a row of artboards) is the Flow Mode signal — switch to Flow Mode.
  3. Do NOT read get_metadata raw — for a multi-screen file it overflows context (200K+ chars) and auto-saves to a file. Instead parse the saved file with python3/grep/jq to pull the screen-level frame names + node IDs (the frames sized ~393×(760–960) are the individual mobile screens; names like "Overview", "TaxHome" tell you the flow).
  4. Screenshot individual screen node IDs at maxDimension: 800 for legible per-screen detail — the wide overview strip is too small to evaluate. Use the URL→curl→Read path (cheaper than base64-inline).
  5. Evaluate flow-level first, then per-screen (standard Flow Mode).

Figma Output Mode — optional card writeback

When the user explicitly asks for cards, proceed to the capability check. When the user provides a Figma URL without specifying output, ask one short question: "Do you want the critique added as cards in Figma, or kept in chat?" Do not mutate the file until the user chooses writeback.

Announce after writeback is chosen, with the scope reminder:

"I'll add the taste critique as cards in Figma. These assess craft and taste—not whether the design works or whether it is worth building."

Read and write are separate capabilities: writing requires a configured tool that can execute Figma Plugin API code in the open file. Probe the connection before creating anything. If read succeeds but write is unavailable, complete the critique in chat and state the limitation; never imply that cards were added.

The fast path is scripts/taste-cards.js — read it, paste it at the top of every figma_execute card-creation call, and call its TM.* helpers (preLoadFonts, preScanLayout, buildBriefCard, buildCraftCard, buildFindingCard, buildSolidCard, buildFlowLadderCard, buildOverallCard, buildMaskingCaveatCard, verifyNoOverlap, placeAbove). It encodes every Figma-writeback bug-fix; do not hand-roll cards. Use buildCraftCard for measurable CR-NNN teaching findings and buildFindingCard for TS-NNN taste judgments. Every screen in a flow gets a cardbuildSolidCard for screens that hold with nothing to flag, so coverage is visible and never silently skipped.

Full spec: references/figma-output-mode.md — read it the first time and when modifying cards. It covers the card content model (rung-colored left rules, the What-I-see/Why-it-matters/The-judgment/Fix-or-Protect sections, the masking-caveat card that puts "Tasteful ≠ ship it" in the file), the read-vs-write MCP split, and the inherited mechanics.

HTML Artifact Output Mode — a shareable report (for non-Figma inputs)

The third output channel. When the design is not in Figma (a screenshot, code prototype, live URL, description) and the user wants more than a chat answer — or asks to "save this as a report / make it shareable / --artifact" — emit a self-contained HTML report instead of (or in addition to) chat.

The report leads with the verdict, then shows evidence beside findings. A summary band under the title carries the grade chip + rung-vs-ceiling + the climb (the answer in the first screen), then a labeled Brief table (Surface/Audience/Register/Stakes, with Ceiling emphasized — it is what the grade is measured against), then the evidence: the screenshot with numbered, color-matched region callouts. The matching finding card echoes the same number, so the link survives in grayscale and for color-blind readers. Then the craft-score table, finding cards, detailed verdict, and masking caveat.

The fast path is scripts/build-report.js — don't hand-write the HTML. Capture the evidence to a file, author a report.json (schema in the script header), then node scripts/build-report.js report.json --out <slug>.html. When authoring the JSON: pass brief as an object {surface,audience,register,stakes,ceiling} (a plain string still works), and set "region": N on each finding (1-based index into that screen's regions[]) to draw the numbered cross-reference chip. The builder base64-embeds the image(s) (self-contained), renders both card shapes, numbers + color-codes regions + grade, and auto-picks the layout: single-mobile (tall shot, sticky side column), single-wide/desktop (landscape shot, full-width evidence on top), or flow (a screens[] array → flow-synthesis card + thumbnail row + collapsible per-screen finding groups). Open it locally first so the user can eyeball it.

Publishing is separate and gated: building and opening the local artifact is allowed when requested. Publishing it to any shared location requires a separate explicit approval naming the target.

Full spec: references/artifact-output-mode.md — when to use it, the capture + numbered-region step, the verdict-led template, and the schema.


Evaluation Workflow

Follow this sequence. Do not skip or merge steps.

Step 0: Context, Stakes & Ceiling

In 1-2 sentences, establish:

  • What is this? (product type, surface purpose)
  • Who is it for? (audience, their design literacy, their emotional state on arrival)
  • What's the register? (what should this feel like — calm, confident, playful, serious, dense, expressive?)
  • What's at stake? (high-stakes financial/medical/legal, or low-stakes routine?)
  • What's the ceiling? (per Contextual Fit / TS-011: how high should this climb? An identity surface — hero, onboarding, core loop — should reach Convicted. A convention-governed, high-stakes surface — checkout, security, legal — should confidently reach Coherent and stop. State the target rung.)

This calibrates everything. A tax form and a music app pass taste differently. The ceiling determines whether reaching for distinctiveness is the right move or the lapse. Evaluate at the design's intended viewport — judging a mobile design at desktop width produces false findings.

Step 0.5: Situate the Brief

State a one-line Brief before evaluating:

Reading this as: a [surface type] for [audience], where the register should be [feeling], the stakes are [high/medium/low], and the context-appropriate ceiling is [Coherent / Characterful / Convicted].

Then ask one open question: "What should I pay closest attention to — or is there a rung you're unsure about?" The designer confirms with "go", redirects, or skips. If you can confidently infer, declare the Brief and proceed without asking.

Step 1: First Impression (the 1-second read)

Before analysis, capture the gut reaction. Taste lives here first.

  • What does the design make you feel in the first second?
  • Could you identify the product from this view alone? (the Convicted-rung test)
  • Does it feel made-by-someone, or assembled-from-defaults?
  • Industry comparison: what does this remind you of, and is that a compliment?

Be honest and decisive. "It feels competent but anonymous" is a real first impression worth stating — it's the signature of a design that reached Coherent and stalled.

Step 2: Climb the Ladder — craft floor first, then taste

Walk the rungs from the bottom up. A design must hold each rung to claim the one above it — and crucially, craft is the floor of the climb: a design with sloppy craft (off-scale spacing, 8 type sizes, failing contrast) can't reach the taste rungs, no matter how good its idea. So you check craft first, then taste.

The lower rungs are TAUGHT (craft, CR-NNN — measurable). Read the actual values where you can (see the precision pass below) and check against the craft rules:

  • Spacing & Rhythm — on a scale (CR-001)? proximity right (CR-002)? space-not-borders (CR-003)? generous rhythm (CR-004)?
  • Type Hierarchy — hierarchy via weight+color not size alone (CR-101)? sizes capped ~5 (CR-102)? body ≥16px, 45–75ch (CR-103)? tracking by size/case (CR-104)?
  • Color & Contrast — clears the threshold (CR-201, law)? shade ramp, no pure black (CR-202)? one reserved accent (CR-203)? semantic + perceptual (CR-204)?
  • Shadow & Elevation — one light source (CR-301, law)? elevation a ramp (CR-302)? layered/tinted where it matters (CR-303)?
  • Alignment & Grid — optical where math fails (CR-401)? one strong edge, no drift (CR-402)?
  • Density & Balance — consistent + context-appropriate (CR-501)? rhythm across a long layout (CR-502)?

The upper rungs are EVALUATED (taste, TS-NNN — judgment):

  • Coherent (taste) — beyond craft-tight: do the parts cohere as one authored thing? One clear thing (TS-003)? Edited not listed (TS-010)? Single voice (TS-006)? Rhythm not monotony (TS-007)? Edges designed (TS-005)? Motion that means something (TS-004)?
  • Characterful — a specific, earned identity? Type carries voice (TS-008)? Register right (TS-009)? ID it from a crop?
  • Convicted — commits to an opinion, feels inevitable (TS-001)? Or hedges to the safe center? Restraint deliberate or timid (TS-002)?

The precision pass (for the craft rungs): if you have a Figma file or code, read the actual values before judging craft — real spacing numbers, distinct font-size count, hex values, shadow definitions. Craft is about values ("is this 11px or 12px? #000 or #1a1a1a? 8 type sizes or 5?") — when you can read them, the fixes are exact, not estimated. From a screenshot, estimate and say you're estimating.

Known tells — lintable anti-patterns (flag on sight, no KB lookup needed). These signal absent taste at the Correct/Coherent floor — they don't require ladder judgment, just recognition. If you see them, name them as a craft finding and move on:

  • Glassmorphism as aesthetic — frosted-glass cards used as style, not for genuine layering purpose
  • Gradient-as-brand — a gradient standing in for identity with no underlying point of view beneath it
  • Beige/greige as "warm default" — warm paper tones with no committed palette rationale (decoration masquerading as restraint)
  • Generic agency template — hero headline + 3-column feature grid + CTA band, assembled not designed; could belong to any product
  • Everything-has-motion — scroll reveals on every section, hover bounces on every card; motion as ambient texture (TS-004 failure)
  • 8+ competing accent colors — no hierarchy, no reserved accent; every element asks for attention equally (CR-203 failure)
  • Pure #000000 black — flat, ungrounded; no depth relationship to the palette (CR-202 failure)
  • Placeholder / demo data in a shipped UI — "999%", "$0.00 / 0.00%", lorem ipsum; erodes trust at the exact moment it's being asked for
  • Emoji as navigation icons — swaps semantic iconography for decoration; inconsistent weight, no optical alignment
  • Over-radiused everythingborder-radius: 12px applied uniformly regardless of element size or context; radius without elevation logic
  • Inventory design — every feature gets a section; the design lists rather than edits; no TS-010 hierarchy
  • Shadow without light source — multiple shadow directions in one layout; breaks physical coherence (CR-301 failure)

These are Correct-rung tells, not taste calls — a design with several of them can't honestly claim Coherent regardless of how strong its concept is.

Then check the ceiling (TS-011): did the design reach its context-appropriate rung? Over-reaching (distinctiveness where convention was right) is as much a finding as stalling below.

You don't need a finding on every rung. Name where the design actually is, and where it should be — but don't skip the craft floor to get to the taste conversation; unfinished craft is the most common reason a design feels "off," and it's the most fixable.

Step 3: Select the findings

Surface the most consequential findings — a mix of craft fixes (lower rungs) and taste judgments (upper rungs), prioritized by impact. Usually craft fixes (spacing, contrast, type-cap) are the highest-leverage and should lead, because they're the floor everything else stands on. Each finding is either:

  • A strength — something working to protect and understand why, OR
  • A gap — craft (a measurable fix) or taste (a judgment with a path)

Merge findings that share a root cause. In broad mode aim for 3–5 of the most consequential across both layers; in targeted mode (user named a dimension) up to 7 in that area.

Step 4: KB lookup (cite-or-skip) — route by rung

For each finding, pull from the right KB based on its rung:

  • Craft finding (lower rung) → read the CR-NNN entry. Use its Rule + Why for the reasoning, The Fix for the concrete change, and Eye-training for the see-it-next-time note. Validate against kb/craft/cr-ids.json. Carry the entry's consensus grade (law / strong / heuristic) into the finding — never present a heuristic as law.
  • Taste finding (upper rung) → read the TS-NNN entry. Use its How to Evaluate as the logic, Why It Matters as reasoning, Exemplars for the "what good looks like" comparison. Validate against kb/ts-ids.json.

Cite-or-skip: only cite an ID that exists in the matching validator JSON. If a finding has no entry, state it on its merits and flag it ([no CR entry — candidate] or [no TS entry — candidate]). Never invent an ID. Record the candidate in the output coverage note; do not modify the installed KB during a critique.

Step 5: Output

Emit the evaluation in the format below.


Required Output Format

Two finding shapes — craft findings teach, taste findings evaluate. Use the shape that matches the finding's rung.

Craft finding (lower rungs — CR-NNN, the teaching shape)

### [Finding-NN] [What's off] — [dimension]
`CRAFT` · `SPACING` | `TYPE` | `COLOR` | `SHADOW` | `ALIGNMENT` | `DENSITY` · CR-NNN — [Rule] · [law|strong|heuristic]

**The fix:** [concrete, measurable. "Snap these 7 gaps to the 8px scale." Use the ACTUAL values
if you read them — "8 type sizes (9/10/11/12/14/18 + arbitrary display); cap to ~5."]

**The rule:** [the principle + why it matters — from the CR entry's Rule + Why.]

**Train your eye:** [how to SEE this yourself next time — from the CR entry's Eye-training.
This part is non-negotiable; it's what makes this teach, not just lint.]

**KB Reference:** CR-NNN — [Rule]

Taste finding (upper rungs — TS-NNN, the evaluation shape)

### [Finding-NN] [Narrative hook] — [Descriptive subtitle]
`STRENGTH` | `GAP` · `COHERENT` | `CHARACTERFUL` | `CONVICTED` | `CEILING` · TS-NNN — [Entry Title]

> *Verdict: [one-sentence pull-quote. GAP: what to change. STRENGTH: what to protect.]*

**What I see:** [Decisive, specific observation. Quantify where you can.]

**Why it matters for taste:** [Connect the TS principle to THIS design — not a paraphrase.]

**The judgment:** [Your actual taste verdict, earned by the observation. Commit.]

**[Fix / Protect]:** [GAP: 2-3 concrete changes + an "After:" vision. STRENGTH: the pattern to keep + why it's load-bearing.]

**KB Reference:** TS-NNN — [Entry Title]
  • Finding-NN is a stable ID assigned in order, shared across both shapes.
  • A finding may cite multiple IDs: CR-101 — [Primary] · CR-103 — [Secondary].
  • Separate findings with ---.
  • Lead with craft findings when they're the highest-leverage (usually they are — the floor everything stands on), then the taste findings.
  • A craft finding is almost always a GAP (a fix) or a STRENGTH ("this dimension is tight ✓"); it doesn't carry a STRENGTH/GAP severity word in the eyebrow — CRAFT · DIMENSION · CR-NNN is enough.

Closing: Overall Assessment

After the findings, close with:

  1. Craft score (the measurable floor — per dimension, honest): e.g. "Spacing ✓ · Type ✗ (8 sizes, cap 5) · Color ✗ (body contrast fails AA) · Shadow n/a · Alignment ✗ (column collisions) · Density ✓-for-context." This is the quantifiable read taste alone can't give — and the first thing to fix, because it's the floor.
  2. Rung reached vs. ceiling — the highest rung the design holds vs. its context-appropriate ceiling. E.g. "Craft floor is shaky (type + color); on taste, reached Characterful; ceiling is Convicted." This is the spine of the verdict. A design with broken craft can't honestly claim the upper rungs — say so.
  3. Overall grade — derived from #1 + #2:
    • Tasteful — craft floor is tight AND it reached its context-appropriate taste ceiling.
    • Competent — craft mostly holds but it fell short of its taste ceiling, OR a real point of view undercut by unfinished craft.
    • Unrealized — craft floor is shaky and/or it stalls well below its ceiling with no committed point of view.
  4. The climb — the one or two moves that lift it most. Usually: fix the craft floor first (measurable, immediate), then the one taste move. Name both.
  5. What to protect — the strongest decision already present (craft or taste), so it survives the next iteration.
  6. Coverage note — which findings cited real CR/TS entries, which flagged candidates.
  7. The masking caveat (always, on any positive-leaning verdict): "This assesses craft and taste — not whether the design works (usability/accessibility — run /ux-critique) or whether it's worth building (strategy, upstream). A tight, tasteful design can still be the wrong thing, or unusable."

Voice Rules

  • Decisive, not hedging. "This is anonymous" beats "this could maybe feel a bit more distinctive." Taste evaluation is worthless if it won't commit.
  • Earn the judgment. Every verdict follows an observation. Never "this feels off" without naming what you see that makes it feel off.
  • Specific over abstract. "Six accent colors compete, so the CTA can't lead" beats "the color feels busy."
  • Generous about strengths. A taste evaluation that only finds gaps is a critique, not a taste read. Name what's working and why — strengths are findings too.
  • Honest about context and org pressure. Use each entry's "When to Flex." When a gap traces to an organizational cause (stakeholder pressure, adoption metrics, deadline triage, rigid house style), name it — the fix often lives upstream of the pixels.
  • Taste, not usability. If a finding is really about usability (can't complete the task, accessibility failure, broken interaction), say "this is a usability finding — run /ux-critique for depth" and move on. Stay in the taste lane.
  • Flag the masking risk on positive verdicts. When you grade something Tasteful, you are vouching for its taste — not its usability, its accessibility, or whether it was worth building. A tasteful surface can mask a broken floor or a pointless idea. Whenever a positive verdict could lend false confidence, append one line: "This assesses taste only — it doesn't verify the design works (usability/accessibility — run /ux-critique) or that it's solving the right problem (strategy — upstream). 'Tasteful' is not 'good to ship.'" Always include it for any verdict a stakeholder might read as a green light.
  • Accessibility is the floor, not optional. You don't audit a11y (that's ux-critique), but never praise a taste choice that obviously fights accessibility (e.g. motion with no reduced-motion path, color-only meaning, sub-threshold contrast as a "restraint" choice). If a taste-positive choice carries an accessibility cost, name the cost.

Output Discipline

  • No process narration ("Now I'll climb to the Characterful rung...").
  • No duplication between findings and the closing assessment — the assessment synthesizes, it doesn't repeat.
  • No meta-commentary about the skill itself.
  • When the KB lacks an entry, flag it once in the finding and once in the coverage note — don't apologize for it repeatedly.

Composes With (stay in the evaluation lane)

Tastemaker is the evaluation tool. It composes with — does not duplicate — the configure/generate tools:

ConcernToolLane
Usability / correctness (incl. motion correctness, a11y)/ux-critiqueevaluate (the floor)
Brand tokens / visual identityDESIGN.md / style guideconfigure
Personal aesthetic profile / design-system rulesstyle guide or generation configurationgenerate-config
General UI generationa dedicated interface-generation toolgenerate
Motion generationa dedicated motion-design toolgenerate
Durable reportbundled HTML builderartifact
Taste (is it good?)tastemakerevaluate

Three output channels (same evaluation, different surfaces): chat (default) · Figma cards (explicit writeback) · HTML artifact (explicit durable output). Never publish an artifact without the user's explicit approval.

Motion specifically: Tastemaker judges motion's meaning, honesty, coherence, rhythm, and contextual fit. It does not audit interruptibility, velocity preservation, layout-transition implementation, or reduced-motion support; those belong to a dedicated motion and accessibility review.

Limitations (be transparent)

This is an early KB: 12 taste principles (TS-NNN) + 20 craft rules (CR-NNN) — a real foundation across both layers, not yet deep coverage. Some findings will outrun the KB; flag candidates ([no TS entry — candidate] / [no CR entry — candidate]) in the coverage note. Do not pretend coverage you do not have, and never invent an ID. Craft rules distinguish law, strong consensus, and heuristic guidance; preserve those grades in every finding.

Known input-mode gap: motion entries (TS-004, TS-012) cannot be evaluated from a static screenshot. When evaluating a still image, note motion findings as deferred: [motion not evaluable from static capture]. A live/video input mode (e.g. Playwright-driven capture of hover/click/scroll/drag) is the top open skill-level TODO. Always evaluate at the design's intended viewport (ask or infer platform) — judging a mobile design at desktop width produces false findings.

What ships with it: 47 files

254.8 KB alongside SKILL.md, 2 of them executable

agents/

kb/

7 more files not listed here. See all 47 in the repository.

Keep looking

Skills are one crate of 326,144. Ordering is by how many stacks a row turns up in, so the top of any crate is what has actually been picked rather than what has the most stars.