Phase 01 · Research and benchmark

What we learned before drawing anything

Phase 1 of 12 · five research stages, signed off 2 Sep 2026 · re-verified 9 Sep 2026 · source brief: CLAUDE.md

Five stages of research for AI Stack Builder — a workspace where an AI practitioner keeps their own building blocks and assembles them into projects that export as a working archive. We surveyed the market, captured how live products behave, went looking for what actually hurts, scored the best of them against a rubric, then compared five shapes for our key flow and chose one. Who the person is and what they hire this for is phase 02personas and JTBD.

Stages
5all closed
Products
15on five axes
Flows
1210 closed
Captures
143screens & logs
Cells scored
154 flows × rubric
Issues read
7863two trackers
START

Overview

Five stages in four days, and the two lists they produced.

What we saw

  • Every live competitor sells to an organisation. Governance, drift, compliance — control over other people’s work. The individual practitioner was unoccupied ground among funded vendors, though open-source tooling has since filled the item-level half of it — and on 10 Sep 2026 four of those tools were installed and pointed at a deliberately broken set. None of them has a set-level unit at all.
  • Trust has moved from social proof to measurement. A composite score, an uplift multiplier, a scan — a number derived from running the thing has replaced stars and downloads. We can run nothing.
  • Nobody authors a graph; everybody derives one. Not one of fifteen products asks a human to draw an edge between two objects.
  • The loudest pain is environmental, not compositional. The archive lands on a machine and does not run — PATH, node version managers, platform paths, processes dying at startup.
  • Our own thesis is real and quiet. Two servers claiming one key, “the chat always chooses the first one specified”: the first wins and nobody is told. Thirteen reactions against 182.
  • Handover is the industry’s weakest flow. Fifteen cells scored, no B4 cell above 4, and not one of them scores what a product says about the machine its artefact lands on. Narrowed on 10 Sep 2026, after four open-source tools were installed and run: one of them does describe the receiving machine. What nobody does is describe it for a named set.

What we decided

  • An item card claims nothing. Usage facts from the user’s own library — used in 3 projects — never a score, rating or badge.
  • Three severities, and nothing blocks. Problem, note, skipped. Export is never disabled; an unclean set is confirmed, with the consequence named in the present tense.
  • Command-first assembly, a run-centric check. Three surfaces plus projects, with the seam at the check. The library leaves the builder.
  • The setup document is written for an agent, not for a human reader — which is what aims at the 182-reaction pain without a platform matrix.
  • A curated public shelf ships with the product, so there is material from the first second while the user’s own library starts honestly empty.
  • Composition and versioning stay out, deferred rather than refused, with the reason written down.

What this phase could not do. It read vendors, artefacts and trackers. It never asked a person anything. That is phase 02 — personas and JTBD — and it changed several things on this page, each marked where it happened.

STAGE 01

Competitors

Who else is in this space, what do they sell, and to whom?

Fifteen products in three groups. Hard — aiming at the same job as ours. Soft — a different product doing the same work somewhere else. Aspirational — the craft bar. Each read on five axes: who it is sold to, what its atomic asset is, the mechanism nothing else has, how it earns trust, and how it charges.

ProductProduct baseKey mechanismTrust
Hard — same product, same audience
TesslSkill as a registry artifact pinned to repo + path + commitPublishing gated by evaluation — a skill ships only after it scoresComposite 93, uplift 1.40×, Quality / Impact %, Snyk scan
PackmindA versioned engineering playbook of standardsOne source distributed into each agent's format, drift detected backGovernance framing, pre-commit blocking
AgentmanSkill as an executable, shareable unitVisual assembly of skills into a working agentCompliance badges + four permission tiers
SmitheryMCP server as a hosted, connectable endpointConnect once, reuse everywhere — it holds auth and sessionsScore /100, verified badge, usage counts
Continue HubA block: model, context, docs, MCP, rule, promptCompose blocks into a custom assistantOpen source and forkable — trust by inspection
Soft — different product, same job
BackstageA typed entity declared in YAML beside the codeRelations derived and filtered, never drawnOwnership: owner, lifecycle, source file on every entity
PortEntity in a Context Lake, plus scorecardsStandards applied as scorecards over the catalogPass, warn or block — with the reason attached
TerraformA versioned module with declared constraintsplan — a legible dry run before the irreversible stepVersion constraints, provider signing, download counts
FigmaA component in a library, and its instancesLive link with a deliberate detach, plus propagationAn instance always names its main component
NotionA page, and a template as a duplicatable pageDuplicate as the distribution primitiveCreator profiles, usage counts, official gallery
Aspirational — the craft bar
LinearAn issue in an opinionated workflowSpeed as a feature — keyboard-first, command menuOpinion: the product tells you how to work
RaycastA command; an extension is a bundle of themLauncher-first — everything one keystroke awayOpen source, author identity, install counts
VercelA deploymentThe deploy as a designed event with a log you actually readThe log itself — nothing hidden behind a spinner
StripeAn API object with explicit relationshipsDocumentation as interface — the reference is the productPrecision: versioned docs, errors that name the cause
GitHubA repositoryFork and pull request — copying is a first-class social actStars, contributors, visible history, dependency graph

Monetisation is the fifth axis and is omitted here for width; it is in the source document. Packmind's is marked unverifiedno source — because app.packmind.com does not resolve and the product is sales-gated.

Evidence · click to enlarge

Tessl skill detail page showing a composite score of 93 and an uplift multiplier
tessl · skill-detail-scored — the reference screen of the survey: composite 93, uplift 1.40×, Quality / Impact / Security bars, and identity as repo + path + commit.
Smithery registry home listing MCP servers with scores
smithery · registry-home — a catalog solving cold start with volume, every server carrying a number.
Backstage catalog table with name, system, owner, type, lifecycle and tags
backstage · catalog-table — the closest thing to our Library that exists today.

Three market patterns

  1. The composition layer is consolidating, fast. Of five hard competitors, one had its hosted layer switched off (Continue, acquired by Cursor in June 2026 — the open client survives, frozen) and one was acquired mid-survey (Smithery is now part of Arcade.dev; we found the banner on their own page, not in any article). Note precisely what was absorbed in Continue's case: the registry, the accounts and the subscriptions. The local client and the open format were left standing.
  2. Trust has moved from social proof to measurement. Tessl scores every skill on quality, impact across n eval scenarios and a security scan, then shows a composite and a multiplier. Smithery shows a score out of 100. Port shows pass / warn / block. A number derived from running the thing has replaced stars and downloads.
  3. Nobody authors a graph; everybody derives one. Backstage's graph is read-only and filtered by depth. Terraform resolves its graph from declared constraints. GitHub derives dependents. Not one product asks a human to draw an edge.

Three differences we can hold

  1. Every live hard competitor sells to an organisation; the practitioner is not the buyer. The value proposition is always control over other people's work — governance, drift, compliance. A personal library that answers to nobody is unoccupied ground. Narrowed by stage 6: that was measured on funded vendors. At the open-source practitioner tier the item-level ground is occupied — what still looks empty is the set-level half.
  2. They all trust the network; we can trust the file. Every catalog solves cold start with curation and volume — 17,500 MCP servers, 3,000 skills. None of them makes your own accumulated material better.
  3. Assembly is a side effect for them and the whole product for us. Smithery composes so it can host, Tessl so it can govern, Backstage so it can scaffold. Nobody treats does this set actually hold together as the product.

Stage 1 — conclusion

The market has converged on measured trust, and we cannot measure anything — we have no way to run a skill. So the honest move is to claim nothing per item: an item card carries usage facts from the user's own library (used in 3 projects, 2 items require this) and never a score, rating or badge.

Rejecting the node canvas turns out to be the market consensus rather than a compromise — fifteen products, none of which asks a human to draw an edge.

And the unoccupied ground is real but unproven. Continue is the only data point on whether an individual practitioner pays, and it reads both ways: the monetised hosted half was switched off and erased; the free local half is still on 1.58M machines.

STAGE 02

Flows

Twelve mechanisms, captured from products that are actually running — behaviour, not appearance.

Each flow was captured from live software: public web where possible, the owner's own signed-in sessions where a product needed one, local CLIs in scratch directories. Read-only throughout — nothing was created, published or left changed. Ten closed, one declined, one deliberately handed forward to the design phase.

Open any row for what it taught and the captures behind it.

01Item detail and trustWhat evidence does a catalog put on a single item?declined

Trust has moved from social proof to measurement — a number derived from running the thing. Tessl scores per eval scenario with a delta; Smithery shows a score out of 100 and a verified badge. We can run nothing, so we ship none of it. Held open only by Agentman's login-walled page, then declined: its lesson is legible from their marketing, and version history is out of scope.

Captures · click to enlarge

Tessl Evals tab showing per-scenario scores
tessl-skill-evals-tab — per-scenario results, 84% with a ↑34% delta, each with a Details affordance.
Smithery server configuration requirements
smithery-server-github-config — a server that actually requires configuration, and how it asks.

research/2-flows/01-item-detail-and-trust/

02Library browse, filter, searchFinding one thing in a large personal collectionclosed

Typing flattens a filter tree into breadcrumbs — the path is shown rather than traversed, which is how two leaves with the same name stay distinguishable. Our mcp kind and an mcp tag will collide in exactly that way. A filter chip is written as a sentence, with Clear and Save in the same corner.

And the one to copy outright: when a filter empties the list, Linear says “4 issues hidden by filters · Clear Filters ✕”. That is the difference between there is nothing and you are not looking at it, stated as a number you can act on.

Captures · click to enlarge

Linear command palette running commands and content search in one field
linear-command-palette-default — one field doing commands and content, with a keyboard-hint bar computed from the current state.
Linear filter typeahead flattening its hierarchy into breadcrumbs
linear-filter-typeahead-breadcrumbs — the filter tree flattened into a path as you type.
Linear filtered to zero results naming how many issues are hidden
linear-filtered-to-zero-hidden-count — the count of what is hidden, and the remedy beside it.

research/2-flows/02-library-browse-filter-search/NOTES-linear.md

03Relations without a canvasShowing dependencies nobody drew by handclosed

Nobody authors a graph. It is derived from declared relations and made legible with a depth filter — the same graph at depth 1 and depth 3 is the whole legibility mechanism. GitHub's Dependents view is the reverse direction: blast radius rather than dependencies, which is the shape of our used in 3 projects.

Captures · click to enlarge

Backstage catalog graph at depth 1
backstage-graph-depth-1 — legible, because it is shallow.
Backstage catalog graph at depth 3, much denser
backstage-graph-depth-3 — the same data three hops out. The control is the design.
GitHub dependents view showing blast radius
github-dependents-blast-radius — who depends on me, not what I depend on.

research/2-flows/03-relations-without-canvas/

04Validation — pass, warn, blockOur wow moment, and the best-covered flow in the researchclosed

Four products, four different lessons: Terraform on what a result reads like, Vercel on what a process reads like, GitHub Actions on partial failure, Port on a grade.

The structural find is Vercel's: a deployment is a stack of collapsed stages, each carrying its own glyph, verdict and duration, each expandable. A stage that did not run gets a clock, not a failure.

The second find decided our severity model. Port does not block — it gates a level. A failed rule stops an entity climbing a cumulative Basic → Low → Good → Great ladder rather than forbidding anything; the demo has no blocking surface at all.

Captures · click to enlarge

Vercel deployment page showing collapsed build stages with durations
vercel-deployment-stages — the structural model for our validation pass.
Port scorecard rule expanded to its predicate and observed value
port-scorecard-rule-expanded — where "Open Critical Vulnerabilities" = 0 · Value: 1. The condition required and the value found.
GitHub Actions run annotations reading process completed with exit code 1
github-actions-run-failed-annotations — the anti-reference. A header counting “11 errors and 6 warnings”, over a slot that says nothing.

research/2-flows/04-validation-check-results/ — NOTES-vercel, NOTES-port, NOTES-github-actions, terraform-plan-output

05Linked vs detached, and blast radiusThe single most spec-bearing capture in the folderclosed

Drift is computed precisely and displayed nowhere. Figma's context menu names the exact overridden property — Reset fill, not a generic “reset overrides” — so the diff is already computed. And yet a modified instance is pixel-identical to a clean one in the layers tree, the properties panel and on the canvas. The only evidence is two rows appearing in a menu you must open on an object you must first select.

The failure is display, not modelling — which is exactly the trap our detached, locally modified state is waiting to fall into. Two things to take whole: revert at two granularities (whole item, single field), and Reset kept next to Detach as the two halves of one axis.

And the part with no prior art at all: a path back. Figma erases the origin at detach and offers nothing afterwards.

Captures · click to enlarge

Figma context menu on a clean component instance
figma-context-menu-clean — no reset rows. They are absent, not greyed.
Figma context menu on an overridden instance showing Reset instance and Reset fill
figma-context-menu-overridden-reset — Reset fill names the exact property that differs. The product knows; the canvas never says.
Figma library updates modal showing per-row instance counts including 423
figma-library-updates-instance-counts — 5, 2, 70, 423 instances before you accept. The count appears exactly when the action reaches beyond what you can see.

research/2-flows/05-linked-vs-detached/NOTES.md

06Export and target adaptationOne source, many agent formatsclosed

External repos are always instructions, never vendored. Ruler — an open-source CLI that syncs one directory into 20+ agent formats — was run locally as a substitute for the sales-gated competitor, and it supplies two mechanics: a managed START/END fenced block so a re-run finds and replaces exactly what it wrote, and a .bak beside every overwritten file. create-next-app arrived at the same two mechanics independently, which collapses our four export targets into two mechanics plus a naming convention.

What neither does: say what it is about to overwrite.

Captures · click to enlarge

Backstage scaffolder template parameter form
backstage-scaffolder-template-form — parameter form → review → artefact, which is our export flow in another domain.

research/2-flows/06-export-and-target-adaptation/ruler-per-agent-output.md

07Env variables and secretsHow to ask for a key without leaking itclosed

Ask the type before the value, and default to the irreversible option. Vercel's drawer offers Secret and Config as two explained radio cards with Secret — the one you cannot read back — as the default. The Note field's placeholder poses the real question: “Where to rotate, or who to contact”. Two bulk paths exist, including pasting a whole .env into the Key field.

The counter-example sits one screen away and is this research's standing example of waste: the env-variables empty state keeps a search box, four filter dropdowns and a sort control on screen, all filtering nothing.

Captures · click to enlarge

Vercel add environment variable drawer with Secret and Config radio cards
vercel-add-env-variable-drawer — type before value, irreversible by default.
Vercel environment variables empty state still showing filters
vercel-env-vars-empty-state — four dropdowns filtering nothing.

research/2-flows/07-env-and-secrets/NOTES.md

08Empty state and cold startThe gap the research carried from the beginningclosed

A brand-new Linear workspace does not open on an illustration inviting you to create your first issue. It opens on the Issues list, already populated with four real issuesGet familiar with Linear, Set up your teams, Connect your tools, Import your data — ordinary objects you can edit, complete or delete. The onboarding checklist is the data model, exercised on itself, and the product is never empty at any point.

Second: three registers of emptiness, and a rule that picks between them. A concept you may never have used defines itself in a sentence. Routine emptiness says one line. Filtered to zero counts what is hidden. Verbosity scales with the chance the reader does not know what the object is.

Captures · click to enlarge

A brand new Linear workspace opening on four real seeded issues
linear-first-run-seeded-issues — the real first five minutes, which no populated system could have shown us.
Linear Projects empty state defining what a project is
linear-projects-empty-state — the object's name as the heading, then a definition rather than an apology, with the shortcut inside the button.

research/2-flows/08-empty-state-and-cold-start/NOTES-linear.md

09Duplicate and forkCopying as a first-class actclosed

Duplication is a distribution primitive, and the copy dialog states what will and will not come along. This is the mechanism behind our second supporting moment — duplicating a project to re-tune it for a new context.

Captures · click to enlarge

GitHub create a new fork form
github-create-fork-form — opened and abandoned; nothing was forked.
Notion page menu with Duplicate
notion-page-menu-duplicate — duplication offered where the object lives.

research/2-flows/09-duplicate-and-fork/NOTES.md

10Dark design languageDeliberately not spent in this phasehanded forward

The one flow about appearance rather than behaviour, so it does not close here — it opens lesson 06, concept, with its material already gathered. Linear at close range: few surfaces, with the sidebar sharing the content ground; elevation as a hairline plus a few percent of lightness rather than a shadow ramp; three foreground tones with the accent reserved for meaning; key caps as a real component with a two-cap form for chords; unset properties written as imperatives.

Captures · click to enlarge

Linear issues list in dark theme
linear-issues-list-dark — few surfaces; the sidebar shares the content ground.
Linear command palette in dark theme, detail of key caps
linear-command-palette-dark-detail — key caps as a component, not a text hint.

research/2-flows/10-dark-design-language/NOTES-linear.md

11Copy — errors, warnings, refusalsEvery sentence the validation pass will ever writeclosed

Cause and consequence are two rows, and you group by the action that fixes a failure rather than the check that found it. Nobody opens a 102-job run to find seven red squares — so the best example in the survey sends the report to the reader instead: a bot posts a distilled failure summary into the pull request, grouped by the command that reproduces each failure, stamped with the commit, with raw output collapsed.

The anti-reference is industry-wide. Fourteen of seventeen annotations on one run read, in full, Process completed with exit code 1 — and the same was found in denoland/deno, withastro/astro, vitejs/vite and rust-lang/rust-clippy. The failure channel carries less information than the deprecation channel.

Captures · click to enlarge

GitHub pull request bot comment listing failing test suites grouped by reproduction command
github-pr-bot-failing-test-suites — the best failure copy in the survey.
Stripe error codes reference
stripe-error-codes — error copy that names the cause, as documentation.

research/2-flows/11-copy-and-error-language/ — ci-failure-copy, dependency-conflict-copy

12Visibility and portfolioPublic or private, and what a profile says about youclosed

Public/private is a low-ceremony decision, and the profile is the portfolio. GitHub puts the visibility flip in a Danger Zone with a typed confirmation; Notion separates sharing with people from publishing to the web entirely.

Captures · click to enlarge

GitHub danger zone repository visibility row
github-danger-zone-visibility — read-only; nothing was clicked.
Notion publish to web dialog
notion-publish-to-web — publishing kept separate from sharing.

research/2-flows/12-visibility-and-portfolio/NOTES.md

Stage 2 — conclusion

Four findings reached back into the specification and changed it.

  • Cold start is a product decision, not a demo problem. Ship real objects, not an empty state — so the validation pass has something to run on before the user has typed anything.
  • The drift indicator is a display problem, not a modelling one. The diff is already computed, so putting it on the card costs nothing. Figma proves it can be computed and still never shown.
  • There is no prior art for the return path. Every product either keeps the link or erases the origin. Bringing a locally modified item back into the library is something we invent, not copy.
  • An action that cannot act is not shown — on primary surfaces. Linear suppresses its toolbar over an empty list; Figma hides a reset with nothing to reset; Vercel keeps four dropdowns over nothing and is simply wrong.

What the flows could not see: density. Every browsing mechanic was captured in a workspace holding four issues. All of it is mechanics; none of it is scale. not established

STAGE 03

Pain

The first evidence in this research about users rather than vendors — and the first that argues with us.

Everything up to here compared companies. Nothing had established what actually hurts. Our whole product rests on one asserted pain: nothing tells you that a skill needs a particular MCP server, that two items write to the same file, or that a key is missing — you find out at runtime. That claim had never been checked.

Two public issue trackers, queried through the GitHub API: continuedev/continue (6,677 issues — the closest dead competitor, same object model) and modelcontextprotocol/servers (1,186 issues — the ecosystem our mcp kind lives in), ranked by reactions.

What this instrument cannot see

An issue tracker records breakage, not friction. People file an issue when something is broken, not when something is tedious. This is not a caveat at the bottom of the page — it governs every number below.

Candidate painVisible here?Why
A · LossNo“I can't find the prompt I wrote three months ago” never becomes an issue
B · Reassembly costNo“I copied the same four files into a new project again” produces silent tedium
C · Silent breakageYesThis is what trackers are made of

The loud pain is environmental, not compositional

The single most-reacted issue in either repository, by a factor of eight over anything configuration-related, is MCP Servers Don't Work with NVM — the app tries to use the wrong Node and fails. The top of that tracker is almost entirely it will not start on my machine.

MCP Servers Don't Work with NVM — node version managers182
Storing API keys in plain text (Continue)32
Environment variables not respected in server-memory23
Security proposal: credential management22
GitHub MCP server fails to start — npx error19
Two server-postgres entries in one config — our own thesis13

Reactions per issue, scaled to the largest. Red — environmental failure. Teal — the failure our product exists to catch.

Our thesis is observed, but quietly

It is there. It is just not loud:

“It's not possible to run multiple instances to connect to different databases. When doing this, the chat always chooses the first one specified in order of mcp.json.”

modelcontextprotocol/servers #1219 — 13 reactions

That is our duplicate-key collision, reported by a user, in the exact file we generate, with the exact failure mode we predicted: the first one wins and nobody is told. The same behaviour sits in Continue's own source, where a duplication detector correctly identifies the clash and the merge silently discards the loser. Two independent sightings of one bug — once in code, once in a user's words.

And nobody is asking for a composition layer

The uncomfortable one. In a 6,677-issue tracker belonging to a product that shipped a hub for exactly this, a search for reuse blocks assistant returns 0 results. Set beside the post-mortem — the hosted hub is the half that was switched off, the local client is the half still running on 1.58M machines — that is two independent signals pointing the same way.

Stage 3 — conclusion

Keep the validation pass. The failure is real, currently silent, and sighted in the wild. Nothing here says stop.

Re-weight the export. Our specification gave the generated setup instructions a single line, while the loudest pain in the ecosystem lives exactly there: the archive lands on a machine and does not run. Most of that is beyond our reach — we validate a set, we cannot fix someone's PATH — except at the one place our output meets their machine.

Do not lead with sharing. Two independent signals against it.

And say plainly what is still unknown. Whether loss or reassembly cost is what would make someone adopt this is invisible to trackers by construction. It needs a different instrument — asking people — and we have not done it. not established

STAGE 04

Benchmark

For each of our four core flows: who does it best in the world, how well, and against what standard?

Stages 1–3 produced observations — Linear counts what its filter hides, Port prints the condition and the observed value, Figma computes drift and never draws it. Those are anecdotes until they are turned into a scale. This stage builds the scale, and stage 5 spends it. The five categories are lifted from findings in stages 1–3, so the rubric is grounded rather than invented.

CategoryThe question it asks
C1 · State legibilityCan you read the current state without acting on it?
C2 · ConsequenceBefore an irreversible step, is the cost stated in advance?
C3 · Failure copyDoes a message name the item, the rule and the observed value?
C4 · RecoveryIs there a way back, offered where the problem is?
C5 · EconomyIs anything on screen that cannot act, or missing that must be?

Anchors. 1 — actively misleads. 2 — the information does not exist. 3 — correct, but you must go looking. 4 — present where you need it. 5 — you could not miss it, and it changed what you did next. A dash is not a zero: it means the flow contains no instance of what the category grades, and the cell always says whether that is the product's fault or the method's.

5 4 3 2 1 — actively misleads — no instance
CandidateC1C2C3C4C5
B1 — Find one thing in a large personal collection · our Library
Linear ⌘K + filters5455
GitHub code search5334
Obsidian quick switcher42254
VS Code palette4224
B2 — Assemble a set under constraints · our Project
VS Code workspace + trust5543
npm install34544
Figma instances + library2435
B3 — Check a set and report what is wrong · our validation pass
terraform plan + validate55445
VS Code Problems panel5445
Vercel build log54
GitHub Actions run5123
B4 — Produce an artefact and hand it to another machine · our Export
Vercel deploy53
create-next-app4423
Figma export dialog4433
Ruler per-agent output3244

Who wins. B1 — Linear, and not narrowly. B2 — VS Code, on Workspace Trust. B3 — terraform, with VS Code's Problems panel level on everything but the summary line. B4 — nobody. The highest cell in the whole flow is a 4.

Three mechanisms we are taking into the MVP

01 — The run as a stack of stages · Vercel

A deployment is a stack of collapsed stages, each carrying its own glyph, verdict and duration, each expandable. A stage that did not run gets a clock, not a failure. The log header states 66 lines before you read it, and a timestamp hover gives relative to start and relative to previous — which is how you find the step that hung without doing arithmetic.

Why it fits. Our validation pass is specified as a designed, legible moment rather than a spinner, and this is that moment's structure, already proven. It also gives a home to two things nobody else has a place for: our Skipped severity maps onto Vercel's clock glyph, and the handover disclosure becomes two more stages rather than a new surface. Terraform reinforces it from the other side — no progress theatre; text that appears when ready and reads correctly frozen.

02 — N of M, and a count that is a link · VS Code

Workspace Trust shows two columns — In a Trusted Folder against In Restricted Mode — with the current one outlined. Two of its lines are counted and hyperlinked: “95 workspace settings are not applied”, “10 extensions are disabled or have limited functionality”. It is the best consequence disclosure in the entire benchmark, and it beats Figma's 423 instances on the one axis Figma leaves open: the number is a link to the list.

Why it fits. We refuse to block an unclean export and confirm it instead, which puts the whole weight on the confirmation sentence. This is the shape that sentence takes: name the count, and let the user open it. The same grammar showed up in four unrelated products — 4 issues hidden by filters, Showing 0 of 6 placed inside the filter input, 0 of 0 selected, 10 files (324 ms) in repo ✕. The number you can act on, next to the control that produced it — and never the word “none”.

Captures · click to enlarge

VS Code Workspace Trust dialog with counted, hyperlinked consequences
vscode-workspace-trust — the cost of each state, counted and clickable, before you choose.
VS Code Problems panel filtered to zero showing Showing 0 of 6 inside the input
vscode-problems-filtered-zero — Showing 0 of 6 inside the input, and the remedy as a link.

03 — The failure line, written where the rule is known · npm and Port

npm's ERESOLVE answers all four questions a conflict message owes: what was required, what was found, who required it, and who asked for what was found — which is precisely our added manually versus added as a dependency. Then both escape hatches, with the cost in the same sentence: “to accept an incorrect (and potentially broken) dependency resolution.”

Why it fits, and it is not a copywriting decision. Compare the 1-grade message: Process completed with exit code 1, under a header counting “11 errors and 6 warnings”. The difference is architecture. The 5-grade messages are emitted where the requirement is known; the 1-grade one is emitted by a process that only knows it stopped. So: the check that knows the rule must be the thing that writes the sentence.

Two refinements ride along. Count problems the way the user counts them — Terraform emitted four diagnostics for one defect, and VS Code reported a duplicate key as two unjoined rows that never mention each other. And a check that could not run must say so, which is why Skipped exists as a third severity rather than a silently missing row.

One mechanism that will not work for us

Figma's export dialog with nothing selected reads 0 of 0 selected beside a greyed Export button, explained by one good sentence: “No selected layers have export settings. Click + in the export section of the properties panel to add one.” It scores respectably — the state and the remedy are both stated precisely.

It works there for three reasons, and we have none of the three:

In FigmaFor us
The blocker is one named action awayOur problem may be four items away, on another surface
The condition belongs to this second's selectionOur problems are properties of the project
Re-exporting costs nothingThe archive is the product's whole point

A greyed primary action is a dead end on the most important control in the product, and our promise is the system checked, not the system forbade. So the mechanism is refused outright: export is never disabled; an unclean set is confirmed, not blocked.

Captures · click to enlarge

Figma export dialog showing 0 of 0 selected with a disabled Export button
figma-export-empty-selection — precise, well-written, and exactly what we cannot do.
Obsidian quick switcher with no match, offering to create the note
obsidian-quick-switcher-no-match — the opposite failure: zero results is the create row, so a typo and an absence look identical. We take the offer and keep the message.

Stage 4 — conclusion

Handover is the industry's weakest flow, and it is ours to win. No cell in B4 scores above 4. Ruler overwrites without warning; create-next-app ends on the word Success! and lists nothing; Figma's dialog disables its own primary action. Our thirty-second archive is aimed at the least-well-served flow of the four, and the specific gap is disclosure before the write.

But this rubric grades craft, not weight. The categories came from surveying products, so every score answers did this interface tell me the thing — never does anyone bleed here. Read against the pain evidence, the weights come out uneven: B1 and B2 are where the craft is, and a tracker is structurally blind to both. B3 and B4 are where the value is — one carries the only sighting of our own thesis, the other carries the loudest pain in the ecosystem.

So a category score alone cannot choose a shape. It has to be read next to which third of the flow that shape owns.

STAGE 05

Patterns

Five radically different shapes for the key flow — assemble a set, check it, export — scored on the rubric stage 4 built.

Not five layouts of one idea. Five framings, each carrying prior art from the earlier stages, so the comparison is between things that exist rather than between guesses. Each answers the same five questions in the same words, and each is scored on does this shape give the category a natural home.

ShapeC1C2C3C4C5The idea
P1 · Two-pane drag43342Library left, project right, items dragged in
P2 · Command-first34545No library pane; ⌘K adds by name, the project is a list
P3 · Document45534The project is an editable manifest, validated like a linter
P4 · Staged wizard34422Target first, one step per kind, then resolve and export
P5 · Run-centric55545Check takes the whole surface; export is its final stage

P5 takes straight fives on four categories and still cannot be the answer alone, because it does not assemble anything. P3 scores best on assembly and dies outside the scores. The constraints that were never up for negotiation: desktop-first, dark from day one, not a node canvas, single user, local only, custom design system.

The choice — a hybrid, stated as a choice

P2 wins the spine. P5 becomes the check-and-export surface. P3 donates one mechanism.

Neither half is a compromise — the two shapes win different thirds of the same flow, and the seam between them is a single control: Check.

Surface 1

Library

Browse, filter, add, edit. Two scopes — your own items and a curated read-only public shelf. Usage facts per item. Not part of the builder flow.

Surface 2

Project

The set as a list, not a canvas. ⌘K adds by name and opens on related items. Every row carries its own state, including detached with the differing fields named.

Surface 3

Run

Entered by Check, takes the whole screen. Stages with verdicts and durations. Export is the final stage, not a button beside the check.

Why this one, for our context specifically

  1. It is the only shape that gets better as the library grows. A palette is indifferent to list length; a pane is not. Drag fails the concrete test — item #250 dragged to a target scrolled out of view — and every mitigation for that is click-to-add, at which point the two-pane shape has quietly become the command-first one with an extra pane.
  2. It obeys our own economy rule instead of breaking it. We refuse controls that cannot act — that is why the visibility toggle is not shipped. A library pane cannot act while a validation pass is running. Applying our own rule to our own design is what removed the pane.
  3. It puts the wow moment on a whole surface. Export as the final stage makes the unclean-export confirmation the next row in a list the user is already reading, rather than a dialog interrupting a flow.
  4. It renders the six item states once, not twice. In a two-pane shape every item renders as a library row and as a member, so detached — a per-project fact — is invisible in the pane where you browse.

Stage 5 — conclusion

The substantive change: the library leaves the builder, and drag stops being the verb. The earlier specification put a filtered library sidebar inside the builder and made drag the mechanism. It does not survive 300 items, and a pane that cannot act during a check is the mistake we refuse everywhere else.

The cost, stated plainly. With no library pane you cannot see what you are not using. Three things now carry that weight, and if all three fail this choice was wrong: the palette opening cold on related items rather than an alphabetical index; the Library one keystroke away and remembering where you were; and per-item usage facts doing the work a visible pane would otherwise do.

CLOSE

What we decided, and why

Ten decisions came out of the phase. Each is a product commitment with evidence behind it, not a preference.

Trust signalAn item card carries usage facts, never a score
The market converged on measured trust — and we have no way to run an item, so any number we invented would be decoration. Used in 3 projects is a usage fact, computable locally, and it is also the blast-radius number the research spent a whole flow chasing.
ValidationThree severities, and nothing blocks
Problem, Note, Skipped. Port gates a level and forbids nothing; Continue shipped the same fatal/non-fatal binary we had specified and filed a missing dependency as non-fatal. Both poles have shipped, and we chose neither.
ExportNever disabled — an unclean set is confirmed
Almost nothing makes an archive impossible to produce; a duplicate command still zips, it is simply wrong inside. So blocking is a choice, and we do not make it. Pressing Export on a set with problems names the consequence in the present tense: “Two items write to .mcp.json. The archive will contain only one of them.”
CyclesInformation, not errors
A project is a set, not an execution order — we never ask what runs first. If two items require each other, both are added and the set is correct. Report it as a fact: “these three always travel together.”
VersioningThe field is cut; external references are pinned
Our model is Figma's, not Terraform's: an item lives in the library once and projects link to it live, so there are never two versions to disagree about. But external repos are not ours — without a pinned commit, the setup instructions hand the user whatever HEAD is that day.
ShapeCommand-first assembly, run-centric check
Three surfaces plus Projects, with the seam at Check. The library leaves the builder because drag does not survive 300 items and an idle pane during a check breaks our own economy rule.
Cold startA curated public library ships with the product
The Library gains a scope switch — your own items, and a read-only public shelf — so there is material from the first second while your own library starts honestly empty. Plus one example project composed to contain a real problem and a real note, because a first run showing six green ticks teaches nothing about what the product is for.
HandoverThe setup document is written for an agent
Not for a human who reads it. It states, per item, what that item requires — dependencies, servers, keys, repos at their pinned commit, target paths — so the agent that opens the project performs the setup. That aims at the 182-reaction pain without a platform matrix and without a new artefact.
Detached itemsEdited in the project, promoted as a new item
Editing in place is what makes detaching a feature rather than a toggle. Promotion creates a new library item and re-links the row — updating the original would change every other project that links to it, which is blast radius spent on an action taken inside one project.
CompositionDeferred, not refused
A project inside a project is out of the MVP. No pattern variant needed it, and it risks unbounded recursion, so the depth rule gets decided before the feature, not after. The second reason given at the time — nobody is filing for it — did not survive stage 6: the same friction is filed on live trackers as feature requests. The decision now stands on recursion alone.
CLOSE

What we still have to find out

Nine gaps, each with a hypothesis written so it could be shown false. Where nothing in the research supports a claim in either direction, it says so rather than filling the space.

G1
Nobody was ever asked anything
not established

Both trackers are structurally blind to loss and to reassembly cost; the phase read vendors and artefacts, never users.

HypothesisReassembly cost, not loss, is what converts — a practitioner adopts to stop re-copying the same four files, and finds searching my own corpus valuable only afterwards. Falsifiable in five conversations.

from Stage 03 Pain · Stage 04 Benchmark

G2
Handover has no prior art to copy

No candidate scored above 4, and not one cell in the matrix scores what a product says about the machine its artefact lands on — because no candidate had such a surface. Narrowed 10 Sep 2026: four open-source tools were installed and run, and one of them checks git, authentication, the Node version, 21 writable directories and PATH shadowing before anything else happens. The gap is real and its explanation has changed: nobody describes the receiving machine for a named set.

HypothesisDisclosure before the write is the whole opportunity. If the last stages of the run state what the archive contains and what the receiving machine must still do, the archive-does-not-run pain drops without us running anything there. Falsifiable: if users hit environment failures at the same rate, the disclosure was theatre.

from Stage 04 Benchmark

G3
The return path has no prior art

Figma erases the origin at detach and offers nothing afterwards. No product in the survey lets a local override become a first-class object again.

HypothesisPromotion as a new item is safe and update the original is not, because the second spends blast radius on an action taken inside one project. Falsifiable: if users routinely promote and then immediately delete the original, they wanted a merge and we built the wrong verb.

from Stage 02 Flows, flow 05

G4
Density was never observed
not established

The whole browsing grammar was captured against a workspace of four issues, and the benchmark's Obsidian cell against a vault of 27 notes. The 300-item claim behind the chosen shape is reasoned, not measured. Superseded by stage 6: four public collections were counted through the tree API at 11, 25, 47 and 48 items, two of them above forty. The 300 is refuted by measurement; density at any size is still unobserved. Design for fifty.

HypothesisThe palette holds at 300 items and the failure mode is discovery, not search — people find what they can name and stay blind to what they cannot. That is precisely the cost already accepted, so it is testable the moment a seeded library exists.

from Stage 02 Flows · Stage 05 Patterns

G5
The public library does not exist yet

It is a decision with no content behind it: no items, no verified sources, no composed example project.

HypothesisA seed of 8–12 items that produces at least one problem and one note teaches the product better than 30 clean ones. Falsifiable on first use: if six green ticks leave people unable to say what the product is for, the seed was decorative.

from Stage 01 Competitors · Stage 02 Flows, flow 08

G6
Licensing for redistributed items
not established

Nothing in the research covers the terms under which someone else's skill may ship inside our public library.

HypothesisA pinned reference plus visible provenance is necessary but may not be sufficient. This needs a licence review before the shelf is built — it is not a design decision.

from Stage 01 Competitors · Stage 05 Patterns

G7
Two of five hard competitors were never seen
not established

Agentman is login-walled, Packmind sales-gated. Their mechanisms and Packmind's monetisation are read from marketing, not observed.

HypothesisNeither changes the picture, because both sell to organisations and that difference already covers the whole group. Falsifiable if either turns out to sell to individuals.

from Stage 01 Competitors

G8
Item versioning and history were never observed

Flow 01 was declined. Tessl pins a commit, but a version change was never seen anywhere.

HypothesisNot needed: detach plus overrides does the job for our own items, and a pinned reference does it for external ones. Falsifiable the first time someone asks what did this item look like last month.

from Stage 02 Flows, flow 01

G9
Our own thesis is real but quiet

The duplicate-key collision is sighted in the wild with 13 reactions, against 182 for an environmental failure. Re-weighted by stage 6: the loudest number in the evidence base is 6,592, and it is about one instruction source feeding several agents.

HypothesisSilent breakage is a retention argument, not an acquisition one — it is what makes the product trusted once adopted, not what makes anyone try it. A positioning claim; the product is the same either way.

from Stage 03 Pain

What is settled, so it does not get re-opened

  • The node canvas. Rejected on market consensus, not taste — not one of fifteen products asks a human to draw an edge.
  • Blocking. Three severities, export never disabled, an unclean set confirmed.
  • A composition layer. Two independent signals against it, plus the recursion risk.
  • Per-item scores or badges. The market converged on measured trust and we can measure nothing.
  • Visual direction. Deliberately not decided here. The material is gathered and the design-system phase spends it.
CLOSE

Captures

What was actually looked at, and where it is kept.

Every screen on this page is embedded in it as a data URI, so the page opens from disk with nothing to fetch. The originals, the ones that did not make the page, and the machine-readable logs live in the repository beside the source documents.

Image captures
14334 on this page
Source documents
47across seven stages
Raw data logs
12queries and results
Products captured
15three groups
Flows
1210 closed
Sign-in walls
5labelled as such

Where each kind lives. Stage 1’s 38 captures are catalogued with a line each, sign-in walls marked, in the screens index. The flow captures sit in twelve numbered folders, each with a notes file that cites the frames it read. The two issue-tracker pulls are raw JSON. The Hacker News threads were downloaded whole and read locally, and the queries that produced them are logged with their result counts, so anyone can re-run them and get the same corpus.

Two limits printed on the evidence rather than buried. Every browsing mechanic on this page was captured against a workspace of four issues and a vault of 27 notes — density under load was never observed at all. And five of the fifteen products could only be read from marketing pages and public documentation, because their product surfaces sit behind a sign-in wall; those cells say so.