Overview
Five stages in four days, and the two lists they produced.
What we saw
- Every live competitor sells to an organisation. Governance, drift, compliance — control over other people’s work. The individual practitioner was unoccupied ground among funded vendors, though open-source tooling has since filled the item-level half of it — and on 10 Sep 2026 four of those tools were installed and pointed at a deliberately broken set. None of them has a set-level unit at all.
- Trust has moved from social proof to measurement. A composite score, an uplift multiplier, a scan — a number derived from running the thing has replaced stars and downloads. We can run nothing.
- Nobody authors a graph; everybody derives one. Not one of fifteen products asks a human to draw an edge between two objects.
- The loudest pain is environmental, not compositional. The archive lands on a machine and does not run — PATH, node version managers, platform paths, processes dying at startup.
- Our own thesis is real and quiet. Two servers claiming one key, “the chat always chooses the first one specified”: the first wins and nobody is told. Thirteen reactions against 182.
- Handover is the industry’s weakest flow. Fifteen cells scored, no B4 cell above 4, and not one of them scores what a product says about the machine its artefact lands on. Narrowed on 10 Sep 2026, after four open-source tools were installed and run: one of them does describe the receiving machine. What nobody does is describe it for a named set.
What we decided
- An item card claims nothing. Usage facts from the user’s own library — used in 3 projects — never a score, rating or badge.
- Three severities, and nothing blocks. Problem, note, skipped. Export is never disabled; an unclean set is confirmed, with the consequence named in the present tense.
- Command-first assembly, a run-centric check. Three surfaces plus projects, with the seam at the check. The library leaves the builder.
- The setup document is written for an agent, not for a human reader — which is what aims at the 182-reaction pain without a platform matrix.
- A curated public shelf ships with the product, so there is material from the first second while the user’s own library starts honestly empty.
- Composition and versioning stay out, deferred rather than refused, with the reason written down.
What this phase could not do. It read vendors, artefacts and trackers. It never asked a person anything. That is phase 02 — personas and JTBD — and it changed several things on this page, each marked where it happened.
Competitors
Who else is in this space, what do they sell, and to whom?
Fifteen products in three groups. Hard — aiming at the same job as ours. Soft — a different product doing the same work somewhere else. Aspirational — the craft bar. Each read on five axes: who it is sold to, what its atomic asset is, the mechanism nothing else has, how it earns trust, and how it charges.
| Product | Product base | Key mechanism | Trust |
|---|---|---|---|
| Hard — same product, same audience | |||
| Tessl | Skill as a registry artifact pinned to repo + path + commit | Publishing gated by evaluation — a skill ships only after it scores | Composite 93, uplift 1.40×, Quality / Impact %, Snyk scan |
| Packmind | A versioned engineering playbook of standards | One source distributed into each agent's format, drift detected back | Governance framing, pre-commit blocking |
| Agentman | Skill as an executable, shareable unit | Visual assembly of skills into a working agent | Compliance badges + four permission tiers |
| Smithery | MCP server as a hosted, connectable endpoint | Connect once, reuse everywhere — it holds auth and sessions | Score /100, verified badge, usage counts |
| Continue Hub | A block: model, context, docs, MCP, rule, prompt | Compose blocks into a custom assistant | Open source and forkable — trust by inspection |
| Soft — different product, same job | |||
| Backstage | A typed entity declared in YAML beside the code | Relations derived and filtered, never drawn | Ownership: owner, lifecycle, source file on every entity |
| Port | Entity in a Context Lake, plus scorecards | Standards applied as scorecards over the catalog | Pass, warn or block — with the reason attached |
| Terraform | A versioned module with declared constraints | plan — a legible dry run before the irreversible step | Version constraints, provider signing, download counts |
| Figma | A component in a library, and its instances | Live link with a deliberate detach, plus propagation | An instance always names its main component |
| Notion | A page, and a template as a duplicatable page | Duplicate as the distribution primitive | Creator profiles, usage counts, official gallery |
| Aspirational — the craft bar | |||
| Linear | An issue in an opinionated workflow | Speed as a feature — keyboard-first, command menu | Opinion: the product tells you how to work |
| Raycast | A command; an extension is a bundle of them | Launcher-first — everything one keystroke away | Open source, author identity, install counts |
| Vercel | A deployment | The deploy as a designed event with a log you actually read | The log itself — nothing hidden behind a spinner |
| Stripe | An API object with explicit relationships | Documentation as interface — the reference is the product | Precision: versioned docs, errors that name the cause |
| GitHub | A repository | Fork and pull request — copying is a first-class social act | Stars, contributors, visible history, dependency graph |
Monetisation is the fifth axis and is omitted here for width; it is in the source document. Packmind's is marked unverified — no source — because app.packmind.com does not resolve and the product is sales-gated.
Evidence · click to enlarge
Three market patterns
- The composition layer is consolidating, fast. Of five hard competitors, one had its hosted layer switched off (Continue, acquired by Cursor in June 2026 — the open client survives, frozen) and one was acquired mid-survey (Smithery is now part of Arcade.dev; we found the banner on their own page, not in any article). Note precisely what was absorbed in Continue's case: the registry, the accounts and the subscriptions. The local client and the open format were left standing.
- Trust has moved from social proof to measurement. Tessl scores every skill on quality, impact across n eval scenarios and a security scan, then shows a composite and a multiplier. Smithery shows a score out of 100. Port shows pass / warn / block. A number derived from running the thing has replaced stars and downloads.
- Nobody authors a graph; everybody derives one. Backstage's graph is read-only and filtered by depth. Terraform resolves its graph from declared constraints. GitHub derives dependents. Not one product asks a human to draw an edge.
Three differences we can hold
- Every live hard competitor sells to an organisation; the practitioner is not the buyer. The value proposition is always control over other people's work — governance, drift, compliance. A personal library that answers to nobody is unoccupied ground. Narrowed by stage 6: that was measured on funded vendors. At the open-source practitioner tier the item-level ground is occupied — what still looks empty is the set-level half.
- They all trust the network; we can trust the file. Every catalog solves cold start with curation and volume — 17,500 MCP servers, 3,000 skills. None of them makes your own accumulated material better.
- Assembly is a side effect for them and the whole product for us. Smithery composes so it can host, Tessl so it can govern, Backstage so it can scaffold. Nobody treats does this set actually hold together as the product.
Stage 1 — conclusion
The market has converged on measured trust, and we cannot measure anything — we have no way to run a skill. So the honest move is to claim nothing per item: an item card carries usage facts from the user's own library (used in 3 projects, 2 items require this) and never a score, rating or badge.
Rejecting the node canvas turns out to be the market consensus rather than a compromise — fifteen products, none of which asks a human to draw an edge.
And the unoccupied ground is real but unproven. Continue is the only data point on whether an individual practitioner pays, and it reads both ways: the monetised hosted half was switched off and erased; the free local half is still on 1.58M machines.
Flows
Twelve mechanisms, captured from products that are actually running — behaviour, not appearance.
Each flow was captured from live software: public web where possible, the owner's own signed-in sessions where a product needed one, local CLIs in scratch directories. Read-only throughout — nothing was created, published or left changed. Ten closed, one declined, one deliberately handed forward to the design phase.
Open any row for what it taught and the captures behind it.
01Item detail and trustWhat evidence does a catalog put on a single item?declined
Trust has moved from social proof to measurement — a number derived from running the thing. Tessl scores per eval scenario with a delta; Smithery shows a score out of 100 and a verified badge. We can run nothing, so we ship none of it. Held open only by Agentman's login-walled page, then declined: its lesson is legible from their marketing, and version history is out of scope.
Captures · click to enlarge
research/2-flows/01-item-detail-and-trust/
02Library browse, filter, searchFinding one thing in a large personal collectionclosed
Typing flattens a filter tree into breadcrumbs — the path is shown rather than traversed, which is how two leaves with the same name stay distinguishable. Our mcp kind and an mcp tag will collide in exactly that way. A filter chip is written as a sentence, with Clear and Save in the same corner.
And the one to copy outright: when a filter empties the list, Linear says “4 issues hidden by filters · Clear Filters ✕”. That is the difference between there is nothing and you are not looking at it, stated as a number you can act on.
Captures · click to enlarge
research/2-flows/02-library-browse-filter-search/NOTES-linear.md
03Relations without a canvasShowing dependencies nobody drew by handclosed
Nobody authors a graph. It is derived from declared relations and made legible with a depth filter — the same graph at depth 1 and depth 3 is the whole legibility mechanism. GitHub's Dependents view is the reverse direction: blast radius rather than dependencies, which is the shape of our used in 3 projects.
Captures · click to enlarge
research/2-flows/03-relations-without-canvas/
04Validation — pass, warn, blockOur wow moment, and the best-covered flow in the researchclosed
Four products, four different lessons: Terraform on what a result reads like, Vercel on what a process reads like, GitHub Actions on partial failure, Port on a grade.
The structural find is Vercel's: a deployment is a stack of collapsed stages, each carrying its own glyph, verdict and duration, each expandable. A stage that did not run gets a clock, not a failure.
The second find decided our severity model. Port does not block — it gates a level. A failed rule stops an entity climbing a cumulative Basic → Low → Good → Great ladder rather than forbidding anything; the demo has no blocking surface at all.
Captures · click to enlarge
where "Open Critical Vulnerabilities" = 0 · Value: 1. The condition required and the value found.research/2-flows/04-validation-check-results/ — NOTES-vercel, NOTES-port, NOTES-github-actions, terraform-plan-output
05Linked vs detached, and blast radiusThe single most spec-bearing capture in the folderclosed
Drift is computed precisely and displayed nowhere. Figma's context menu names the exact overridden property — Reset fill, not a generic “reset overrides” — so the diff is already computed. And yet a modified instance is pixel-identical to a clean one in the layers tree, the properties panel and on the canvas. The only evidence is two rows appearing in a menu you must open on an object you must first select.
The failure is display, not modelling — which is exactly the trap our detached, locally modified state is waiting to fall into. Two things to take whole: revert at two granularities (whole item, single field), and Reset kept next to Detach as the two halves of one axis.
And the part with no prior art at all: a path back. Figma erases the origin at detach and offers nothing afterwards.
Captures · click to enlarge
Reset fill names the exact property that differs. The product knows; the canvas never says.research/2-flows/05-linked-vs-detached/NOTES.md
06Export and target adaptationOne source, many agent formatsclosed
External repos are always instructions, never vendored. Ruler — an open-source CLI that syncs one directory into 20+ agent formats — was run locally as a substitute for the sales-gated competitor, and it supplies two mechanics: a managed START/END fenced block so a re-run finds and replaces exactly what it wrote, and a .bak beside every overwritten file. create-next-app arrived at the same two mechanics independently, which collapses our four export targets into two mechanics plus a naming convention.
What neither does: say what it is about to overwrite.
Captures · click to enlarge
research/2-flows/06-export-and-target-adaptation/ruler-per-agent-output.md
07Env variables and secretsHow to ask for a key without leaking itclosed
Ask the type before the value, and default to the irreversible option. Vercel's drawer offers Secret and Config as two explained radio cards with Secret — the one you cannot read back — as the default. The Note field's placeholder poses the real question: “Where to rotate, or who to contact”. Two bulk paths exist, including pasting a whole .env into the Key field.
The counter-example sits one screen away and is this research's standing example of waste: the env-variables empty state keeps a search box, four filter dropdowns and a sort control on screen, all filtering nothing.
Captures · click to enlarge
research/2-flows/07-env-and-secrets/NOTES.md
08Empty state and cold startThe gap the research carried from the beginningclosed
A brand-new Linear workspace does not open on an illustration inviting you to create your first issue. It opens on the Issues list, already populated with four real issues — Get familiar with Linear, Set up your teams, Connect your tools, Import your data — ordinary objects you can edit, complete or delete. The onboarding checklist is the data model, exercised on itself, and the product is never empty at any point.
Second: three registers of emptiness, and a rule that picks between them. A concept you may never have used defines itself in a sentence. Routine emptiness says one line. Filtered to zero counts what is hidden. Verbosity scales with the chance the reader does not know what the object is.
Captures · click to enlarge
research/2-flows/08-empty-state-and-cold-start/NOTES-linear.md
09Duplicate and forkCopying as a first-class actclosed
Duplication is a distribution primitive, and the copy dialog states what will and will not come along. This is the mechanism behind our second supporting moment — duplicating a project to re-tune it for a new context.
Captures · click to enlarge
research/2-flows/09-duplicate-and-fork/NOTES.md
10Dark design languageDeliberately not spent in this phasehanded forward
The one flow about appearance rather than behaviour, so it does not close here — it opens lesson 06, concept, with its material already gathered. Linear at close range: few surfaces, with the sidebar sharing the content ground; elevation as a hairline plus a few percent of lightness rather than a shadow ramp; three foreground tones with the accent reserved for meaning; key caps as a real component with a two-cap form for chords; unset properties written as imperatives.
Captures · click to enlarge
research/2-flows/10-dark-design-language/NOTES-linear.md
11Copy — errors, warnings, refusalsEvery sentence the validation pass will ever writeclosed
Cause and consequence are two rows, and you group by the action that fixes a failure rather than the check that found it. Nobody opens a 102-job run to find seven red squares — so the best example in the survey sends the report to the reader instead: a bot posts a distilled failure summary into the pull request, grouped by the command that reproduces each failure, stamped with the commit, with raw output collapsed.
The anti-reference is industry-wide. Fourteen of seventeen annotations on one run read, in full, Process completed with exit code 1 — and the same was found in denoland/deno, withastro/astro, vitejs/vite and rust-lang/rust-clippy. The failure channel carries less information than the deprecation channel.
Captures · click to enlarge
research/2-flows/11-copy-and-error-language/ — ci-failure-copy, dependency-conflict-copy
12Visibility and portfolioPublic or private, and what a profile says about youclosed
Public/private is a low-ceremony decision, and the profile is the portfolio. GitHub puts the visibility flip in a Danger Zone with a typed confirmation; Notion separates sharing with people from publishing to the web entirely.
Captures · click to enlarge
research/2-flows/12-visibility-and-portfolio/NOTES.md
Stage 2 — conclusion
Four findings reached back into the specification and changed it.
- Cold start is a product decision, not a demo problem. Ship real objects, not an empty state — so the validation pass has something to run on before the user has typed anything.
- The drift indicator is a display problem, not a modelling one. The diff is already computed, so putting it on the card costs nothing. Figma proves it can be computed and still never shown.
- There is no prior art for the return path. Every product either keeps the link or erases the origin. Bringing a locally modified item back into the library is something we invent, not copy.
- An action that cannot act is not shown — on primary surfaces. Linear suppresses its toolbar over an empty list; Figma hides a reset with nothing to reset; Vercel keeps four dropdowns over nothing and is simply wrong.
What the flows could not see: density. Every browsing mechanic was captured in a workspace holding four issues. All of it is mechanics; none of it is scale. not established
Pain
The first evidence in this research about users rather than vendors — and the first that argues with us.
Everything up to here compared companies. Nothing had established what actually hurts. Our whole product rests on one asserted pain: nothing tells you that a skill needs a particular MCP server, that two items write to the same file, or that a key is missing — you find out at runtime. That claim had never been checked.
Two public issue trackers, queried through the GitHub API: continuedev/continue (6,677 issues — the closest dead competitor, same object model) and modelcontextprotocol/servers (1,186 issues — the ecosystem our mcp kind lives in), ranked by reactions.
What this instrument cannot see
An issue tracker records breakage, not friction. People file an issue when something is broken, not when something is tedious. This is not a caveat at the bottom of the page — it governs every number below.
| Candidate pain | Visible here? | Why |
|---|---|---|
| A · Loss | No | “I can't find the prompt I wrote three months ago” never becomes an issue |
| B · Reassembly cost | No | “I copied the same four files into a new project again” produces silent tedium |
| C · Silent breakage | Yes | This is what trackers are made of |
The loud pain is environmental, not compositional
The single most-reacted issue in either repository, by a factor of eight over anything configuration-related, is MCP Servers Don't Work with NVM — the app tries to use the wrong Node and fails. The top of that tracker is almost entirely it will not start on my machine.
Our thesis is observed, but quietly
It is there. It is just not loud:
“It's not possible to run multiple instances to connect to different databases. When doing this, the chat always chooses the first one specified in order of
modelcontextprotocol/servers #1219 — 13 reactionsmcp.json.”
That is our duplicate-key collision, reported by a user, in the exact file we generate, with the exact failure mode we predicted: the first one wins and nobody is told. The same behaviour sits in Continue's own source, where a duplication detector correctly identifies the clash and the merge silently discards the loser. Two independent sightings of one bug — once in code, once in a user's words.
And nobody is asking for a composition layer
The uncomfortable one. In a 6,677-issue tracker belonging to a product that shipped a hub for exactly this, a search for reuse blocks assistant returns 0 results. Set beside the post-mortem — the hosted hub is the half that was switched off, the local client is the half still running on 1.58M machines — that is two independent signals pointing the same way.
Stage 3 — conclusion
Keep the validation pass. The failure is real, currently silent, and sighted in the wild. Nothing here says stop.
Re-weight the export. Our specification gave the generated setup instructions a single line, while the loudest pain in the ecosystem lives exactly there: the archive lands on a machine and does not run. Most of that is beyond our reach — we validate a set, we cannot fix someone's PATH — except at the one place our output meets their machine.
Do not lead with sharing. Two independent signals against it.
And say plainly what is still unknown. Whether loss or reassembly cost is what would make someone adopt this is invisible to trackers by construction. It needs a different instrument — asking people — and we have not done it. not established
Benchmark
For each of our four core flows: who does it best in the world, how well, and against what standard?
Stages 1–3 produced observations — Linear counts what its filter hides, Port prints the condition and the observed value, Figma computes drift and never draws it. Those are anecdotes until they are turned into a scale. This stage builds the scale, and stage 5 spends it. The five categories are lifted from findings in stages 1–3, so the rubric is grounded rather than invented.
| Category | The question it asks |
|---|---|
| C1 · State legibility | Can you read the current state without acting on it? |
| C2 · Consequence | Before an irreversible step, is the cost stated in advance? |
| C3 · Failure copy | Does a message name the item, the rule and the observed value? |
| C4 · Recovery | Is there a way back, offered where the problem is? |
| C5 · Economy | Is anything on screen that cannot act, or missing that must be? |
Anchors. 1 — actively misleads. 2 — the information does not exist. 3 — correct, but you must go looking. 4 — present where you need it. 5 — you could not miss it, and it changed what you did next. A dash is not a zero: it means the flow contains no instance of what the category grades, and the cell always says whether that is the product's fault or the method's.
| Candidate | C1 | C2 | C3 | C4 | C5 |
|---|---|---|---|---|---|
| B1 — Find one thing in a large personal collection · our Library | |||||
| Linear ⌘K + filters | 5 | — | 4 | 5 | 5 |
| GitHub code search | 5 | — | 3 | 3 | 4 |
| Obsidian quick switcher | 4 | 2 | 2 | 5 | 4 |
| VS Code palette | 4 | — | 2 | 2 | 4 |
| B2 — Assemble a set under constraints · our Project | |||||
| VS Code workspace + trust | 5 | 5 | — | 4 | 3 |
npm install | 3 | 4 | 5 | 4 | 4 |
| Figma instances + library | 2 | 4 | — | 3 | 5 |
| B3 — Check a set and report what is wrong · our validation pass | |||||
terraform plan + validate | 5 | 5 | 4 | 4 | 5 |
| VS Code Problems panel | 5 | — | 4 | 4 | 5 |
| Vercel build log | 5 | — | — | — | 4 |
| GitHub Actions run | 5 | — | 1 | 2 | 3 |
| B4 — Produce an artefact and hand it to another machine · our Export | |||||
| Vercel deploy | 5 | — | — | — | 3 |
create-next-app | 4 | 4 | — | 2 | 3 |
| Figma export dialog | 4 | — | 4 | 3 | 3 |
| Ruler per-agent output | 3 | 2 | — | 4 | 4 |
Who wins. B1 — Linear, and not narrowly. B2 — VS Code, on Workspace Trust. B3 — terraform, with VS Code's Problems panel level on everything but the summary line. B4 — nobody. The highest cell in the whole flow is a 4.
Three mechanisms we are taking into the MVP
01 — The run as a stack of stages · Vercel
A deployment is a stack of collapsed stages, each carrying its own glyph, verdict and duration, each expandable. A stage that did not run gets a clock, not a failure. The log header states 66 lines before you read it, and a timestamp hover gives relative to start and relative to previous — which is how you find the step that hung without doing arithmetic.
Why it fits. Our validation pass is specified as a designed, legible moment rather than a spinner, and this is that moment's structure, already proven. It also gives a home to two things nobody else has a place for: our Skipped severity maps onto Vercel's clock glyph, and the handover disclosure becomes two more stages rather than a new surface. Terraform reinforces it from the other side — no progress theatre; text that appears when ready and reads correctly frozen.
02 — N of M, and a count that is a link · VS Code
Workspace Trust shows two columns — In a Trusted Folder against In Restricted Mode — with the current one outlined. Two of its lines are counted and hyperlinked: “95 workspace settings are not applied”, “10 extensions are disabled or have limited functionality”. It is the best consequence disclosure in the entire benchmark, and it beats Figma's 423 instances on the one axis Figma leaves open: the number is a link to the list.
Why it fits. We refuse to block an unclean export and confirm it instead, which puts the whole weight on the confirmation sentence. This is the shape that sentence takes: name the count, and let the user open it. The same grammar showed up in four unrelated products — 4 issues hidden by filters, Showing 0 of 6 placed inside the filter input, 0 of 0 selected, 10 files (324 ms) in repo ✕. The number you can act on, next to the control that produced it — and never the word “none”.
Captures · click to enlarge
Showing 0 of 6 inside the input, and the remedy as a link.03 — The failure line, written where the rule is known · npm and Port
npm's ERESOLVE answers all four questions a conflict message owes: what was required, what was found, who required it, and who asked for what was found — which is precisely our added manually versus added as a dependency. Then both escape hatches, with the cost in the same sentence: “to accept an incorrect (and potentially broken) dependency resolution.”
Why it fits, and it is not a copywriting decision. Compare the 1-grade message: Process completed with exit code 1, under a header counting “11 errors and 6 warnings”. The difference is architecture. The 5-grade messages are emitted where the requirement is known; the 1-grade one is emitted by a process that only knows it stopped. So: the check that knows the rule must be the thing that writes the sentence.
Two refinements ride along. Count problems the way the user counts them — Terraform emitted four diagnostics for one defect, and VS Code reported a duplicate key as two unjoined rows that never mention each other. And a check that could not run must say so, which is why Skipped exists as a third severity rather than a silently missing row.
One mechanism that will not work for us
Figma's export dialog with nothing selected reads 0 of 0 selected beside a greyed Export button, explained by one good sentence: “No selected layers have export settings. Click + in the export section of the properties panel to add one.” It scores respectably — the state and the remedy are both stated precisely.
It works there for three reasons, and we have none of the three:
| In Figma | For us |
|---|---|
| The blocker is one named action away | Our problem may be four items away, on another surface |
| The condition belongs to this second's selection | Our problems are properties of the project |
| Re-exporting costs nothing | The archive is the product's whole point |
A greyed primary action is a dead end on the most important control in the product, and our promise is the system checked, not the system forbade. So the mechanism is refused outright: export is never disabled; an unclean set is confirmed, not blocked.
Captures · click to enlarge
Stage 4 — conclusion
Handover is the industry's weakest flow, and it is ours to win. No cell in B4 scores above 4. Ruler overwrites without warning; create-next-app ends on the word Success! and lists nothing; Figma's dialog disables its own primary action. Our thirty-second archive is aimed at the least-well-served flow of the four, and the specific gap is disclosure before the write.
But this rubric grades craft, not weight. The categories came from surveying products, so every score answers did this interface tell me the thing — never does anyone bleed here. Read against the pain evidence, the weights come out uneven: B1 and B2 are where the craft is, and a tracker is structurally blind to both. B3 and B4 are where the value is — one carries the only sighting of our own thesis, the other carries the loudest pain in the ecosystem.
So a category score alone cannot choose a shape. It has to be read next to which third of the flow that shape owns.
Patterns
Five radically different shapes for the key flow — assemble a set, check it, export — scored on the rubric stage 4 built.
Not five layouts of one idea. Five framings, each carrying prior art from the earlier stages, so the comparison is between things that exist rather than between guesses. Each answers the same five questions in the same words, and each is scored on does this shape give the category a natural home.
| Shape | C1 | C2 | C3 | C4 | C5 | The idea |
|---|---|---|---|---|---|---|
| P1 · Two-pane drag | 4 | 3 | 3 | 4 | 2 | Library left, project right, items dragged in |
| P2 · Command-first | 3 | 4 | 5 | 4 | 5 | No library pane; ⌘K adds by name, the project is a list |
| P3 · Document | 4 | 5 | 5 | 3 | 4 | The project is an editable manifest, validated like a linter |
| P4 · Staged wizard | 3 | 4 | 4 | 2 | 2 | Target first, one step per kind, then resolve and export |
| P5 · Run-centric | 5 | 5 | 5 | 4 | 5 | Check takes the whole surface; export is its final stage |
P5 takes straight fives on four categories and still cannot be the answer alone, because it does not assemble anything. P3 scores best on assembly and dies outside the scores. The constraints that were never up for negotiation: desktop-first, dark from day one, not a node canvas, single user, local only, custom design system.
The choice — a hybrid, stated as a choice
P2 wins the spine. P5 becomes the check-and-export surface. P3 donates one mechanism.
Neither half is a compromise — the two shapes win different thirds of the same flow, and the seam between them is a single control: Check.
Surface 1
Library
Browse, filter, add, edit. Two scopes — your own items and a curated read-only public shelf. Usage facts per item. Not part of the builder flow.
Surface 2
Project
The set as a list, not a canvas. ⌘K adds by name and opens on related items. Every row carries its own state, including detached with the differing fields named.
Surface 3
Run
Entered by Check, takes the whole screen. Stages with verdicts and durations. Export is the final stage, not a button beside the check.
Why this one, for our context specifically
- It is the only shape that gets better as the library grows. A palette is indifferent to list length; a pane is not. Drag fails the concrete test — item #250 dragged to a target scrolled out of view — and every mitigation for that is click-to-add, at which point the two-pane shape has quietly become the command-first one with an extra pane.
- It obeys our own economy rule instead of breaking it. We refuse controls that cannot act — that is why the visibility toggle is not shipped. A library pane cannot act while a validation pass is running. Applying our own rule to our own design is what removed the pane.
- It puts the wow moment on a whole surface. Export as the final stage makes the unclean-export confirmation the next row in a list the user is already reading, rather than a dialog interrupting a flow.
- It renders the six item states once, not twice. In a two-pane shape every item renders as a library row and as a member, so detached — a per-project fact — is invisible in the pane where you browse.
Stage 5 — conclusion
The substantive change: the library leaves the builder, and drag stops being the verb. The earlier specification put a filtered library sidebar inside the builder and made drag the mechanism. It does not survive 300 items, and a pane that cannot act during a check is the mistake we refuse everywhere else.
The cost, stated plainly. With no library pane you cannot see what you are not using. Three things now carry that weight, and if all three fail this choice was wrong: the palette opening cold on related items rather than an alphabetical index; the Library one keystroke away and remembering where you were; and per-item usage facts doing the work a visible pane would otherwise do.
What we decided, and why
Ten decisions came out of the phase. Each is a product commitment with evidence behind it, not a preference.
.mcp.json. The archive will contain only one of them.”What we still have to find out
Nine gaps, each with a hypothesis written so it could be shown false. Where nothing in the research supports a claim in either direction, it says so rather than filling the space.
Nobody was ever asked anything
not establishedBoth trackers are structurally blind to loss and to reassembly cost; the phase read vendors and artefacts, never users.
HypothesisReassembly cost, not loss, is what converts — a practitioner adopts to stop re-copying the same four files, and finds searching my own corpus valuable only afterwards. Falsifiable in five conversations.
from Stage 03 Pain · Stage 04 Benchmark
Handover has no prior art to copy
No candidate scored above 4, and not one cell in the matrix scores what a product says about the machine its artefact lands on — because no candidate had such a surface. Narrowed 10 Sep 2026: four open-source tools were installed and run, and one of them checks git, authentication, the Node version, 21 writable directories and PATH shadowing before anything else happens. The gap is real and its explanation has changed: nobody describes the receiving machine for a named set.
HypothesisDisclosure before the write is the whole opportunity. If the last stages of the run state what the archive contains and what the receiving machine must still do, the archive-does-not-run pain drops without us running anything there. Falsifiable: if users hit environment failures at the same rate, the disclosure was theatre.
from Stage 04 Benchmark
The return path has no prior art
Figma erases the origin at detach and offers nothing afterwards. No product in the survey lets a local override become a first-class object again.
HypothesisPromotion as a new item is safe and update the original is not, because the second spends blast radius on an action taken inside one project. Falsifiable: if users routinely promote and then immediately delete the original, they wanted a merge and we built the wrong verb.
from Stage 02 Flows, flow 05
Density was never observed
not establishedThe whole browsing grammar was captured against a workspace of four issues, and the benchmark's Obsidian cell against a vault of 27 notes. The 300-item claim behind the chosen shape is reasoned, not measured. Superseded by stage 6: four public collections were counted through the tree API at 11, 25, 47 and 48 items, two of them above forty. The 300 is refuted by measurement; density at any size is still unobserved. Design for fifty.
HypothesisThe palette holds at 300 items and the failure mode is discovery, not search — people find what they can name and stay blind to what they cannot. That is precisely the cost already accepted, so it is testable the moment a seeded library exists.
from Stage 02 Flows · Stage 05 Patterns
The public library does not exist yet
It is a decision with no content behind it: no items, no verified sources, no composed example project.
HypothesisA seed of 8–12 items that produces at least one problem and one note teaches the product better than 30 clean ones. Falsifiable on first use: if six green ticks leave people unable to say what the product is for, the seed was decorative.
from Stage 01 Competitors · Stage 02 Flows, flow 08
Licensing for redistributed items
not establishedNothing in the research covers the terms under which someone else's skill may ship inside our public library.
HypothesisA pinned reference plus visible provenance is necessary but may not be sufficient. This needs a licence review before the shelf is built — it is not a design decision.
from Stage 01 Competitors · Stage 05 Patterns
Two of five hard competitors were never seen
not establishedAgentman is login-walled, Packmind sales-gated. Their mechanisms and Packmind's monetisation are read from marketing, not observed.
HypothesisNeither changes the picture, because both sell to organisations and that difference already covers the whole group. Falsifiable if either turns out to sell to individuals.
from Stage 01 Competitors
Item versioning and history were never observed
Flow 01 was declined. Tessl pins a commit, but a version change was never seen anywhere.
HypothesisNot needed: detach plus overrides does the job for our own items, and a pinned reference does it for external ones. Falsifiable the first time someone asks what did this item look like last month.
from Stage 02 Flows, flow 01
Our own thesis is real but quiet
The duplicate-key collision is sighted in the wild with 13 reactions, against 182 for an environmental failure. Re-weighted by stage 6: the loudest number in the evidence base is 6,592, and it is about one instruction source feeding several agents.
HypothesisSilent breakage is a retention argument, not an acquisition one — it is what makes the product trusted once adopted, not what makes anyone try it. A positioning claim; the product is the same either way.
from Stage 03 Pain
What is settled, so it does not get re-opened
- The node canvas. Rejected on market consensus, not taste — not one of fifteen products asks a human to draw an edge.
- Blocking. Three severities, export never disabled, an unclean set confirmed.
- A composition layer. Two independent signals against it, plus the recursion risk.
- Per-item scores or badges. The market converged on measured trust and we can measure nothing.
- Visual direction. Deliberately not decided here. The material is gathered and the design-system phase spends it.
Captures
What was actually looked at, and where it is kept.
Every screen on this page is embedded in it as a data URI, so the page opens from disk with nothing to fetch. The originals, the ones that did not make the page, and the machine-readable logs live in the repository beside the source documents.
- Image captures
- 14334 on this page
- Source documents
- 47across seven stages
- Raw data logs
- 12queries and results
- Products captured
- 15three groups
- Flows
- 1210 closed
- Sign-in walls
- 5labelled as such
Where each kind lives. Stage 1’s 38 captures are catalogued with a line each, sign-in walls marked, in the screens index. The flow captures sit in twelve numbered folders, each with a notes file that cites the frames it read. The two issue-tracker pulls are raw JSON. The Hacker News threads were downloaded whole and read locally, and the queries that produced them are logged with their result counts, so anyone can re-run them and get the same corpus.
Two limits printed on the evidence rather than buried. Every browsing mechanic on this page was captured against a workspace of four issues and a vault of 27 notes — density under load was never observed at all. And five of the fifteen products could only be read from marketing pages and public documentation, because their product surfaces sit behind a sign-in wall; those cells say so.