START
How to read this
Every claim below carries a mark. The mark says how it is known, never how much we like it.
✓Confirmed. Another person can re-run the instrument and get the same answer — a logged query with its count, a file counted in a public repository, a source read at origin. Checkable without trusting anyone’s memory.
*Practitioner-reported. Somebody said it in an interview, from memory, about their own work. There has been one interview, so every * on this page is one person. It is the only evidence here that reaches motive, cost or feeling at all — and it is not a fact about anybody else.
?Unknown. No instrument has established it in either direction. Every one of these is restated as a hypothesis at the end, in the form we assume X; we would check it by Y.
Two rules that stop the marks inflating. A * never becomes a ✓ by repetition — five people’s memory is five people’s memory, and only an instrument promotes it. And a re-runnable query proves that something was said, not that it is true: an issue body, a forum comment and a tool’s own README are all people describing their own behaviour, so the ✓ covers the utterance and the behaviour underneath stays self-report.
The evidence base, in one paragraph. Five public issue trackers, the largest of them 89,923 issues. 2,036 Hacker News comments downloaded and read in full, including one 274-comment thread on exactly this subject. A GitHub repository search, and four real agent-material repositories counted file by file through the tree API. About 140 captured frames of other people’s products. And one thirty-minute conversation with one practitioner. That is everything. Reddit, named across community surveys as where this population actually talks, blocks our crawler and is missing entirely.
Why there are no names, ages or photographs. The trap this stage most easily falls into is inventing a persona from the specification and then citing it back at the specification as evidence. The audience sentence in our own spec — design engineers, 25+, visually literate, living in Linear, Vercel, Raycast, Figma — was written before a single user was observed, so it is a claim under test, not a source. Each persona below is named by what it does, and the passport is left blank because there is nothing honest to put in it.
The blanket label is gone and every mark stayed exactly where it was. This page used to carry one word — provisional — over all of it, lifting on five practitioner conversations, of which one happened and four became unavailable. On 15 Sep 2026 the register dropped the word rather than meeting its trigger, on the rule the evidence system is built on: the mark is per claim, not per document. One word both over-warned about the primary persona, which carries 20 confirmed against 5 unknown, and under-warned about the empty-handed one, which carries 2 against 6 and is the persona two shipped surfaces already serve. Nothing was promoted: every one practitioner said so is still that, every unknown is still unknown, and the four conversations are still owed and still unavailable. Read each card’s own what this card is allowed to settle line instead.
STAGE 06
The evidence
What was read before any of this was written, and the seven things it turned up.
Phase 01 established what vendors sell and what breaks. None of it established who the person is. This phase reads the public record with new instruments and then, for the first time, asks somebody. It is unfinished: four more conversations are owed.
What was read, and what could not be
The Claude Code issue tracker — 89,923 issues, which is the tracker of the product this audience actually uses and which stage 3 never opened; the Codex and Gemini CLI trackers; 1,762 Hacker News comments downloaded from eight threads and read in full; Stack Overflow; GitHub repository search. Reddit blocks our crawler, so the largest venue in the space is missing by technical accident rather than by judgement.
Seven findings
One source, many agents ✓The loudest demand in the entire evidence base
6,592 reactions on a single request to support a shared instruction file, because “it doesn’t work as well when collaborating with developers who aren’t using the same agent”. Thirteen issues across three vendors say the same. That is thirty-six times the 182 stage 3 called the loudest pain in the ecosystem — and it is what our export target selector already aims at.
Reassembly cost ✓Filed after all — stage 3 looked in the wrong shape
The 0 results that concluded “nobody is asking for this” was a bug-shaped query on a product whose hosted half was already dead. On live trackers the friction is filed as feature requests: “each plugin must duplicate these resources… maintenance burden… copies can drift out of sync”, and “no modular reuse across projects. A developer working across 10+ repos…”
Eleven to forty-eight ✓Counted, not reported
Four public collections were read through the tree API and counted: 11, 25, 47 and 48 items, two of them above forty. That is the first time this research measured somebody’s collection instead of asking about it, and it contradicts the 300-item figure the chosen shape’s accepted cost was priced against. It also retires our own “twenty to forty”, which blended one organisation’s management threshold with one person’s file count. Tens, not hundreds — design for fifty. The instrument sees only people who publish; the one practitioner interviewed keeps his private.
A git repo, symlinked ✓ *Where the material actually lives
XDG compliance at 446 reactions; a request for a git repo as the source of truth at 151 — and a request is evidence that this is not yet how that person works. Twenty-five comments say they symlink one source into many formats; nobody was watched doing it, so that half stays self-report however often it is repeated. The private repo symlinked into the agent directories is one practitioner’s own setup. What people do today is cp -r, symlinks, include directives and hand-written scripts — including one person who built the automation and still copied by hand.
The real doubt ✓ *Not “is my set broken” but “does any of this do anything”
“Mostly useless… 50/50 or less that it even reads the file.” “I can never quite tell if it’s helping anything.” An entire thread titled I am morally opposed to updating my instructions file. And from the interview: “Maybe half of it, if you want the real answer… I’ve never A/B’d anything.” Our check answers does this set cohere. That is a different question, and it is now an open register entry.
Other people’s skills carry real risk ✓ published13.4% critical, on somebody else’s scan
A security study scanned 3,984 public skills across two registries: 534 with critical issues, 76 confirmed malicious payloads. We read the study at origin; we did not re-run it and cannot — the vendor sells security tooling, and its own post does not reconcile the headline figure with the table beneath it. The order of magnitude is what this carries. “The barrier to publishing? A markdown file and a GitHub account one week old. No code signing. No security review. No sandbox by default.” We plan to ship a curated public shelf; checked sources now has to mean checked for content.
The ground is no longer empty ✓— at the item level, and since 10 Sep 2026 run rather than read
Open-source skill managers with 425 to 4,526 stars say in their READMEs that they audit duplicates, pin registry commits, score trust out of 100 and detect version drift. One of them already claims our own conflicts never block, you choose principle. That was self-description until 10 Sep 2026, when four of them were installed and pointed at a deliberately broken set. The set-level half held: none of them has a set-level unit at all, so none detects a duplicate command name, a shared config file or a missing key across a set. Two claims did not survive contact. One tool describes the receiving machine — a surface the benchmark said nobody had. And the duplicate audit, given one skill copied into three places with one copy edited, skipped a whole directory it does not scan and called the other two “identical copies” — 8,072 bytes against 8,088, different hashes — then offered to delete one.
Then every claim in the personas and the jobs was audited one by one — 238 of them, 163 confirmed, 48 hypothesis, 27 invented — and the five questions that audit raised were taken back to the public record. That second round counted four real agent-material repositories file by file, which is the first instrument here that measures somebody’s collection instead of asking them how big it is. It moved five marks and changed one persona’s character.
STAGE 06
The three personas
Split on behaviour we can actually see, never on demographics.
Two axes separate them, and both have evidence at both ends. Does anyone else ever open this material — 182 ranked issues across two trackers contain not one about a team, against 6,592 reactions on a request whose stated motive is collaborating with developers on other tools. And how many tools the same material has to serve — 245 reactions for bring it to the editor I already use, against 6,592 for one source for all of them. A third axis, author against consumer, produces the third persona and carries a warning printed on it.
P1 · Primary
The keeper who runs several agents
solo, until it is not2–4 toolswrites their own
Standing: ✓ and * agreeing from independent instruments — the only persona of which that is true.
Context — where they arrive from
They did not set out to build a library. They copied one file into a second project, then a third, and at the fourth they made a repository. *
Fourteen months later it holds tens of items, not hundreds — the counted band is 11 to 48 across skills, agents, commands, instruction files and configs. ✓ Four public agent-material repositories read through the GitHub tree API: 11 items, 25, 47, 48. This card said 20 to 40 until the count was run, and two of the four exceed 40.
They arrive from one of three situations, and only three, because those are the only three anyone has described: copying something out into a new project; adding a rule right after an agent did something annoying; hunting for something they know they wrote and cannot find. * — and they never go in to read or review, because nothing prompts it.
Environment
- Two to four tools installed, two used on a given day. ✓ Counted, not reported: one repository configures Claude, Devin, Copilot and Pi; another Claude, Codex and Antigravity, with
CLAUDE.md, AGENTS.md and MULTI-AGENT.md side by side.
- The material lives in a git repository they own, symlinked into the places each tool expects. ✓ Visible in the trees, and described independently by six people:
chezmoi, a .agents/skills/ directory plus a symlink, two source-of-truth repos syncing to every tool. Bias printed: this counts only people who publish, and the one practitioner interviewed keeps his private.
- Three machines or more, counting containers — WSL, dev containers, Remote-SSH, Docker. ✓ *
- Nothing is public. ✓ 0 of 1,762 comments mention a portfolio; the same search in the tracker returns two unrelated issues.
Jobs — what they are trying to do
- Get the same rules to hold across every tool, without maintaining N copies by hand. ✓ 6,592 reactions — the loudest demand in the entire evidence base. 161 more on users have to add the same skill twice.
- Start a new project without spending the same forty minutes again. ✓ * Filed as a feature request, not a bug, which is why a bug-shaped search missed it.
- Stop the copies from diverging. ✓ *
- Reuse something they already built without dragging its old project with it. * One person, and the only evidence this question has ever had.
- Delete half of it with confidence. *
Pains
- It lands on a fresh machine and quietly does not run. ✓ 182 reactions, 91 comments. * A container on Node 18 against a server needing 20+: detection took twenty-five minutes, the fix took two.
- The failure does not announce itself. ✓ Instructions past 32 KB silently dropped, “with no warning anywhere”; skills present and not noticed, 75 reactions.
- Two things claim the same slot and the first silently wins. ✓ 13 reactions — our own thesis, sighted in the exact file we generate. Quieter than the environmental pain by a factor of fourteen, and that ratio is the honest weighting.
- Keys and secrets. ✓ 192, 54, 45, 32, 23, 22 across two ecosystems — the largest crowd of its kind in the corpus.
- Not knowing whether any of it does anything. ✓ * “Mostly useless… 50/50 or less that it even reads the file.” An entire thread titled I am morally opposed to updating my instructions file. And: “maybe half of it, if you want the real answer… I’ve never A/B’d anything.”
Trust triggers
Convinces. A usage fact about their own material — and this is the one thing a practitioner asked for in our own words, unprompted, with our vocabulary forbidden in the question: “this file was loaded in 40 sessions, this one in 2, this one never. That alone would let me delete half of it with confidence.” Then he extended it past our spec: last actually useful, counted per session. * — the strongest form a single report can take, and still one person.
Convinces. A number that is also a link to the list behind it, and a consequence named in the present tense before the irreversible step. ? Evidence of what good products chose, not that it works on this person.
Repels. A score we invented, when we can run nothing. A green tick that was not earned. Being reported on. And — the sharpest — a check that answers a question they did not ask: asked what one thing a genie should fix, before the product was ever described, he said “show me what actually loaded and what actually mattered… I don’t need it to fix anything, I need to see it.” *
Quote
“I have a ~200 line file of style rules that I copy and paste between all my projects (There’s got to be a better way to manage files like that!) but I can never quite tell if it’s helping anything.”
Hacker News, 22 Jun 2026 ✓ — both halves of this persona in one sentence: the reassembly that built the collection, and the doubt that it is worth anything.
What this card may settle
The run’s stage list and the verdict each stage carries · the copy register for a problem and for a note · what an item card carries, specifically last actually useful rather than last exported · the disclosure stages before the irreversible step · how prominent the tool selector is. It may not settle anything about the library at scale, or about positioning — both rest on unknowns.
P2 · Secondary
The receiver
receives somebody else’susually a different toolconsumer here
Standing: ✓ for the population — six independent senders and three issue threads. * and second-hand for the receiving end.
Context
They are handed a repository, a config or a project and expected to work in it, usually on a different tool from the person who wrote it. ✓ The most-reacted issue anywhere in this evidence base — 6,592 — states exactly this as its motive.
The sending side is confirmed and says it is hard. ✓ Company-managed skills pushed to every developer by a bootstrap script, with the sender adding “distributing them is currently awkward”; skills committed to git “so complete team leverages them”; a public skill registry with profile-based syncing for the Norwegian Government. Separately: 44 comments matched, 6 concerned more than one person and 38 were solo setups.
Jobs
- Get the thing running today, without becoming an expert in somebody else’s setup. ✓
- Find out what is actually load-bearing. * What he had to explain, twice, on a call: “which files are load-bearing and which are aspirational. Which rules are real constraints from the client versus my personal taste.” None of it is written down anywhere.
Pains
- The handover fails silently and the receiver does not know it failed. * A contractor got the repo: paths broken, one server not starting, and “he assumed that was normal and worked around it for two days without mentioning it.” Second-hand — the contractor was never asked.
- Half of what makes the setup work was never written down at all. * “Roughly half of what makes a project go well is stuff I’ve never written down because writing it down felt too obvious.” If it holds, it bounds the ceiling of the whole product, because we validate the written part.
- A rule from the author’s world colliding with the receiver’s. * “My rule says one thing, the client’s linter says another, and there’s no precedence anywhere.”
Trust triggers
Convinces. Being told what the machine still has to do, before anything is written. ? No product in the fifteen surveyed has such a surface, and it has never been tested on a receiving agent or a receiving person.
Repels. A set that appears without consent — 142 reactions on servers silently synced “without any opt-in, notification, or consent”. ✓ And severity that lies: the nearest competitor filed a block that failed to resolve as non-fatal, so a missing dependency degraded the result quietly. ✓
Quote
“Codex, Amp, Cursor, and others are starting to standardize around a shared instruction file… It doesn’t work as well when collaborating with other developers who aren’t using the same agent.”
6,592 reactions, 389 comments ✓ — the most-reacted issue anywhere in this evidence base is about the receiver.
And the sentence that is this product’s thesis written by somebody else: “skills need to be edited across projects and across team members in a controlled way. Git is of course required for this but is not enough.” ✓
What this card may settle
What the setup document states and in what order · which stages sit before the irreversible step and what they disclose · the wording of a note about a missing key. It may not settle anything about how the receiver feels: six voices arrived and every one of them is a sender. The receiving end is exactly as unobserved as it was.
P3 · Secondary · changed 8 Sep
The empty-handed
has nothing yettools unknownconsumer
Standing: ? that the persona exists in numbers — ✓ that the nearest observed people are sceptical of the premise behind it.
Read this card as a warning, not as a portrait. The specification already ships two surfaces and a content commitment for this person. This card used to say that nobody in this shape had ever been observed. That stopped being true on 8 September, and the first evidence runs against the persona rather than for it.
Context
The supply is confirmed and enormous. ✓ 19,703 repositories match a search for published skill collections, topped by curated lists at 74,686, 25,709 and 15,006 stars. Against that, personal agent-material repositories that are public number 93.
The demand, from the first five practitioners ever observed on the question, is sceptical. ✓ Four of five refuse it, minimise it, or prefer their own. The one positive was distributing his own material through a marketplace mechanism — which is P1’s job, not this one’s.
Jobs
Unestablished, and the first opinions point the other way. ? The plausible ones — get something working without composing it myself, see what good looks like before writing my own — remain untested, and the observed preference is a few items from named authors over curated volume.
Pains
- Taking other people’s material is measurably risky. ✓ Published, not verified by us: 3,984 public skills scanned, 13.4% with critical security issues, 76 confirmed malicious payloads, and “the barrier to publishing? A markdown file and a GitHub account one week old. No code signing. No security review. No sandbox by default.” We did not re-run the scan and cannot. The order of magnitude is what it carries, and it is enough.
- Everything else. ?
Trust triggers
Convinces, on the best first-run capture in the phase. A brand-new workspace that opens already holding four real, deletable records, so the onboarding is the data model exercised on itself. ? that it convinces this person.
Convinces, presumably: provenance. Origin shown and a pinned version, which is what keeps a shelf with no server behind it honest — this is the version we checked, not this is current. ?
Quote
“I’m not sure I’ve ever used any of them, and when I’ve looked at them it’s been some YouTuber trying to make money.”
Hacker News, Ask HN: How do you manage skills files? ✓ — this box was empty until 8 September and the absence was the finding. What filled it is a refusal.
Read what this is and is not. Four sceptics on Hacker News are not a market, that venue is the one most predisposed to distrust a marketplace, the thread self-selects for people already managing their own skills, and nobody asked a beginner anything. Everyone quoted keeps their own material, which makes them sceptical consumers, not empty-handed newcomers. The persona is still unobserved; what is observed is the attitude of the people nearest to it.
What this card may settle
Not whether the shelf ships — that is a decision taken on reasoning. What it may now be cited for is a warning: the nearest observed population prefers provenance over volume, which bears on what the shelf contains and how it sorts. And it settles one thing outright: the shelf needs a stated content-review standard, because the one hard number attached to this persona is 13.4%.
STAGE 06
Why P1 is primary
The contest we expected did not happen, and the reason is worth more than the answer.
The plan expected a fight between the person whose pain is sighted — an archive that lands and does not run — and the collector the specification is positioned for. The one practitioner asked says they are the same person:
“Already had the collection, and that’s the point. When it was one file I knew what was in it. At forty files with overlapping instructions I have no working model of what’s active on a given run. The collection created the problem.” *
That is one sentence from one person and it is the only causal claim anybody has made. If it holds, the collector is the breaker, and P1 is not a choice between two candidates but the merge of them.
1. Both instruments agreeMost of all on P1
The persona whose jobs and pains most often have ✓ and * pointing the same way from independent instruments — the multi-tool job at 6,592, the drift job at 48, the environmental pain at 182, the collision at 13, the doubt across four threads and one interview.
2. Where the value already isCheck and handover
The benchmark re-weighted four flows by pain and put the value in check and produce. Those are P1’s two moments.
3. Higher risk, fewer leversThe method’s own rule
If we are wrong about P1, the run and the item card are both wrong, and they are the product. If we are wrong about P3, one scope switch and one seeded project are wrong.
4. P2 depends on P1Not an independent product
The receiver’s pain begins with an archive P1 produced. Designing for P2 without P1 has nothing to design against.
What would overturn this — stated in advance, so it can happen
- Four more answers that separate the collection from the breakage. If practitioners two to five say the breakage came from a single server on a new laptop and had nothing to do with the size of their collection, the merge collapses, P1 splits back into two personas, and the primary becomes whichever the remaining evidence favours.
- An answer that says adoption is driven by getting material, not by keeping one’s own. Then P3 becomes primary, and the library becomes the centre of the product rather than a browse screen.
- Four more people answering the genie question with show me what ran. Then the primary survives but the core-value sentence in the specification does not.
The bias, printed on the page. Every public source is a person who chose to write in public; the one practitioner interviewed is also a filer, so a single conversation cannot correct it. Reddit is missing. Every number spoken in the interview is a recollection — forty-something files, one in three fresh environments, about half of it — and the respondent flagged his own unreliability twice, unprompted.
STAGE 07
The jobs
A job survives a change of product. A feature does not.
The specific way this goes wrong here is not inventing jobs but laundering the specification into job form. Our spec already names four surfaces and all the mechanics, so “when my set is complete, I want the system to run a validation pass” would be the spec with a when in front of it, and the whole exercise would confirm the spec by construction. Every motivation clause below was machine-checked against a list of our own feature names; twenty-two formulations, zero hits.
The main job
When something I have already got working has to live somewhere else — a second tool, a new machine, a container, a colleague’s laptop — I want it to keep working there without me rediscovering everything it quietly depended on, so that the move costs me minutes instead of an afternoon of finding out what silently did not load.
✓ * — it merges the three loudest evidenced demands, which turned out to be one job along two axes: one source across several tools (6,592), it does not run on my machine (182), and reassembly filed as feature requests (48 and 21). The loudest number in the repository — 6,592 — sits inside it, and so does the one stage 3 called the loudest pain in the ecosystem. Carried by P1, with P2 on the other end of the same sentence.
Where this differs from the specification, said once. Our spec words the value as assembly with validation. The main job is a transfer job: assembly is what you do in order to move something, and the check is how you find out before the move rather than after. The second contains the first, and it puts the weight on the far side of the handover.
A second candidate survived, and the method’s rule says what that means. Knowing whether any of my material does anything did not fold into the main job — you could close one perfectly and leave the other untouched, and both have ✓ and * behind them. Two surviving main jobs mean two products. It is recorded rather than resolved, because it is the one we cannot build: we run nothing on anyone’s machine.
Related
four, on the way to the main one
RJ-1 · P1 sending, P2 receiving
When I am about to hand a working setup to another machine or another person, I want to know in advance what that side will still have to have, so that they do not find out by watching things quietly fail.
✓ that the need exists and is unserved — fifteen products scored, none above 4 on handover, and not one scores what a product says about the machine its artefact lands on. * for what it costs when it fails, one person, second-hand.
RJ-2 · P1
When I put together the pieces a project needs, I want to learn now what else has to come with them and where two of them will quietly fight over the same thing, so that I do not learn it three days later from an agent behaving strangely.
✓ 13 reactions — “the chat always chooses the first one specified”: the first one wins and nobody is told, in the exact file we generate. Plus the quieter family of I set it and it was ignored at 6, 3, 2, 0, 0, 0. * a formatting rule carried in from another project pushed 1,100 lines of noise into a colleague’s review branch. 13 against 182 is quieter, and that is all the count says.
RJ-3 · P1
When the same thing exists in several projects and I correct it in one of them, I want the correction to reach the others, so that I do not pay for the same mistake a second time months later.
✓ 48 reactions — “maintenance burden… copies can drift out of sync”. * the same off-by-one fixed in one of three projects, found six weeks later by a client, three hours to re-debug.
RJ-4 · P1 moving, P2 receiving
When I take something from one place to another, I want my keys and the parts that belong to a particular client to stay behind, so that I am never the reason something private turns up somewhere it should not be.
✓ the largest crowd of its kind: 192, 54, 45, 32, 23, 22 across two ecosystems. * and the part no count could give — the private and the reusable are tangled in the same files, and separating them is “a couple of hours that are never the most urgent couple of hours”.
Emotional
three — how they want to feel
EJ-1 · P1
When a tool makes a choice I did not make, I want it to tell me it made one, so that I am not left inventing theories about why my own setup behaves the way it does.
✓ 142 reactions — servers “silently synced… without any opt-in, notification, or consent”; instructions past 32 KB dropped “with no warning anywhere”. * “my guesses are unfalsifiable”.
EJ-2 · P1
When I am told everything is fine, I want to know that something was genuinely looked at, so that I am not carrying a private suspicion that the reassurance is decoration.
✓ 10 Sep 2026: a named person builds a reproduction repository, chases it for ten days, calls the rules system “rather unusable”, and asks for the one thing this job is about — “it would Really be useful if we could inspect the raw full context window… to tell if it is an issue with the AI not loading the rule vs just deciding not to follow them.” The vendor confirms rules apply inconsistently. ✓ Earlier: “Blackbox oracles make bad workflows, and tend to produce a whole lot of cargo culting” — comment scores are not exposed, so that one carries no weighting.
EJ-3 · P1
When I have been adding to something for over a year, I want to stop suspecting that most of it does nothing, so that I can keep it because it works rather than because I am afraid to remove it.
✓ * “deleting feels riskier than keeping… so the folder only grows, which is a bad property for a thing whose job is to be precise.” This is the emotional face of the second main job, and nothing we can build closes it.
Social
two, thin on purpose
This is a single-user desktop tool, and the population evidence for a social layer is an absence that was looked for twice. Inventing a social layer to fill the template is the thing the plan told this stage not to do.
SJ-1 · P1 as author, P2 opening it
When somebody else opens something I made, I want them to get moving without twenty minutes of me on a call, so that what I built is a setup rather than something only I know how to run.
✓ that the situation occurs — 6,592, 168, 151. * for the feeling: “the config had become tribal knowledge rather than a setup”. One person, and the receiver has never been asked.
SJ-2 · P1 · post-MVP
When I think about showing my work, I want it to be in a state I would not have to apologise for, so that the prospect of being seen is a reason to tidy it rather than a reason to keep it hidden.
✓ the absence, searched for deliberately in two instruments: 0 of 1,762 comments, and two unrelated issues. * a qualified yes for a different motive — “publishing forces cleanup and I need external pressure to do that”. They disagree and both stay. Out of scope either way.
STAGE 07
The matrix
Jobs against personas. An importance nobody has told us is ?, never an averaged 2.
1 = it comes up · 2 = it costs them something · 3 = it is a reason they would change how they work. Read the P2 and P3 columns first. No receiver has ever been asked, and after four rounds and five venues none has ever spoken in the first person — every account of a handover is written by the sender. P2 still carries a number in five cells: three of them are second-hand, and two were filled by watching what receivers did in public diffs, because somebody who never posts still leaves one. P3 carries a number in none of the ten rows, and four instruments have now hunted for that person and failed. That is a fact about our instruments rather than about those people, and it is the most useful thing this table produces.
Iteration history· one audit and four collection passes, 9–10 Sep 2026
How this matrix reached the state below. Each pass is dated — because the numbers moved in both directions, and which way they moved is part of the evidence.
09 Sep · the audit. Subtracted only. Two cells lost their numbers — RJ-1 and EJ-2 for the keeper — because each held an importance the cell itself called an inference, or described as carrying no weighting. Four were lowered from 3 to 2, because a 3 means a reason they would change how they work and in every case the person on record demonstrably did not: the one receiver worked around a broken handover for two days and said nothing; the one keeper still copies with cp -r; the one person afraid of leaking client material has never spent the couple of hours to separate it. One cell gained anything at all — the empty-handed column of H-J4 — and what it gained argues against the feature.
10 Sep · the experiments. A handover test, and four competing tools installed and pointed at a deliberately broken set. Not one importance moved, and the reason is structural: these instruments see machines and artefacts, and the columns of this matrix are people. Nine cells changed in the two right-hand columns instead.
10 Sep · the forums. Two vendor community sites this research had never used, where the vendor answers in public. Three cells up: RJ-2 to a 3, EJ-2 and H-J3 from unknown to 2. The first upward movement this matrix has had.
10 Sep · the diffs. An earlier pass had read 214 forks with their patches; nobody had asked that corpus about this table. A receiver who never posts still leaves a diff. Four more cells, two of them in the receiver’s column and both on behaviour rather than on somebody’s account of it: four receivers stripping the author’s credentials out of what they inherited, and 30 of 214 writing 21,266 lines of the setup manual the sender never shipped.
10 Sep · the shelves. 18 of the largest public collections, 568 issues. One more cell — and then a wall. Three real people, blocked by exactly what this product exists to prevent: a 4,386★ collection whose agents need an undeclared MCP server nobody documents, so new users cannot run them at all; a first-timer meeting a file-collision prompt he cannot answer, repeated for every agent. Not one of them fills a cell, and that is the finding. A public artefact shows an act; a persona is defined by a situation. You can see that somebody installed a stranger’s agents and could not run them; you cannot see whether they had anything of their own.
Net: thirty-six unknown cells that morning, thirty by the end. The keeper’s column from seven unknowns to three. The empty-handed column unchanged at ten — and the two clearest public beginners wrote their own material on day one, or asked to shadow a human. Neither reached for a library; neither got a reply.
Ten sourced jobs, six hypothesis jobs · wide table, scrolls sideways on a narrow screen
| Job | P1 · keeper | P2 · receiver | P3 · empty-handed | What closes it | Do competitors close it? |
| MAIN — make it keep working somewhere else |
36,592 and 182 ✓ · 25 min to detect * |
2lowered by the audit · the 6,592 motive ✓, filed by an author · second-hand * |
?nothing to move |
Tool selector; setup document written for the receiving agent; pinned versions; the disclosure stages |
No. Four handover cells, none above 4, and structurally so. Item-level sync is claimed in READMEs we never ran |
| RJ-1 — know what the other side needs, before I send it |
?the number was an inference — withdrawn |
2two days working around it, and he changed nothing * |
? |
The setup document per item; the missing-keys list; the pinned versions |
No — still the clearest open cell, and narrower since 10 Sep 2026. One of these tools, installed and run, does describe the receiving machine — git, Node, authentication, 21 writable directories, PATH shadowing. What nobody does is describe it for a named set |
| RJ-2 — what my pieces drag in, and where two will fight |
3raised 10 Sep 2026 — the vendor confirms duplicates do not collapse, and two people rebuilt their tooling around it |
?a receiver does not assemble |
? |
The dependency walk, declared conflicts, duplicate names, same-path collisions |
Partly, never for this material. The nearest dead competitor computed the duplicate correctly and discarded the loser in silence. The nearest live one, installed and run on 10 Sep 2026, calls two files with different hashes “identical copies” and offers to delete one. Adjacent domains do it excellently |
| RJ-3 — fix once, have it reach every copy |
3raised 10 Sep 2026 — a second person names the job, then tries five mechanisms and fails at all five ✓ |
? |
? |
The live link — an item lives once and projects link to it |
Solved elsewhere, dead here. Excellent for design components; the one product that shipped it for this material was switched off |
| RJ-4 — move the work without moving the secrets |
2192·54·45·32·23·22 ✓ — the largest crowd on a fear, and he has never done the separation |
2raised 10 Sep 2026 — four receivers watched stripping the author’s credentials out of inherited material ✓ |
? |
Keys collected across the set and written out as an example file before export |
The runtime half is occupied — credential proxies, vaults, key brokers, secret scanners. The handover half is not: none is about what a set carries when it leaves |
| EJ-1 — not be quietly overruled |
3142 + two more ✓ |
2overruled by absence * |
? |
Three severities, nothing blocks, an unclean export confirmed with its consequence in the present tense |
The market has the mechanism and the nearest competitor failed at it — a block that failed to resolve, filed as non-fatal, in production |
| EJ-2 — believe a clean result was earned |
2restored 10 Sep 2026 — a reproduction repo, ten days, and a request to see the raw context window |
? |
? |
A neutral glyph for a check that had nothing to check; no invented score |
They close a different job. Everyone ships a number; whether a number makes anyone believe is unknown |
| EJ-3 — stop suspecting half of it is dead weight |
3four threads ✓ · “maybe half” * |
? |
? |
Nothing — and nothing can. Usage facts answer is it used, not did it change anything. We run nothing on anyone’s machine |
The market is forming here and we are not in it — four attempts in nine months to measure whether this material works, at 524, 364, 135 and 79 points |
| SJ-1 — not be the missing manual |
2“tribal knowledge” * |
2raised 10 Sep 2026 — 30 of 214 receivers wrote 21,266 lines of the manual the sender never shipped ✓ |
? |
A setup document addressed to the agent that opens the project |
No. The same structural gap as RJ-1 |
| SJ-2 — something I would put my name to · post-MVP |
1absence found twice ✓ |
? |
? |
Nothing in the MVP, by decision |
Yes, all of them — and all of them need a server |
| H-J1 — lay hands on something I wrote months ago |
2raised 10 Sep 2026 — his prompt library is a folder of notes in Telegram ✓ |
? |
? |
The library, its filters and its search; the per-item usage facts |
Yes, and well. This is where the market’s craft is concentrated — we would compete where they are strongest and our evidence is thinnest |
| H-J2 — start from something I did before and re-tune it |
?nobody has said it |
? |
? |
Duplicating a project |
Yes, cheaply — a keystroke in one product, a form with a prefilled name in another |
| H-J3 — change one copy without touching the rest |
2raised 10 Sep 2026 — “each user should be able to slightly modify the rules” |
? |
? |
Local overrides, in-project editing, promotion as a new item |
Half. One product models overrides precisely — and erases the origin, so there is no way back. Our return path has no prior art at all |
| H-J4 — get moving with material somebody else made |
? |
? |
?the persona this was built for — and the first evidence runs against it: four of five refuse, minimise or prefer their own. Provenance over volume |
The public/private scope switch, the curated shelf, the example project |
Yes — the most crowded space in the whole survey. Every catalog solves cold start this way, and 19,703 repositories supply it |
| H-J5 — watch what actually ran · the surviving second main job |
3importance is not in doubt |
? |
? |
Nothing, and nothing can be. Same cell as EJ-3 — counted once, not twice |
Forming |
| H-J7 — the unwritten half travels too |
2raised 10 Sep 2026 — four people say it, one prices it at 10–30 minutes every switch ✓ |
?one aside — and it bounds the whole product |
? |
Nothing, and nothing can be |
No. Nobody in the survey attempts it |
CLOSE
What to build first
Importance 3 for the primary persona and not closed by the market. After the audit that rule chose one buildable job, not three — and as of 10 Sep 2026 it chooses three again, because two more cells reached a 3.
Updated 10 Sep 2026, and recorded rather than applied. Two more cells reached a 3, so five jobs now score 3 for the keeper — and with one of them a constraint on everything else rather than a thing to build, and one of them the job this product cannot close at all, the rule now selects three buildable jobs by itself. It has stopped needing the judgement the box below was written to make honest. But they are not the same three: RJ-2 enters and RJ-4 is displaced. Changing the shortlist is the owner’s call, not a research round’s, so what follows stands as written.
This section used to say “applied mechanically”, and it no longer can. Three jobs still score 3 for the keeper — the main job, not be quietly overruled, and stop suspecting half of it is dead weight — but the second is a constraint on how everything else behaves rather than a thing to build, and the third is the one job the product cannot close at all. The other two below now score 2 and are kept on stated grounds: the market is open in both, and the specification has already spent a mechanism on each. That is a judgement with the matrix as its input, not a result the matrix produces.
01
Make it keep working somewhere else
3 · 2 — one of only two rows carrying a number in both people-columns, and the only functional one. Nobody above 4 on handover, and structurally so: no candidate has a surface that says anything about the machine its artefact lands on.
02
Move the work without moving the secrets
2, kept on stated grounds — the largest crowd on a fear, plus the part no count could give: the private and the reusable are tangled in the same files. The runtime half of this problem is occupied; the handover half is not.
03
Fix once, have it reach every copy
2, kept on stated grounds — 48 reactions with a price tag attached: six weeks, a client complaint, three hours re-debugging a bug already fixed elsewhere. The one product that shipped it here was switched off, and the live link is the one mechanism the spec already has for it.
The two other 3s, and neither is a row to build — this matters more than the shortlist
- Not be quietly overruled. It is not a feature to build; it is a constraint on how every other feature behaves, and the specification already encodes it. It belongs in the core as a rule, not as a row.
- Stop suspecting half of it is dead weight. The feature cell is empty because nothing we can build fills it. This is the highest-importance job in the matrix that the product cannot close, and its honest form is a positioning question rather than a backlog item.
What might not be worth building — a list of hypotheses, not a cut list
Read the label before the list. Every feature below closes a job whose every cell is ?. That is not evidence that nobody wants it — it is an absence in instruments that cannot see presence. The trackers are blind to find and reassemble by construction, and nobody in the empty-handed shape has ever been observed. Cutting a feature on this list would be acting on our own blindness. Nothing is removed.
- The scope switch and the curated shelf — closes only H-J4, in the most crowded space in the survey, for a persona who has now been observed for the first time and four of whose five voices refuse, minimise or prefer their own material. The sharpest finding here — and it is about what the shelf holds and how it sorts, not about whether it ships.
- The example project — weaker as a feature, stronger as an argument.
- Duplicating a project — cheap, closes nothing evidenced, and the market closes it in a keystroke. A candidate to postpone, not to cut.
- Promotion of a locally modified item — after the audit it stands on nothing rather than on one aside: the aside cited is about a conflict outside the set, not a wish for a per-project override. It stands or falls with in-project editing.
- Library-wide search and filters — the market’s best-served flow, our thinnest evidence, and priced against a library size nothing supports. Build for fifty, not for three hundred.
- Per-item usage facts — not an orphan. The only specified feature a real person asked for in our own words, unprompted, and then extended.
CLOSE
Hypotheses
Every ? above, restated so none of it can be read as a finding.
Read the third column knowing the instrument is gone. Most rows below say four more conversations, and as of 9 Sep 2026 those conversations cannot be run. An unavailable instrument removes an event; it does not lower a bar, and nothing here is promoted because of it. What is reachable without a person was written down as a third round of research and run on 10 Sep 2026: the handover test on our own machine, forks read as handovers — the only way to observe a receiver without asking one — drift measured in git history, whether imported material is ever touched again, and the competing tools installed and pointed at a deliberately broken set. It closed H6, measured the behaviour under H7 and H8 without reaching the person in them, and reached none of the rows about motive or feeling. Where a row below has moved, it says so.
| # | We assume | We would check it by | Which decision it holds up |
| H1 | P1 arrives irritated, so the tone of a failure message matters more than its completeness | The same question, four more times, listening for the register rather than the three modes | The copy register for a problem and a note |
| H2 | People want a previous project back, and are stopped by entanglement rather than by not finding it | Four more conversations, asked as a situation | Duplication, and item granularity |
| H3 | A count that links to its list convinces this audience | The five conversations, shown two variants; or first use of a seeded library | Used in 3 projects on an item card |
| H4 | A score would repel rather than reassure, now that free tools at this tier ship one | The same open question, four more times, plus asking directly about the tools that score | Usage facts, never a score |
| H5 | The receiver has a different tool and cannot easily ask the author | Half closed. Senders confirmed at population. Ask a receiver — nobody has, and the guide still does not recruit one | What the setup document assumes about its reader |
| H6 | A receiving agent performs the setup from the document alone | Closed 2026‑09‑10 — the half-day was spent. Ten real pinned items with five planted faults, handed to three receivers with one instruction. Performs: yes, three of three, including cloning an external item at its pinned version. Correctly: no. With the faults undisclosed, one found them, one decided a real conflict was not one and closed it on a false claim, one saw none — and one copied the machine’s live credential into a plain file, unasked. | The specification’s most load-bearing bet — and it half holds |
| H7 | The empty-handed persona exists in numbers that justify shipping a shelf | Split in two, 10 Sep 2026. The behaviour is now measured and it is large — thousands of repositories carry verbatim copies of public skills, and 32% of 446 traced copies were edited after import, at a median of 27 days. The person is still unknown: no instrument here can say who any importer was, and someone who keeps nothing of their own remains unfound | The scope switch and the shelf |
| H8 | Their job is get something working without composing it myself | The same conversation, asked as a situation, never as a pitch. One datum against the “without composing it myself” half: a third of imported copies get edited, and the median edit lands four weeks after the import — that is use, not install-and-forget | Whether the shelf is a browse surface or a starter kit — and now also what it contains and how it sorts |
| H9 | A first check showing one real problem and one real note teaches better than six green ticks | First use of the seeded example, watched | The example project’s composition |
| H10 | Provenance and a pinned version are what make a read-only shelf trustworthy without a server | The same first use, and whether anyone clicks through to an origin. Two measurements, 10 Sep 2026: provenance does not survive a copy — only 10% of 446 copies name their origin at all — and a pinned version did real work in front of us: the receiving agent used the pins to fetch back both items the archive had silently lost | Public item cards, and the licence field the model lacks |
| H11 | The specification’s audience sentence — 25+, visually literate, living in Linear, Vercel, Raycast, Figma — describes these people | No instrument here can reach it, and none has. It stays a hypothesis through the design system unless somebody asks | Visual direction, and who the craft bar is set for |
The half-day was spent, and it paid twice. Two rounds of research, six instruments, 19,703 repositories and 2,036 comments could not touch the single most load-bearing claim in the specification — a receiving agent performs the setup from the document alone. On 2026‑09‑10 a third round stopped reading the public record and ran it. The claim holds: three receivers out of three did the work. The half nobody had thought to doubt does not: leaving the faults for the receiver to notice produced one correct diagnosis, one confident wrong one, and one silence. The disclosure the sender gets before export is the disclosure the receiver turns out to need.