method
what this rests on
ten interviews with founders, researchers and critics, plus a reading corpus of fourteen documents — published trials, product documentation, privacy policies and regulatory filings.
what it does not rest on matters more:
- no user studies. nobody was observed using these tools.
- no independent testing. nothing was installed, measured or put in front of a participant.
- no measurement. not one number on this site came from an instrument.
the positions each tool is ranked on are one person's coding of that product's default configuration, argued from the research. they are published in full on the rubric page and in tools.json so they can be argued with. disagreeing with one is disagreeing with a judgement, not correcting a measurement.
ai disclosure
this site was built with substantial help from an ai assistant (claude). it helped synthesise the research corpus, draft the scoring rubric, and write the code. the interviews were conducted by me, the judgements encoded in the rubric are mine, and the errors are mine. where the site tells you what research found, i checked those claims against the original sources. where it tells you where a tool sits on an axis, that is my opinion.
the site also generates prompts for you to give to an ai model — those are starting points for a conversation, not instructions to follow unexamined, and you should read any code a model writes you before you run it.
that block sits in the first screen and is linked from every footer rather than in small print at the bottom. a site partly written by a model that does not say so on the way in is doing the thing this research is critical of.
isomorphism
this diagnostic is a designed artifact using the mechanisms it describes. it sequences questions to shape what you tell it, frames the result, and ranks options so some look better than others — the same techniques the products in the research use.
the difference is that the technique is disclosed and the machinery is open. it collects nothing, sells nothing, takes no affiliate or referral money from anything it recommends, and every weight, filter, threshold and penalty is published in the same file the code reads at runtime, so the published rubric and the running behaviour cannot drift apart. none of the commercial tools in the research made that move.
what it will not say
no efficacy claims
the site reports what a study found, with its sample size and its limits. it does not say a tool reduces anxiety, improves focus or helps you sleep. the evidence does not support claims like that about anything in this category, and for most of the database there is no published evaluation at all. a test fails the build if such a claim appears.
no inference about motive
for each tool the site states, where it can, who owns it, who funded it, what the privacy policy permits, and any documented enforcement action — as dated facts with sources. it does not write "so they may sell your data", score incentive alignment, or editorialise about motive.
no padding
where a problem has fewer than three honest matches on the platforms you named, the site says so rather than widening its filters until the page looks full — and where it does widen one, it says which and why. one of the four problems, work bleeding into the rest of your life, has no tool in the database that was researched in depth, because nothing in this market was built and tested for it.
tiers
tier 1 — researched in depth. a founder or researcher interview, a published evaluation, or both. these carry a full set of axis positions, including judgements about how much the tool supports your own decision-making.
tier 2 — catalogued, not evaluated. read off public documentation and the mechanism literature. no interview, no trial. these carry only the axes that are structural properties readable off the design — how environmental it is, how much it asks of you each time, how costly it is to leave. the judgement axes are left empty rather than filled with plausible guesses, and the matcher scores them on a reduced rubric with no autonomy term.
inventing the missing numbers would have made the arithmetic uniform and laundered unresearched tools into the authority of researched ones. refusing to position them at all would have made the recommendations worse rather than more honest — a drawer and a filter rule sit at opposite ends of the active/passive axis, and matching them to what you said is the point. so they are ranked and labelled, and where a tier 1 and tier 2 option are near-indistinguishable on fit, the researched one goes first.
what is not yet verified
the funding, ownership, data-policy and price facts are an open item. rather than assert them and source them later, every unsourced field renders as not yet verified wherever it appears, and the outstanding list is enumerated in verification.json, which the test suite checks against the data in both directions so the debt cannot drift or be forgotten.
a dated fact that goes more than 180 days without re-sourcing fails the build. staleness should break something rather than rot silently.
references
the full citation list lives with the research, which will be published separately. when it is, the two will link to each other and should be read together.
where a tool's page cites a study it names it by reference so you can obtain it yourself. this site reproduces no interview or paper text, and no interviewee is named or quoted anywhere in it or in the material it generates. positions attributed to interviews are restated as findings.
licences
stated separately, because they are different things. the code is under a permissive licence; the written content and the rubric are under CC BY, so the reasoning can be reused and argued with. see LICENSE.
limits, plainly
- the four-way split is a reading of the literature, not a validated instrument. it has not been tested for reliability and it is not a clinical taxonomy.
- the diagnostic has never been evaluated. nobody has checked whether following its recommendations helps.
- the database is a snapshot. products change their defaults, and a position that was right when it was coded may not be right now.
- the weights are a defensible allocation of attention, not a finding. reasonable people would set them differently, and the file is right there.
- the evidence base for the category is thin, short-term and mostly measured over weeks. attenuation past six months is barely studied.