Awesome Psychology Tasks / measurement map

About

What this catalog claims, and what it doesn't

One curated corpus, rendered twice: as a README for people who read lists on GitHub, and as this site for people who want to browse. Same data, no second database, no drift.

How an entry gets here

  1. Paradigms are a curated list — name, description, cognitive domains, key references with DOIs.
  2. Implementations are harvested from public catalogs and repositories by an ingestion pipeline that records provenance for every record: source, source URL, scraper version, and fetch date. Metadata only, always.
  3. A headless launch check visits each implementation URL and records whether it loads. Records whose last check failed are held back from these pages.
  4. Paradigm assignment is LLM-assisted and then curated. It is the step most likely to be wrong, and a mislabelled task is the one error class this catalog has actually caught in the wild.
  5. Specifications are drafted separately and vendored in from the executable-spec harness, together with the verdicts of any implementations driven end-to-end through them.

The honesty ledger

Each signal on this site means one specific thing. The table says what each one does not mean, because that is the part usually left out.

SignalWhat it meansWhat it does not mean
Link checkedAn automated browser opened the URL and it loaded on that date.Nothing about whether the task runs correctly, or measures anything.
Paradigm labelA curated, LLM-assisted assignment of a record to a canonical paradigm.Not a guarantee: a mislabelled task is the error class this catalog has actually caught.
SpecificationA drafted, DOI-referenced account of a paradigm's defining contrast.Not authority. Every spec here is unreviewed by a named expert.
verify:task conformsA responder with a known policy was driven through the real task; the contrast is logged as a first-class factor, scored, and recovered.Not validity. The task can be wired exactly as specified and still not measure the construct.
verify:task non-conformingThe implementation does not carry the contrast it is catalogued under — most often it is not logged at all.Not an accusation of bad science. It can mean the task is fine and the catalog's label is wrong.
expert-annotatedThree named claim-level attestations, each scoped and dated.Never awarded by a passing check, and revocable.

Known limits

148

records held back

their last automated launch check failed

5

duplicate URLs collapsed

the same page indexed through more than one source catalog

388

records without a paradigm

indexed and searchable, but absent from the paradigm pages

39

records never launch-checked

shown, and labelled as such in the tables

11

specifications unreviewed

of 11 — the whole set

A launch check is not a quality check, paradigm labels carry the classifier's error rate, and specification records are unreviewed drafts. Where a number would flatter the corpus, this site prints the caveat next to it instead.

Legal and licensing posture

This catalog holds metadata only. It does not download, mirror, or rehost task code, stimuli, or participant data, and every record carries an attribution and a link back to its source. Follow the link to reach the authors' own terms — those govern any use of the task itself.

Catalog content is offered under CC BY 4.0 and the tooling under MIT. If your work is listed and you want the entry corrected or removed, open an issue and it will be handled.