Awesome Psychology Tasks / measurement map

A map of measurement claims

Every task here is catalogued. Almost none of it has been checked.

This is a catalog of experimental psychology paradigms and the implementations that claim to run them — and a running record of how much scrutiny each claim has actually had. Two kinds of scrutiny, kept apart on purpose.

Human review

0 / 11

specifications reviewed by a named expert.

Every specification in the catalog is sourced: drafted, referenced, and unreviewed. Raising that state takes a person; no machine check can do it.

Take one spec →

the wall

Machine checks

7 / 9

audited implementations carry the contrast they claim.

Each was driven end-to-end by a responder with a known policy. The 2 that failed look like the paradigm and are not wired as it.

See the verdicts →

The wall is the point. A machine check confirms a task is wired the way a spec says. It cannot tell you the spec is the right way to measure the thing. That judgement is human, so it is tracked separately and never inferred from a passing test.

Open call

The first expert review is the one that matters

The catalog's value is not that it lists things — software can list things. It is that a named person, on the record, has judged whether a measurement claim holds. Nobody has done that here yet, and the site says so on every page rather than dressing a green check up as review.

The ask is bounded: one specification, three separable claims, an afternoon. Sign-offs are named, scoped, dated, and revocable.

Start here

Browse 269 paradigms

Grouped by cognitive domain, each with its description, key references, and every indexed implementation.

23 domains

Search the corpus

One box across paradigms, specifications, and 994 implementation records.

Runs in your browser

Read the specifications

What it means to implement a paradigm correctly: the defining contrast, the design parameters that move results, and an executable test.

11 specs · 9 audited instances

Compare platforms

Where to build and host: open-source and commercial platforms, timing behaviour, pricing, and the large task collections.

21 platforms

What this catalog does not claim

A link that loads is not a task that works

994 records survive the headless launch check; 148 failed it and are held back. Neither result says anything about the science.

Paradigm labels carry a classifier's error rate

388 records are not confidently mapped to any paradigm and stay out of the paradigm pages. Of those that are mapped, the one confirmed error found so far was a mislabel, not a broken task.

Specifications are drafts

All 11 are unreviewed. They are useful as drafts and dangerous as authority, which is why their state is printed next to every one of them.

How entries get here, in full →