Wednesday Labs
Methods note

Two bugs we found in our own IAT and sociogram scoring, and how we fixed them

We build Archives by Wednesday Labs, a research platform for psychologists running Implicit Association Tests, questionnaires and sociograms. Before we shipped the scoring code, we checked it against the published algorithms it's supposed to implement, and found two real errors. Here is what they were, why they matter, and what we changed.

September 2026 · Wednesday Labs

1. The IAT counterbalancing bug: order B skipped a flip

A standard IAT has seven blocks. Two are practice on single categories, three are the combined "critical" blocks that produce the D-score, and the whole test is normally run in two counterbalanced orders so that which side of the screen carries which category doesn't bias the result for everyone the same way.

BlockTrialsContentScored
120Target concept discriminationNo
220Attribute discriminationNo
320Combined, practiceNo
440Combined, testYes
520Target concept, sides switchedNo
620Combined (switched), practiceNo
740Combined (switched), testYes

In "order B," blocks 3/4 and 6/7 swap which pairing comes first, so half your participants see concept+attribute-A before concept+attribute-B, and half see it the other way round. That's the whole point of counterbalancing: it cancels out any advantage that comes purely from going first.

Our first version swapped the content of blocks 3/4/6/7 correctly, but left block 1 (and its mirror, block 5) exactly as in order A. Block 1 sets which side of the screen each target category sits on for the rest of the test. If it doesn't flip along with the combined blocks, order B participants spend blocks 3 and 4 with their target categories on the opposite side from what they just practised in block 1, which adds an unwanted layer of difficulty that has nothing to do with their implicit associations.

The fix flips block 1 (and block 5's key mapping) whenever order B is used, so the side each concept sits on stays consistent with whichever combined-block content that participant is seeing:

if (def.blockNum === 1) {
  // order B starts with conceptB on the left so block 5 switches back
  return {
    ...def,
    label: `${labels.conceptB} or ${labels.conceptA}`,
    leftLabel: def.rightLabel,
    rightLabel: def.leftLabel,
    pools: def.pools.map((p) => ({ ...p, key: (p.key === "e" ? "i" : "e") })),
  };
}

D is scored per Greenwald, Nosek & Banaji (2003): trials over 10,000 ms are dropped, a participant is excluded if more than 10% of their scored responses are under 300 ms, error trials are replaced with that block's own mean correct latency plus a 600 ms penalty, and the two block-pairs are pooled by their own standard deviations before averaging. We ported this from an earlier implementation and checked it line by line against the algorithm's published SAS scoring syntax before finding the counterbalancing issue upstream of it.

We also ran a simulation: 200 runs of random keypresses through the full trial generator showed D centred at zero with no counterbalancing-order effect, and an individual-level standard deviation of about 0.2, which is the reason we never show a single D-score to participants and always debrief that this measures automatic, not considered, associations.

2. Community detection that merges two groups joined by one bridge

Sociograms map who nominates whom in a group and often need to answer "how many subgroups are there, and who's in which one." A very common approach is label propagation: everyone starts holding their own label, then repeatedly adopts whichever label most of their neighbours hold, until it settles.

Label propagation is fast and usually fine, but with a fixed sweep order it has a known failure mode on networks where two otherwise-separate cliques are joined by a single bridging person. We tested this directly on a toy network: two triangles of three people each, connected by one person who has a tie into both.

Label propagation (before)

In a fixed update order, the bridge person's label sweeps across the connection on one pass and both triangles converge on a single label.

Result: 1 group reported, modularity Q = 0

Greedy modularity maximisation (after)

Starts with everyone in their own group and merges the pair of groups that raises modularity most, repeating until no merge helps (Clauset, Newman & Moore, 2004).

Result: 2 groups reported, modularity Q = 0.37

Two groups bridged by one person is exactly the structure a researcher most needs correctly detected, a liaison, a gatekeeper, someone whose absence would split a team. An algorithm that erases that structure defeats the purpose of running a sociogram at all. We replaced label propagation with greedy modularity maximisation throughout, and now test every community-detection change against this bridged-graph case specifically, alongside the underlying eigenvector-centrality routine, which has its own version of the same failure: a plain eigenvector calculation collapses to zero on networks where ties don't loop back (nobody is ever nominated by someone who was themselves nominated), so we iterate on (A + I) instead, which stays well-defined on one-directional networks too, and offer Katz centrality alongside it for the same reason.

Why this is on our website and not just in our code

Neither error would show up by staring at the code. Both only appear when you build a small network by hand, reason about what the correct answer should be, and check the algorithm against it. That's what psychology training is actually good for, and it's why we test every scoring routine in Archives by Wednesday Labs against a hand calculation or an independent formula before it ships, not just against the paper it's copying from.

If you're evaluating tools to run IATs, questionnaires, or sociograms for your own research, the questions worth asking any vendor, including us, are: what happens to error trials in your D-score, what counterbalancing order is your default, and what does your community-detection algorithm return on two cliques joined by one bridge. If the answer is "we haven't checked," that's worth knowing before you collect data.

Archives by Wednesday Labs is built by two psychology researchers (MA Counselling and MA Clinical Psychology), not licensed out of a generic survey tool. It's priced for a student's budget, not a funded lab's.
Request early access to Archives Ask us a methods question