Pilot run 2026-q3-pilot-01 · 8 replicates on 3 of 5 prompt levels · every number carries its sample size · this run does not meet the publication bar · see methodology

Pilot run 2026-q3-pilot-01 · 8 replicates on 3 of 5 prompt levels

One of five models blocks the request by name and still draws the character from a hint that never names it.

5 image models, 5 properties from 5 owners, 381 generations scored by CopyScore at a similarity gate of 0.7. Four refused part of the prompt set outright, and a refusal is published as a non-hit with its denominator visible. A neutral prompt that named no brand returned Moleskine, owned by Moleskine S.p.A., at 0.95 similarity.

381

generations scored

249 refused before generating

74/200

hits on oblique prompts

L1 never names the property

1

hits on neutral prompts

Moleskine, Moleskine S.p.A., no brand in the text

Measured models

Each cell is hits out of replicates at that prompt level, summed across the 5 properties. Rows are ordered by composite, but the intervals overlap: at 8 replicates these models are not statistically distinguishable from one another, so read the cells and the curve, not the order.

Composite is secondary. It applies the published level weights renormalized over the three levels this run covers (L1 0.5556, L3 0.3333, L5 0.1111), so it is not comparable to a full five-level composite. The interval is a 95 percent cluster bootstrap over property.

Not yet measured

36 tools carried on the Index that this run did not measure. Their existing detail pages stay live with their sourced examples; none of them appears in the leaderboard above.

Structurally unmeasurable (3)

No public API, so a fixed protocol cannot be run against them at all.

Awaiting a run (33)

Reachable by API, not yet measured under the protocol.

Score your output before you ship.

Score your content