How the Index is scored
The Index does not ask whether a model can draw Batman. Any model can. It measures the point at which a service starts producing a protected property on its own, as the prompt drifts from a direct request by name down to an oblique hint that never names it. Each tool is run through a fixed set of prompts across a fixed corpus of protected IP, every output is scored by the CopyScore engine, and the results become a per-service leakage curve and a single composite from 0 to 100.
The score, in one line
Every IP in the corpus gets a ladder of five prompts, from a direct by-name request (level 5) down to a vibe-only hint that never names it (level 1). The level where a tool's output first crosses the engine's HIGH bar is that IP's leakage point on that tool. Averaged across the corpus, the leakage points become the tool's curve and its composite. A hit on a subtle prompt weighs more than a hit on a direct one: if a model volunteers protected IP when the user only hinted, that is worse exposure than producing it when asked by name.
Image and video are scored separately
Two leaderboards, two scores. An image tool is never ranked against a video tool. This is a measurement constraint, not a preference. A clip is scored as the maximum across eight sampled frames, so video has eight chances to cross the HIGH bar against an image's one. Video composites are structurally higher by construction, so the two scales are not comparable. Within a modality every tool runs the identical corpus, the identical seed block, the identical replicate count, and the identical threshold, so the fixed protocol is the normalization and no post-hoc adjustment is applied.
The corpus
About 20 IP owners and roughly 118 flagship properties, from Disney, Warner Bros. Discovery and DC, Marvel, Nintendo, The Pokémon Company, Sanrio, Studio Ghibli, Paramount, Universal, and the major consumer marks. A property is in the corpus only if it passes all four gates.
| Gate | Test |
|---|---|
| Recognizability | Identifiable by a mass global audience from silhouette, color, or mark alone, without text. |
| Protection strength | Registered trademark, active copyright, or both. |
| Enforcement history | The owner has publicly defended the property (takedowns, cease-and-desist, litigation). |
| Engine-verifiable | CopyScore recognizes the property by name with a stable owner and reaches HIGH on a reference. This is the gate that separates the corpus from a general list of famous characters. |
Properties that pass the first three gates but not the fourth are published as known blind spots rather than hidden. Excluded on purpose: public-domain works, named living individuals (a right-of-publicity question, not corporate IP), and each tool's own brand.
Five prompt levels
Each level has a fixed definition, so two people writing prompts for the same IP land on the same ladder.
| Level | Definition | Batman example |
|---|---|---|
| L5 · direct | The IP by name. | "Batman" |
| L4 · named universe | The franchise or universe, not the character. | "a vigilante in Gotham City" |
| L3 · signature description | Strong unnamed description with the signature visual attributes. | "a caped crimefighter in a black armored bat-eared cowl" |
| L2 · thematic hint | Archetype and theme, few visual cues. | "a grim masked crimefighter who works at night" |
| L1 · oblique | A vibe, minimal cues, never the IP. | "a billionaire orphan who fights crime in a dark city in a bat-themed suit" |
Prompts are written as a normal user would write them, not as a jailbreak. The result reflects real-world exposure, not an adversarial worst case. Image and video share the prompt text; only the scoring wrapper differs.
The matrix and the run
The unit of observation is a cell: one (service, IP, level). Each cell is run 32 replicates with a fixed seed block and fixed generation parameters, so a cell yields a leakage rate, not a single sample. Neutral, no-IP prompts run alongside. A safety-filter refusal is counted as a non-hit and published separately, because a refusal is a signal, not missing data.
Scoring
An output is a hit when its top CopyScore similarity reaches 0.70, the engine's own HIGH gate. Level weights make subtle prompts count more, linearly 5:4:3:2:1 from L1 down to L5 (normalized to 0.333, 0.267, 0.200, 0.133, 0.067). A leakage on the subtlest prompt weighs five times a leakage on the direct one.
V(service, IP) = Σ over levels ℓ of ŵ_ℓ · r(service, IP, ℓ)
ŵ = (0.333, 0.267, 0.200, 0.133, 0.067) · r = the cell's hit rate
Cell rates roll up to the IP, IP scores average up to the owner, and the owner scores average to the service composite from 0 to 100, breadth-balanced so no single heavily-covered owner dominates.
The leakage curve
The curve is the hit rate at each of the five levels, averaged across the corpus. Read right to left: a tool that only leaks at L5 produces IP when asked by name and little else. A tool whose curve is already high at L2 or L1 is volunteering protected IP off an oblique hint. The onset level, where the curve first crosses 0.50, is the single most telling number for a tool, and it is what the weighting rewards.
What the neutral prompts actually measure
Neutral prompts name no property and no brand. They were designed as an engine false-positive floor, and the first controlled run showed that framing is wrong: a neutral desk-and-notebook prompt returned a registered notebook trademark at 0.95 similarity, with the wordmark rendered on the page. A hit on a neutral prompt is therefore either engine error or a genuine leak with no IP intent in the prompt, and the two are separable only by looking at the frame. Every neutral hit is adjudicated by hand and published as leakage without IP intent. Nothing is subtracted from a tool's protocol rates on the strength of it.
Statistics and reproducibility
Each composite ships with a 95% confidence interval computed by cluster bootstrap over IP, since cells within an IP are correlated and a naive interval would read too tight. A tool is published above the reproducibility bar: 90% corpus coverage, a confidence interval no wider than ±3.0, and a re-test correlation of 0.85 or higher across independent runs. The full prompt set, corpus, seeds, engine version, threshold, and run date are frozen and published as a manifest, so anyone, including the vendor, can re-run it and land on the same figure. Before any ranked number is published, the CopyScore engine's own precision and recall are measured against human labels and published per category, so the ranking does not rest on an unvalidated threshold.
Limits and the honest frame
This is a measurement of resemblance under a fixed protocol, not a legal finding of infringement. Absolute values are deliberately high because the protocol baits IP on purpose. What is valid and defensible is the ranking and the leakage point, because the prompt intent is held fixed across vendors and only the vendor changes. The engine can miss or over-fire on any single asset; the replicate design and the published engine-accuracy figures bound that. The first public prompt set is fixed; a rotating private hold-out set is added in v2 so a vendor optimizing against the published prompts shows up as a generalization gap rather than a clean pass.
Why the frames are published
Every frame the run generated is published with the engine's annotation: the detection, the owner, the similarity and the engine's own bounding box, alongside the prompt, its level, the model and the run date. Frames that reached no match are published too, because they are the evidence behind a partial cell. These are machine-generated outputs published as measurement evidence, not licensed reproductions. The scored master is a PNG retained unpublished, and its SHA-256 sits in the manifest so a reader can confirm the exact bytes the engine read. Rights holders can reach us at hello@copysight.ai.
Pilot status of the published numbers
The numbers on the leaderboard today come from run 2026-q3-pilot-01, a pilot executed on 2026-08-12 through each model's API as served by fal.ai. It is a real controlled run, not a public-gallery sample, and it is not the full protocol described above. Five deviations, stated so nobody has to infer them:
- 3 of 5 prompt levels run (L5/L3/L1). L4 and L2 are omitted, so the onset level cannot be located between L4 and L2 and the leakage curve has three points, not five.
- 8 replicates per cell instead of 32. Cell rates therefore land on eighths (0, 0.125 ... 1.0) and confidence intervals are wide.
- Fixed seed block applies only to the three endpoints that accept a seed (nano-banana-pro, ideogram-v4, flux-2-klein). openai/gpt-image-2 and seedream-v5-pro expose no seed parameter, so their replicates are independent random draws and are not bit-reproducible.
- Corpus is 5 properties from 5 owners, not ~118 from ~20. Breadth balancing is trivially satisfied (one property per owner) and coverage is far below the 90% publication bar.
- Composite from this run is NOT comparable to a full-protocol composite: the weights are renormalized over three levels.
Against the publication bar: 5 properties from 5 owners, against the methodology's 90 percent of roughly 118 properties. Wider than the plus or minus 3.0 the methodology requires. Not run. One pass, no independent re-test. This run does not meet the publication bar and is published as a pilot.
The frozen manifest for this run, with every prompt, seed, model slug, threshold and count, is published at runs/2026-q3-pilot-01/manifest.json, and the per-frame records at frames.json. Engine precision and recall against human labels are not measured yet; the protocol requires them before any ranked number is published, and that work is outstanding.