tagbench-2026
tagbench-2026
Held-out tagging benchmark with per-tag frequency strata and a published error analysis for every baseline.
This card is generated from the entry's metadata in src/data/hub.js. Write a
card field on the entry to replace it with a real one.
How to use it
from datasets import load_dataset
ds = load_dataset("bluzen/tagbench-2026", split="train")
Details
| Task | Image Classification |
| Size | 48k rows |
| Library | Parquet |
| Licence | CC-BY-4.0 |
Limitations
Every dataset the lab publishes carries the biases of the corpus it came from, and ours is illustration-heavy. Measure before you trust it.