character-index
All datasets
Image Feature Extraction96k entitiesParquet5.4k77Updated 29 May 2026
character-index
Character identity index with source attribution, used for retrieval evaluation and near-duplicate detection.
This card is generated from the entry's metadata in src/data/hub.js. Write a
card field on the entry to replace it with a real one.
How to use it
from datasets import load_dataset
ds = load_dataset("bluzen/character-index", split="train")
Details
| Task | Image Feature Extraction |
| Size | 96k entities |
| Library | Parquet |
| Licence | CC-BY-4.0 |
Limitations
Every dataset the lab publishes carries the biases of the corpus it came from, and ours is illustration-heavy. Measure before you trust it.