Skip to content
Bluzen Labs
character-index
All datasets

bluzen/character-index

Open
Image Feature Extraction96k entitiesParquet5.4k77Updated 29 May 2026

character-index

Character identity index with source attribution, used for retrieval evaluation and near-duplicate detection.

This card is generated from the entry's metadata in src/data/hub.js. Write a card field on the entry to replace it with a real one.

How to use it

from datasets import load_dataset

ds = load_dataset("bluzen/character-index", split="train")

Details

TaskImage Feature Extraction
Size96k entities
LibraryParquet
LicenceCC-BY-4.0

Limitations

Every dataset the lab publishes carries the biases of the corpus it came from, and ours is illustration-heavy. Measure before you trust it.