aurora-curated
All datasets
Image Classification1.2M rowsParquet41k366Updated 26 August 2026
aurora-curated
A 1.2M-row quality-filtered slice of Aurora. The subset most of our own training runs actually use.
This card is generated from the entry's metadata in src/data/hub.js. Write a
card field on the entry to replace it with a real one.
How to use it
from datasets import load_dataset
ds = load_dataset("bluzen/aurora-curated", split="train")
Details
| Task | Image Classification |
| Size | 1.2M rows |
| Library | Parquet |
| Licence | CC-BY-4.0 |
Limitations
Every dataset the lab publishes carries the biases of the corpus it came from, and ours is illustration-heavy. Measure before you trust it.