Skip to content
Bluzen Labs

The evaluation kit

What we use to decide whether a change was real: a preference-trained scorer, human preference pairs, a stratified benchmark, and the ranker for eyeballing the result.

All collectionsUpdated 11 August 2026

Models(1)

bluzen/aesthetic-scorer-v2

Image Classification

Preference-trained aesthetic scorer calibrated on paired human judgements rather than scraped star ratings.

88M paramsPyTorchApache-2.030 Jul 202620k190

Datasets(2)

bluzen/tagbench-2026

Image Classification

Held-out tagging benchmark with per-tag frequency strata and a published error analysis for every baseline.

48k rowsParquetCC-BY-4.011 Aug 202623k208

bluzen/pref-pairs-v1

Text Classification

Paired human preference judgements over generated illustrations, with annotator agreement kept in the file.

210k pairsParquetCC-BY-4.027 Jul 202614k141

Spaces(1)

aesthetic-ranker

bluzen/aesthetic-ranker

Image Classification

Upload a folder, get it ranked. Mostly built to sanity-check the scorer against your own taste.

CPUStreamlitApache-2.031 Jul 2026205