A model on its own is half an answer. These group the checkpoint, the data it was trained on, the benchmark it is measured against and the demo that shows it working.
Everything needed to tag an illustration corpus: the tagger, the vocabulary it was trained on, the benchmark it is measured against, and a demo to sanity-check all three.
What we use to decide whether a change was real: a preference-trained scorer, human preference pairs, a stratified benchmark, and the ranker for eyeballing the result.