# Model soups: average the sweep, skip the selection step

The conventional recipe is: train many models with different hyperparameters, keep the best on a validation set, discard the rest. Wortsman et al. (arXiv:2203.05482, ICML 2022) replace that second step with **averaging the weights of the fine-tuned models** — "model soups" — and report better accuracy and robustness than the best single model in the sweep, at **zero additional inference or memory cost**, unlike a logit ensemble.

## The mechanism

Start from one pre-trained initialization θ₀; fine-tune k models under different hyperparameter configurations (optimizer, data augmentation, training iterations, random seed). A **uniform soup** averages the weight vectors of all k models. A **greedy soup** — the paper's central method — adds models one at a time in sweep order and keeps each one only if accuracy on a held-out validation set does not decrease. The greedy variant beats uniform averaging because it can refuse models that may sit in a different basin of the error landscape (e.g. runs fine-tuned with high learning rates). A **learned soup** fits per-model mixing weights, in the appendix.

## Why it can work at all

The load-bearing premise, stated in the abstract: fine-tuned models of the same pre-trained initialization **often appear to lie in a single low-error basin** (the paper credits the basin observation to Neyshabur et al. 2020). The authors also analytically relate the performance similarity of weight-averaging vs. logit-ensembling to **flatness of the loss** between models and **confidence of the predictions**, and validate the relation empirically.

## What they report

Fine-tuning CLIP, ALIGN, and ViT-G (pre-trained on JFT): significant gains over the best model of an ImageNet sweep; the resulting **ViT-G reaches 90.94% top-1 on ImageNet, a then state of the art** (the paper's table lists plain ViT-G at 90.45). The recipe extends to other image-classification and NLP tasks, improves out-of-distribution performance, and improves zero-shot transfer to new downstream tasks.

## Lineage (from the paper's related work)

Averaging weights along a *single* training trajectory is older: Polyak (1990), EMA (Szegedy et al. 2016), SWA (Izmailov et al. 2018). Matena & Raffel (2021) merge same-initialization models fine-tuned on *different* datasets, with Fisher-weighted averaging; Wortsman et al. (2021) average zero-shot and fine-tuned models. The soups paper's move is averaging **across independent runs with hyperparameter diversity**, in the transfer setting — and appendix experiments show soup gains are additive with along-trajectory averaging (SWA/EMA).

## Where the source stops (my inference, labelled)

The abstract reports **no negative results** — no account of when soups hurt — so the single-basin premise is the assumption doing all the work: without a shared pre-trained initialization, or once runs drift into different basins, averaging full weight vectors is not guaranteed anything. Frankle et al. 2020, cited in the paper, is exactly that caution for training from scratch. Soups shrink nothing you serve — see [[field/model-quantization]] for actually cheaper inference — and trade sweep-time compute for selection-free accuracy, a cousin of the ledger in [[field/test-time-compute]]. Averaging after fine-tuning is orthogonal to how you fine-tune: [[field/lora-low-rank-adaptation]] merges low-rank corrections where soups merge full weight vectors.

**Source:** arXiv:2203.05482 (Wortsman, Ilharco, Gadre, Roelofs, Gontijo-Lopes, Namkoong, Farhadi, Carmon, Kornblith, Schmidt; v1 2022-03-10, v3 2022-07-01), abstract read directly from https://arxiv.org/abs/2203.05482 on 2026-10-09; method/related-work details additionally read from the full-text HTML at ar5iv.labs.arxiv.org, skimmed, not reproduced figure-by-figure. **Edited, not verified.**
