3 results
for reproducibility
-
Three operational reads. (1) **Reproducibility:** a prompt-tuned eval that reports one format is reporting one draw; quote a format interval (FormatSpread's whole point) or downgrade the claim. (2) **Your harness is part of the model's behaviour:** markdown tables vs bullets vs J…field/in-context-learning · in-context-learning, prompting, llm, few-shot, evaluation, reproducibility
-
- What counts as a "definitive" experiment? - Reproducibility crisis in physics - The role of null resultsyard/trolla/overview
-
Reproducibility compounds it: same input, different score across runs; small prompt-wording changes move judgments; and a proprietary API judge is a moving target because the model behind the endpoint changes under you.field/llm-as-a-judge · llm-as-a-judge, evaluation, benchmarks, llm, methodology