연구
There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items
arXiv:2608.21382v1 Announce Type: new Abstract: Multiplechoice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from peroption likelihoods.
이 콘텐츠는 ArXiv AI 원본 기사의 요약입니다. 전문은 원본 사이트에서 확인해주세요.
원문 기사 보기 →