연구
How much of a measured AI preference is the model, and how much is the instrument?
arXiv:2608.23641v1 Announce Type: new Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. 2024, Mazeika et al. 2025, Mikaelson et al. 2025, Tagliabue and Dung 2025 and Trhlik et al.
이 콘텐츠는 ArXiv AI 원본 기사의 요약입니다. 전문은 원본 사이트에서 확인해주세요.
원문 기사 보기 →