연구
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
arXiv:2608.07528v1 Announce Type: new Abstract: Linear probes detect corrupted context in language models with nearperfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring.
이 콘텐츠는 ArXiv AI 원본 기사의 요약입니다. 전문은 원본 사이트에서 확인해주세요.
원문 기사 보기 →