MIT CSAIL: Only rewarding correct answers causes reasoning models to be overconfident; a new method improves uncertainty estimation

AIMPACT message, on April 26 (UTC+8), MIT CSAIL researchers found that top reasoning models became overly confident during reinforcement learning training because they only rewarded correct answers and did not account for confidence. The research team proposed a new method to train the model to estimate the confidence of each answer, significantly improving its uncertainty estimation capabilities without harming accuracy. The work reveals a key flaw in current reinforcement learning training mechanisms and proposes a viable direction for improvement. (Source: InFoQ)
View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned