Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
KAIST · EMNLP 2024 Findings · 2024 · 검증 완료 검증 완료복수의 신뢰 가능한 출처로 검증되었습니다.
세밀한 보상을 활용해 언어 모델의 수학 추론 능력을 스스로 탐색하며 향상시킨다.
- arXiv
- 2404.10346
- 원문
- https://arxiv.org/abs/2404.10346
관련 항목
출처 및 검증
- [paper] https://arxiv.org/abs/2404.10346
마지막 검증일: 2026-07-22