Bỏ qua đến nội dung
filmtechXAI
← Quay lại
CẬP NHẬT

Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

AWS Machine Learning Blog

In this post, we explore an idea for generating thinking tokens for datasets that lack reasoning traces in SFT customization. We first examine the reasoning suppression problem, then introduce Self-Distilled Reasoning (SDR), validate it across three benchmarks, and provide practical recommendations.


Đọc đầy đủ tại: AWS Machine Learning Blog