Hello authors,
Thank you for the nice work. I have a few questions regarding some parts of the paper and the experimental setup.
1) Baseline comparison in Table 1
From Table 1, it seems that some baseline results, such as d1, ESPO, SPG, and d TreeRPO, may have been directly taken from the original papers rather than reproduced within the same codebase. My understanding is that the missing entries appear because the corresponding experiments are not reported in the original papers.
However, I am concerned that there may be important configuration differences across these results. For example, the original d1 codebase typically uses quantization plus LoRA training, whereas the codebase used in this paper appears to be based on full finetuning and may also use different reward implementations. In addition, different papers may use different training settings, such as the number of optimization steps, batch sizes, or other hyperparameters.
If the values in Table 1 are directly taken from the original papers, these configuration differences may affect the fairness of the comparison. Could the authors clarify how this issue is handled? Were all baselines reproduced under the same codebase and training configuration, or were some numbers taken directly from prior work? Please let me know if I misunderstood this point.
2) Temperature setting during training
In Figure 8 of Appendix B.1, the authors report the reasoning potential under different temperature settings. I noticed that the arbitrary order strategy, which corresponds to confidence based decoding in LLaDA, outperforms AR by a large margin at temperature 1.0. However, the default temperature in GRPO training appears to be set to 1.0, which does not seem to be the most favorable setting for AR.
Could the authors provide the exact temperature configuration used during training? Also, if temperature 1.0 is indeed the default choice, could the authors explain the reason behind this choice? In particular, why not use sharper temperatures, such as 0.6 or 0.3, where AR seems to outperform the confidence based strategy more clearly?
Thank you for reading my questions. I look forward to your clarification.
Hello authors,
Thank you for the nice work. I have a few questions regarding some parts of the paper and the experimental setup.
1) Baseline comparison in Table 1
From Table 1, it seems that some baseline results, such as d1, ESPO, SPG, and d TreeRPO, may have been directly taken from the original papers rather than reproduced within the same codebase. My understanding is that the missing entries appear because the corresponding experiments are not reported in the original papers.
However, I am concerned that there may be important configuration differences across these results. For example, the original d1 codebase typically uses quantization plus LoRA training, whereas the codebase used in this paper appears to be based on full finetuning and may also use different reward implementations. In addition, different papers may use different training settings, such as the number of optimization steps, batch sizes, or other hyperparameters.
If the values in Table 1 are directly taken from the original papers, these configuration differences may affect the fairness of the comparison. Could the authors clarify how this issue is handled? Were all baselines reproduced under the same codebase and training configuration, or were some numbers taken directly from prior work? Please let me know if I misunderstood this point.
2) Temperature setting during training
In Figure 8 of Appendix B.1, the authors report the reasoning potential under different temperature settings. I noticed that the arbitrary order strategy, which corresponds to confidence based decoding in LLaDA, outperforms AR by a large margin at temperature 1.0. However, the default temperature in GRPO training appears to be set to 1.0, which does not seem to be the most favorable setting for AR.
Could the authors provide the exact temperature configuration used during training? Also, if temperature 1.0 is indeed the default choice, could the authors explain the reason behind this choice? In particular, why not use sharper temperatures, such as 0.6 or 0.3, where AR seems to outperform the confidence based strategy more clearly?
Thank you for reading my questions. I look forward to your clarification.