Skip to content

Clarification on Baseline Configuration Differences and Training Temperature #10

Description

@weiwang1342-wq

Hello authors,

Thank you for the nice work. I have a few questions regarding some parts of the paper and the experimental setup.

1) Baseline comparison in Table 1

From Table 1, it seems that some baseline results, such as d1, ESPO, SPG, and d TreeRPO, may have been directly taken from the original papers rather than reproduced within the same codebase. My understanding is that the missing entries appear because the corresponding experiments are not reported in the original papers.

However, I am concerned that there may be important configuration differences across these results. For example, the original d1 codebase typically uses quantization plus LoRA training, whereas the codebase used in this paper appears to be based on full finetuning and may also use different reward implementations. In addition, different papers may use different training settings, such as the number of optimization steps, batch sizes, or other hyperparameters.

If the values in Table 1 are directly taken from the original papers, these configuration differences may affect the fairness of the comparison. Could the authors clarify how this issue is handled? Were all baselines reproduced under the same codebase and training configuration, or were some numbers taken directly from prior work? Please let me know if I misunderstood this point.

2) Temperature setting during training

In Figure 8 of Appendix B.1, the authors report the reasoning potential under different temperature settings. I noticed that the arbitrary order strategy, which corresponds to confidence based decoding in LLaDA, outperforms AR by a large margin at temperature 1.0. However, the default temperature in GRPO training appears to be set to 1.0, which does not seem to be the most favorable setting for AR.

Could the authors provide the exact temperature configuration used during training? Also, if temperature 1.0 is indeed the default choice, could the authors explain the reason behind this choice? In particular, why not use sharper temperatures, such as 0.6 or 0.3, where AR seems to outperform the confidence based strategy more clearly?

Thank you for reading my questions. I look forward to your clarification.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions