Smaller, Weaker, yet Better: Training LLM Reasoners via Compute-Optimal Sampling by from on 2024-09-03 05:26 (#6QE9J) Comments