Open-Weights Reasoning Models Achieve 94% on MATH Benchmark at 1/10th Latency
Researchers release a 14B parameter reasoning architecture leveraging sparse chain-of-thought distillation, rivaling proprietary 70B models while running locally on consumer hardware.
- 14B parameter model runs at 85 tokens/sec on RTX 4090
- Implements dynamic depth token pruning during reasoning loops
- Weights and training datasets published under Apache 2.0 license