A publication dated September 17, 2026 describes TinyLoRA and reports gains on three benchmarks using RL training with 13 parameters.
A publication dated September 17, 2026 presents TinyLoRA, a method that replaces LoRA’s low-rank matrix with a trainable vector projected by a random tensor. The proposal explores whether reinforcement-learning-based fine-tuning can reduce adapter parameter counts and memory use.
According to the text, training with GRPO and 13 parameters improved Qwen2.5-7B-Instruct scores on the GSM8K, MATH500, and AIME24 benchmarks. To check the results and their details, consult the original publication and verify its setup, metrics, and reported comparisons; the available summary gives no numerical values.