Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #5702

GSM8K

Related event timeline, sources and context from the news index.

EVENT TIMELINE

4

RESEARCH · 1 SOURCE · arXiv cs.AI

GoldiMask: context selection and target weighting for fine-tuning diffusion language models

arXiv:2609.38385v1 introduces GoldiMask, a supervised fine-tuning procedure for discrete diffusion language models that selects which tokens to reveal as context via an approximate submodular objective and weights remaining prediction targets by their benefit and learnability. Across three backbones and three datasets the paper reports higher average accuracy in most settings (including reasoning and code generation) and reduced decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding; ablations show both context selection and target weighting contribute to the gains.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Neurosymbolic router learns DFA to route edge queries, improving small-model accuracy and efficiency

arXiv:2609.35833v1 proposes a neurosymbolic router that classifies incoming queries and dispatches them to the cheapest correct solver, learning a deterministic finite automaton (DFA) with the L* algorithm using a small language model as a membership oracle. On a Raspberry Pi 4B evaluated on 100 unseen prompts from DeepMind Mathematics, GSM8K, and RuleTaker, the learned router achieves 100% routing accuracy and 98.3% overall accuracy with a 512-token reasoning budget (93.3% on word problems), outperforming Program-of-Thought (72.0%) and a tool-calling agent (58.7%), while answering formatted queries in 1–11 ms and, in a 30-token configuration, running 8.8× faster and 2.8× more energy-efficient than Program-of-Thought.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

arXiv paper: generate-transform decomposition explains when small-LLM team scaling helps across orchestration architectures

This arXiv preprint (arXiv:2609.36104v1) evaluates eight agent orchestration architectures across five instruction-tuned 7–9B models and several benchmarks (GSM8K, GSMHard, ARC, GPQA, MMLU, and an executable-code task) up to 30 calls. The authors introduce an exact generate–transform decomposition that splits accuracy change into coverage and transformation effects, finding that returns to adding agents are sharply task- and architecture-dependent: Proposer-Critic scales steeply and yields large gains on arithmetic problems (up to +17 points) but not on multiple-choice benchmarks, and token cost per budget still varies ~2.1×.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Attention-Aware Routing (AAR) for MoE models (arXiv:2609.20974v1)

The paper introduces Attention-Aware Routing (AAR), which augments Mixture-of-Experts routers with temporal and spectral features derived from a sliding window of attention weights while keeping the base transformer frozen and training only routing parameters. On OLMoE, AAR improves GSM8K accuracy by +3.37 percentage points over a routing-only SFT baseline, demonstrates that routing updates propagate to reshape subsequent-layer attention without changing attention weights directly, reduces long diverging generations for incorrect answers, and shows strong depth sensitivity that separates retrieval and reasoning behavior across layers.

7.0