Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #10937

MATH-500

Related event timeline, sources and context from the news index.

EVENT TIMELINE

3

RESEARCH · 1 SOURCE · arXiv cs.AI

Mixture-of-Agents paper measures per-token 'sufficient compute' and shows routing/drafting gains

arXiv:2610.02491v1 introduces a Mixture-of-Agents (MoA) methodology that measures each token's upper-bound 'sufficient compute' by testing which of 15 progressively larger models can reproduce a reference token; a 0.5B agent reproduces 92–95% of tokens, while the costliest 10% of tokens account for an estimated 64–80% of FLOPs. Using the MoA-derived map, model routing and drafting strategies reduce projected latency (e.g., from 7.59s to 5.12s on MATH-500) and cut draft-token use and projected latency versus fixed baselines, indicating remaining headroom for smarter allocation controllers.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

GoldiMask: context selection and target weighting for fine-tuning diffusion language models

arXiv:2609.38385v1 introduces GoldiMask, a supervised fine-tuning procedure for discrete diffusion language models that selects which tokens to reveal as context via an approximate submodular objective and weights remaining prediction targets by their benefit and learnability. Across three backbones and three datasets the paper reports higher average accuracy in most settings (including reasoning and code generation) and reduced decoding iterations on GSM8K and MATH-500 under confidence-threshold parallel decoding; ablations show both context selection and target weighting contribute to the gains.

7.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Separating coverage from specialization in LLM harnesses (arXiv:2609.35873v1)

The paper introduces a controlled evaluation that disentangles answer coverage, repeatable task advantages, and pre-execution selection for generated LLM harnesses. On 386 MATH-500 tasks, comparisons of eight generated harnesses against a baseline with nine identical copies show that repeated identical programs produce 2.16 percentage points of repeat-averaged oracle headroom; generated programs display repeatable score patterns that mainly reveal persistent weaknesses rather than useful persistent wins, the frozen selector yields 0.00 percentage-point gain, and both populations reach 98.70% oracle coverage at 27 harness executions; BIRD traces locate failures in mechanism implementation, activation, and output validity, leading the authors to argue that coverage and repeatability alone do not justify claims of useful specialization and to propose stricter evaluation standards for harness diversity.

5.0