Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #5847

MMLU

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

arXiv paper: generate-transform decomposition explains when small-LLM team scaling helps across orchestration architectures

This arXiv preprint (arXiv:2609.36104v1) evaluates eight agent orchestration architectures across five instruction-tuned 7–9B models and several benchmarks (GSM8K, GSMHard, ARC, GPQA, MMLU, and an executable-code task) up to 30 calls. The authors introduce an exact generate–transform decomposition that splits accuracy change into coverage and transformation effects, finding that returns to adding agents are sharply task- and architecture-dependent: Proposer-Critic scales steeply and yields large gains on arithmetic problems (up to +17 points) but not on multiple-choice benchmarks, and token cost per budget still varies ~2.1×.

7.0

RESEARCH · 1 SOURCE · Hugging Face

Multiverse maps block-removal LLM pruning to an Ising optimization and reports large MMLU gains

Multiverse's new paper, “LLM Compression by Block Removal with Constrained Binary Optimization,” formulates block selection for depth pruning as a constrained binary optimization problem that maps onto an Ising glass; it computes a Hessian of block couplings once on a small calibration set so candidate prunings can be scored cheaply by energy. The authors report that this search over interacting block combinations yields large gains in the deep-compression regime, e.g., nearly 23 percentage points on MMLU at 50% compression of Llama-3.3-70B-Instruct versus the best competing block-removal method.

8.0