Tech Meridian ← ENTITY INDEX
PROMY MERIDIAN RU

COMPANY · ENTITY #7644

PubMedQA

Related event timeline, sources and context from the news index.

EVENT TIMELINE

2

RESEARCH · 1 SOURCE · arXiv cs.AI

HARDEN: constrained evolutionary search to create harder, answer-preserving evaluation cases

HARDEN is a constrained evolutionary search method that adapts inputs of existing evaluation cases into more challenging variants while preserving expected outputs and enforcing feasibility constraints (semantics, realism, execution validity). Applied to FinQA, PubMedQA, and ContractNLI on three Qwen3.5 scales, HARDEN reduced model accuracy by 22.7% on average and up to 49.9% versus single-pass baselines using the same feasibility checks.

6.0

RESEARCH · 1 SOURCE · arXiv cs.AI

Large Knowledge Model (LKM) paper maps literature into reasoning graphs to form a Scientific Reasoning Landscape

The arXiv paper introduces the Large Knowledge Model (LKM), a corpus-scale scientific knowledge infrastructure that represents papers as source-grounded reasoning graphs and aligns questions, claims, and reasoning chains across works to form a three-part Scientific Reasoning Landscape (Question, Workflow, Evidence). The authors describe system design and evaluations showing that LKM retrieval, with a fixed answering model, improves accuracy by 9.30%, 4.20%, and 14.69% on ChemBench, PubMedQA, and SciBench respectively, and enables reasoning-aware search, evidence-grounded QA, comparative evidence analysis, and research planning.

7.0