Tech Meridian ← LIVE FEED
RU

RESEARCH · RESEARCH · #530

Masked-concept SNOMED CT recommendation benchmark using SNOMED CT Entity Linking Challenge v1.2.1

This arXiv preprint introduces a masked‑concept recommendation benchmark built from the SNOMED CT Entity Linking Challenge v1.2.1 dataset derived from MIMIC‑IV‑Note (75,491 annotations across 272 discharge summaries; 204 notes training, 68 test). The paper compares popularity, sparse TF‑IDF prototypes, dense LSA embeddings, sparse‑dense fusion, retrieved‑note evidence, and a retrieval‑augmented hybrid, finding sparse TF‑IDF best (Recall@1 14.81%, Recall@10 33.43%, MRR 0.2114, nDCG@10 0.2297) and showing strong performance degradation for low‑frequency or unseen concepts.

KEY POINTS

  1. This arXiv preprint introduces a masked‑concept recommendation benchmark built from the SNOMED CT Entity Linking Challenge v1.2.1 dataset derived from MIMIC‑IV‑Note (75,491 annotations across 272 discharge summaries; 204 notes training, 68 test).
  2. The paper compares popularity, sparse TF‑IDF prototypes, dense LSA embeddings, sparse‑dense fusion, retrieved‑note evidence, and a retrieval‑augmented hybrid, finding sparse TF‑IDF best (Recall@1 14.81%, Recall@10 33.43%, MRR 0.2114, nDCG@10 0.2297) and showing strong performance degradation for low‑frequency or unseen concepts.
  3. The result provides a reproducible baseline and shows that local lexical context and training-set terminology coverage, not retrieval augmentation, chiefly limit SNOMED CT recommendation performance in low‑resource clinical settings.

WHY IT MATTERS

The result provides a reproducible baseline and shows that local lexical context and training-set terminology coverage, not retrieval augmentation, chiefly limit SNOMED CT recommendation performance in low‑resource clinical settings.

SOURCES & TIMELINE

1