RESEARCH · RESEARCH · #1419
Praxa: an evidence-bound harness for governed LLM agent execution (arXiv:2610.00015v1)
The paper presents Praxa, an agent harness that makes proposal, authority, execution, external verification, reconciliation, and promotion explicit and testable through deterministic admission, brokered execution, external read-back, and reviewed promotion. The authors report four evidence lanes—an author-run repository audit, a Terminal-Bench Core 0.1.1 pilot, a post-debug coordination-proxy comparison, and deployed source/configuration evidence—but note limitations including unavailable raw transcripts, mixed pilot results that do not show superiority, and no demonstrated production outcome or adversarial security guarantees.
KEY POINTS
- The paper presents Praxa, an agent harness that makes proposal, authority, execution, external verification, reconciliation, and promotion explicit and testable through deterministic admission, brokered execution, external read-back, and reviewed promotion.
- The authors report four evidence lanes—an author-run repository audit, a Terminal-Bench Core 0.1.1 pilot, a post-debug coordination-proxy comparison, and deployed source/configuration evidence—but note limitations including unavailable raw transcripts, mixed pilot results that do not show superiority, and no demonstrated production outcome or adversarial security guarantees.
- Praxa matters because it proposes an evidence-bound architecture that makes authority-to-effect transitions explicit and testable, improving auditability even though current results do not demonstrate production benefits or adversarial safety.
WHY IT MATTERS
Praxa matters because it proposes an evidence-bound architecture that makes authority-to-effect transitions explicit and testable, improving auditability even though current results do not demonstrate production benefits or adversarial safety.