ProvenanceGuard detects cross-source conflation in MCP agents

ProvenanceGuard is a source-aware verification layer designed for agents using the Model Context Protocol (MCP). The paper presents it as a post-generation gate that preserves source identities from the agents tool trace and evaluates each claim against the specific source the answer names or implies.

How the verifier works

Rather than pooling evidence, ProvenanceGuard reads the captured MCP trace and carries source IDs through five steps: decompose the answer into discrete claims; route each claim to the most relevant tool output; check whether that source supports the claim (using an NLI-style verifier in the experiments); compare the supporting source with the source the answer names or implies; and emit per-claim source verdicts plus a global allow/block decision. Blocked answers may be sent to a RARR-style repair loop and re-verified.

In the experimental setup reported in the paper, the authors used local models for trace processing: MiniLM for routing, a DeBERTa NLI verifier for support checks, and a local language model to decompose answers. The paper notes these named models were part of the evaluated configuration, not a requirement of the approach.

Evaluation and limitations

The researchers tested the method on MCP traces from a medical agent, collecting 281 real traces. For the main held-out set, human experts annotated 361 claims from 40 answers. According to the paper, experts marked 139 claims as undeserving of passage; ProvenanceGuard blocked 138 of those and allowed one. The verifier also held 67 claims that experts considered supported, sending them for review or repair — a behavior the paper attributes to a deliberately conservative decision policy suitable for data-sensitive review.

When a supporting source could be identified, the system selected the correct source about 86% of the time in this test. On a harder multi-source stress test the system scored 0.846 F1 for the block decision but identified the exact source correctly in 50.3% of claims. In a controlled attribution swap test, ProvenanceGuard detected all 50 cases where the named source had been changed while preserving the supporting evidence.

Repair runs resolved all blocked answers in the reported experiments: a full-trace repair resolved 173 blocked answers though 144 ended in fallback text; a reconstructed multi-source test resolved 59 initially blocked answers with two terminal fallbacks. The offline gate added roughly 0.5 seconds per answer in the local configuration, with NLI and routing calls in the tens of milliseconds.

The paper also notes early adoption of the approach in NVIDIA NVFlows optional grounding-verification stage for a finance agent. The ProvenanceGuard method and detailed results are available in the authors paper (also on arXiv); the work was presented as a poster at Agentic AI Summit 2026.


Original source: Hugging Face Blog

Leave a Comment