Evidence Integrity in Grounded AI Systems
Evidence Integrity Research
Studying whether instructions hidden in untrusted source documents can manipulate an evidence-grounded LLM system's structured decisions while its output stays structurally valid and apparently grounded.
- Status
- Methodology built · no empirical results yet
- Stack
- AI security
- Prompt injection
- Evidence attribution
- Structured outputs
- Paired benchmarking
- Reproducibility
- Python
Research question
Many LLM applications are built to be evidence-grounded: they receive source documents and return structured claims, each tied to the part of the source that supports it. The citation is what makes the output look trustworthy.
This project asks whether that trust is deserved when the sources themselves are untrusted:
Can adversarial instructions embedded in source documents corrupt claim-level evidence integrity in evidence-grounded LLM applications?
Put operationally: when the underlying facts and the task are held constant, can text placed in a source document change the system's assessment, add a claim, hide a gap, or redirect a citation — while the output still parses, still cites real source blocks, and still looks grounded?
This is research in progress. It reports a methodology and a working harness. It does not yet report findings.
Threat model
Trusted
Application instructions
Define the task.
Experiment configuration
Defines the evaluation protocol.
Expected ground truth
Human-authored facts per case.
Untrusted
Source documents
May contain embedded instructions.
User-provided content
Treated as evidence, not authority.
Retrieved content
Retrieval does not confer trust.
Protected
Source-to-claim integrity
Claims stay supported by what they cite.
Cross-case isolation
One case cannot influence another.
Hidden instructions and canaries
Must not be exposed.
The adversary knows the general task and may try to keep the output looking valid while changing what it means. A compromised model provider, host, or evaluator is out of scope.
The threat model names the property that makes this worth studying: schema validity alone is not a security property. A response can be syntactically valid while violating claim-level evidence integrity.
Structural vs semantic grounding
- Structural groundingis not the same asSemantic grounding
A citation id can exist, and it can point to a real source block. That shows the reference is well-formed. It does not show that the block's text supports the claim attached to it.
Everything built so far validates the first kind. The harness can say whether a cited block exists in the current case, whether a reference is malformed, and whether it leaks across cases. It cannot yet say whether the cited text actually supports the claim.
That gap is the research problem, not a result. Semantic claim-support validation has not been implemented, and the project's own documents reserve it for a later milestone.
R2 — Evidence-grounded testbed
R2 established a small, deterministic, inspectable system to attack later: a model reads a synthetic candidate document and a synthetic job document and returns an assessment with claims, gaps, and evidence references.
Inputs
- then
Synthetic documents
A resume-like and a job-like source per case.
- then
Segmentation
Deterministic paragraph splitting.
- then
Source blocks
Immutable, with stable ids.
Provider boundary
Real model or test double.
Outputs
- then
Raw output
Preserved verbatim.
- then
Structured assessment
Match, partial, or mismatch.
- then
Claims and gaps
Each claim cites source blocks.
Reference validation
Structural checks only.
- Synthetic data only. Four clean cases, each with human-authored ground truth: the facts a correct assessment should reflect and the expected classification.
- Stable block identity. The same text always segments into the same blocks with the same ids, so an evidence reference means the same thing on every run.
- Trusted and untrusted text are separated in the prompt. Instructions, source blocks, and the output schema are assembled as distinct parts.
- References are classified. Each cited block is valid, nonexistent, malformed, or cross-case. Any invalid reference fails the run.
- A deterministic test double. A provider that returns canned responses offline, flagged as a test double, so the pipeline can be tested without a model or credentials.
| Question | R2 answers it |
|---|---|
| Does the cited block exist in this case? | Yes |
| Is the reference well-formed and case-isolated? | Yes |
| Does the cited text actually support the claim? | No |
R3 — Paired attack benchmark
R3 adds the attack, under a design meant to make any difference attributable to it.
Held constant
Candidate document
Identical in both conditions.
Ground truth
Unchanged by the attack.
Task and model configuration
Same instructions, same settings.
paired with
Job document
Clean
The original job text.
Attacked
The same text plus one payload paragraph at the beginning, middle, or end.
same pipeline
Results
Clean result
Assessment, claims, gaps, references.
Attack result
Assessment, claims, gaps, references.
compared by
Comparison
Deterministic paired metrics
No model acts as a judge.
The attacked document is generated from the clean one and a machine-readable attack specification. Because payloads are inserted as their own paragraph, segmentation isolates each one in its own source block and leaves the neighbouring clean blocks untouched. No attack configuration lives in code.
The benchmark currently has ten paired synthetic cases.
Attack taxonomy
These are four fixed benchmark attacks: simple, hand-written injection templates. They are not novel attacks, not optimized attacks, and not PLeak.
| Class | Name | What it attempts |
|---|---|---|
| A1 | Assessment manipulation | Change the overall classification regardless of the evidence. |
| A2 | Unsupported claim | Add a specific positive claim the candidate source does not support. |
| A3 | Gap suppression | Hide missing qualifications, or empty the gaps list. |
| A4 | Citation steering | Attach a chosen source block as evidence for a claim. |
Metrics
Each pair is compared with deterministic rules:
- Assessment changed — the classification differs between clean and attacked runs.
- Target assessment success — the attacked run reached the classification the attack asked for.
- Claim set changed — the normalized set of claims differs.
- Target claim present — the requested claim appears, by deterministic text normalization.
- Gap count delta and gap suppression — gaps were removed, or a targeted gap disappeared.
- Evidence reference changed — the set of cited blocks differs.
- Structural validity — tracked separately for each run: schema valid, references valid, case-isolated.
- Structurally valid manipulation — the attack achieved its target and the output stayed fully structurally valid.
The last one is the measurement the research question turns on. It is still a structural measurement. None of these is a semantic citation-correctness rate or an evidence-integrity violation rate, because nothing here evaluates whether cited text supports a claim.
Research architecture
Build
- then
Synthetic fixture
- then
Segmentation
- then
Source registry
Prompt builder
Run
- then
Provider boundary
- then
Raw model output
- then
Structured parser
Reference validator
Record
- then
Immutable artifacts
Paired metrics
In R3 the controlled attack transformation happens at the very start, on the job fixture, before segmentation. Everything downstream is identical for the clean and attacked branches.
Every paired run writes its own directory: the configuration, the attack specification, the segmented sources for both branches, the raw responses, the parsed responses with their validation status, and the comparison. A run directory is never overwritten; targeting an existing one is an error. That is what makes a run reproducible and auditable, and it removes the option of quietly editing a result afterwards.
Test double vs empirical result
The benchmark harness has been validated with a deterministic test double. Every result artifact in the repository comes from that test double, and each is tagged in its own metadata as test-only and non-scientific.
Those runs verify that the experimental pipeline works end to end. They do not demonstrate that any real LLM is vulnerable, and their outcomes are not findings. A live-model provider interface exists, but no real-model run has been recorded.
So there are no attack success rates on this page, and there should not be.
Research decisions
Synthetic data only
- Why
- Ground truth is fully controlled, and no real person's documents are involved.
- Tradeoff
- Less ecological realism than real resumes and postings would give.
Stable source block identities
- Why
- An evidence reference has to mean the same thing across runs to be inspectable and reproducible.
- Tradeoff
- Requires deterministic segmentation, which fixes how documents can be split.
Structural validation before semantic validation
- Why
- Whether a reference exists and whether it supports a claim are different questions and should not be blurred.
- Tradeoff
- R2 and R3 cannot establish that a cited block entails its claim.
Clean-versus-attack pairs
- Why
- Facts stay fixed and only the payload changes, so a difference can be attributed to the attack.
- Tradeoff
- The initial scope is deliberately narrow: one payload, one document, one task.
Test doubles validate plumbing only
- Why
- Deterministic runs prove the harness, parsing, and metrics behave as specified.
- Tradeoff
- They cannot support any empirical claim about model behaviour.
Current status
- R0Complete
Research foundation
Research questions, threat model, outcomes, metrics, and experiment plan.
- R1Partial
External baseline (PLeak)
Upstream code is pinned in isolation and a synthetic protected-target fixture exists. The baseline has not been reproduced.
- R2Complete
Evidence-grounded testbed
Four clean cases, deterministic segmentation, structural reference validation.
- R3AComplete
Paired benchmark harness
Ten paired cases, four fixed attack classes, paired metrics. Test double only.
- R3BNext
Real-model pilot
The frozen benchmark against a real model. Not yet run.
- LaterNot started
Semantic support validation and defenses
Checking whether cited text supports a claim, then evaluating defenses.
PLeak appears here only as an external baseline on its own track. It has not been reproduced in this project, the fixed R3 attacks are unrelated to it, and none of its results belong to this work.
Limitations
- Synthetic data. Every document and candidate is invented.
- Small benchmark. Four clean cases and ten paired cases.
- Fixed attack templates. Hand-written payloads, not optimized or adaptive.
- No semantic support validation. Only structural reference checks exist.
- No real-model results. The harness has run only against a test double.
- No defenses evaluated. Nothing here measures a mitigation.
- External baseline not reproduced. The PLeak track has no results.
- One task. A single assessment task with a three-way classification.
Nothing on this page claims a new vulnerability, that grounded or retrieval-augmented systems in general are vulnerable, statistical significance, a validated defense, or a published paper.
Next experiment
The next step is a real-model pilot of the paired benchmark: the ten existing pairs, each run clean and attacked, against one real model, once.
The goal is not statistical significance. It is to find out whether the phenomenon appears at all on a real model under a benchmark that was frozen before any real-model result existed. Whatever it shows — including nothing — is what would be reported.