Skip to content
Xingchi Guo
All research

Evidence Integrity in Grounded AI Systems

Evidence Integrity Research

Studying whether instructions hidden in untrusted source documents can manipulate an evidence-grounded LLM system's structured decisions while its output stays structurally valid and apparently grounded.

Status
Methodology built · no empirical results yet
Stack
  • AI security
  • Prompt injection
  • Evidence attribution
  • Structured outputs
  • Paired benchmarking
  • Reproducibility
  • Python

Research question

Many LLM applications are built to be evidence-grounded: they receive source documents and return structured claims, each tied to the part of the source that supports it. The citation is what makes the output look trustworthy.

This project asks whether that trust is deserved when the sources themselves are untrusted:

Can adversarial instructions embedded in source documents corrupt claim-level evidence integrity in evidence-grounded LLM applications?

Put operationally: when the underlying facts and the task are held constant, can text placed in a source document change the system's assessment, add a claim, hide a gap, or redirect a citation — while the output still parses, still cites real source blocks, and still looks grounded?

This is research in progress. It reports a methodology and a working harness. It does not yet report findings.

Threat model

  1. Trusted

    • Application instructions

      Define the task.

    • Experiment configuration

      Defines the evaluation protocol.

    • Expected ground truth

      Human-authored facts per case.

  2. Untrusted

    • Source documents

      May contain embedded instructions.

    • User-provided content

      Treated as evidence, not authority.

    • Retrieved content

      Retrieval does not confer trust.

  3. Protected

    • Source-to-claim integrity

      Claims stay supported by what they cite.

    • Cross-case isolation

      One case cannot influence another.

    • Hidden instructions and canaries

      Must not be exposed.

The adversary controls text in at least one untrusted source. It does not control the instructions, the configuration, the ground truth, or the evaluator.

The adversary knows the general task and may try to keep the output looking valid while changing what it means. A compromised model provider, host, or evaluator is out of scope.

The threat model names the property that makes this worth studying: schema validity alone is not a security property. A response can be syntactically valid while violating claim-level evidence integrity.

Structural vs semantic grounding

  • Structural groundingis not the same asSemantic grounding

A citation id can exist, and it can point to a real source block. That shows the reference is well-formed. It does not show that the block's text supports the claim attached to it.

Everything built so far validates the first kind. The harness can say whether a cited block exists in the current case, whether a reference is malformed, and whether it leaks across cases. It cannot yet say whether the cited text actually supports the claim.

That gap is the research problem, not a result. Semantic claim-support validation has not been implemented, and the project's own documents reserve it for a later milestone.

R2 — Evidence-grounded testbed

R2 established a small, deterministic, inspectable system to attack later: a model reads a synthetic candidate document and a synthetic job document and returns an assessment with claims, gaps, and evidence references.

Inputs

  1. Synthetic documents

    A resume-like and a job-like source per case.

    then
  2. Segmentation

    Deterministic paragraph splitting.

    then
  3. Source blocks

    Immutable, with stable ids.

    then
  4. Provider boundary

    Real model or test double.

Outputs

  1. Raw output

    Preserved verbatim.

    then
  2. Structured assessment

    Match, partial, or mismatch.

    then
  3. Claims and gaps

    Each claim cites source blocks.

    then
  4. Reference validation

    Structural checks only.

  • Synthetic data only. Four clean cases, each with human-authored ground truth: the facts a correct assessment should reflect and the expected classification.
  • Stable block identity. The same text always segments into the same blocks with the same ids, so an evidence reference means the same thing on every run.
  • Trusted and untrusted text are separated in the prompt. Instructions, source blocks, and the output schema are assembled as distinct parts.
  • References are classified. Each cited block is valid, nonexistent, malformed, or cross-case. Any invalid reference fails the run.
  • A deterministic test double. A provider that returns canned responses offline, flagged as a test double, so the pipeline can be tested without a model or credentials.
QuestionR2 answers it
Does the cited block exist in this case?Yes
Is the reference well-formed and case-isolated?Yes
Does the cited text actually support the claim?No

R3 — Paired attack benchmark

R3 adds the attack, under a design meant to make any difference attributable to it.

  1. Held constant

    • Candidate document

      Identical in both conditions.

    • Ground truth

      Unchanged by the attack.

    • Task and model configuration

      Same instructions, same settings.

    paired with

  2. Job document

    • Clean

      The original job text.

    • Attacked

      The same text plus one payload paragraph at the beginning, middle, or end.

    same pipeline

  3. Results

    • Clean result

      Assessment, claims, gaps, references.

    • Attack result

      Assessment, claims, gaps, references.

    compared by

  4. Comparison

    • Deterministic paired metrics

      No model acts as a judge.

The two branches differ by exactly one documented payload. The clean text is recoverable from the attacked text.

The attacked document is generated from the clean one and a machine-readable attack specification. Because payloads are inserted as their own paragraph, segmentation isolates each one in its own source block and leaves the neighbouring clean blocks untouched. No attack configuration lives in code.

The benchmark currently has ten paired synthetic cases.

Attack taxonomy

These are four fixed benchmark attacks: simple, hand-written injection templates. They are not novel attacks, not optimized attacks, and not PLeak.

ClassNameWhat it attempts
A1Assessment manipulationChange the overall classification regardless of the evidence.
A2Unsupported claimAdd a specific positive claim the candidate source does not support.
A3Gap suppressionHide missing qualifications, or empty the gaps list.
A4Citation steeringAttach a chosen source block as evidence for a claim.

Metrics

Each pair is compared with deterministic rules:

  • Assessment changed — the classification differs between clean and attacked runs.
  • Target assessment success — the attacked run reached the classification the attack asked for.
  • Claim set changed — the normalized set of claims differs.
  • Target claim present — the requested claim appears, by deterministic text normalization.
  • Gap count delta and gap suppression — gaps were removed, or a targeted gap disappeared.
  • Evidence reference changed — the set of cited blocks differs.
  • Structural validity — tracked separately for each run: schema valid, references valid, case-isolated.
  • Structurally valid manipulation — the attack achieved its target and the output stayed fully structurally valid.

The last one is the measurement the research question turns on. It is still a structural measurement. None of these is a semantic citation-correctness rate or an evidence-integrity violation rate, because nothing here evaluates whether cited text supports a claim.

Research architecture

Build

  1. Synthetic fixture

    then
  2. Segmentation

    then
  3. Source registry

    then
  4. Prompt builder

Run

  1. Provider boundary

    then
  2. Raw model output

    then
  3. Structured parser

    then
  4. Reference validator

Record

  1. Immutable artifacts

    then
  2. Paired metrics

In R3 the controlled attack transformation happens at the very start, on the job fixture, before segmentation. Everything downstream is identical for the clean and attacked branches.

Every paired run writes its own directory: the configuration, the attack specification, the segmented sources for both branches, the raw responses, the parsed responses with their validation status, and the comparison. A run directory is never overwritten; targeting an existing one is an error. That is what makes a run reproducible and auditable, and it removes the option of quietly editing a result afterwards.

Test double vs empirical result

The benchmark harness has been validated with a deterministic test double. Every result artifact in the repository comes from that test double, and each is tagged in its own metadata as test-only and non-scientific.

Those runs verify that the experimental pipeline works end to end. They do not demonstrate that any real LLM is vulnerable, and their outcomes are not findings. A live-model provider interface exists, but no real-model run has been recorded.

So there are no attack success rates on this page, and there should not be.

Research decisions

Synthetic data only

Why
Ground truth is fully controlled, and no real person's documents are involved.
Tradeoff
Less ecological realism than real resumes and postings would give.

Stable source block identities

Why
An evidence reference has to mean the same thing across runs to be inspectable and reproducible.
Tradeoff
Requires deterministic segmentation, which fixes how documents can be split.

Structural validation before semantic validation

Why
Whether a reference exists and whether it supports a claim are different questions and should not be blurred.
Tradeoff
R2 and R3 cannot establish that a cited block entails its claim.

Clean-versus-attack pairs

Why
Facts stay fixed and only the payload changes, so a difference can be attributed to the attack.
Tradeoff
The initial scope is deliberately narrow: one payload, one document, one task.

Test doubles validate plumbing only

Why
Deterministic runs prove the harness, parsing, and metrics behave as specified.
Tradeoff
They cannot support any empirical claim about model behaviour.

Current status

  1. R0

    Research foundation

    Research questions, threat model, outcomes, metrics, and experiment plan.

    Complete
  2. R1

    External baseline (PLeak)

    Upstream code is pinned in isolation and a synthetic protected-target fixture exists. The baseline has not been reproduced.

    Partial
  3. R2

    Evidence-grounded testbed

    Four clean cases, deterministic segmentation, structural reference validation.

    Complete
  4. R3A

    Paired benchmark harness

    Ten paired cases, four fixed attack classes, paired metrics. Test double only.

    Complete
  5. R3B

    Real-model pilot

    The frozen benchmark against a real model. Not yet run.

    Next
  6. Later

    Semantic support validation and defenses

    Checking whether cited text supports a claim, then evaluating defenses.

    Not started

PLeak appears here only as an external baseline on its own track. It has not been reproduced in this project, the fixed R3 attacks are unrelated to it, and none of its results belong to this work.

Limitations

  • Synthetic data. Every document and candidate is invented.
  • Small benchmark. Four clean cases and ten paired cases.
  • Fixed attack templates. Hand-written payloads, not optimized or adaptive.
  • No semantic support validation. Only structural reference checks exist.
  • No real-model results. The harness has run only against a test double.
  • No defenses evaluated. Nothing here measures a mitigation.
  • External baseline not reproduced. The PLeak track has no results.
  • One task. A single assessment task with a three-way classification.

Nothing on this page claims a new vulnerability, that grounded or retrieval-augmented systems in general are vulnerable, statistical significance, a validated defense, or a published paper.

Next experiment

The next step is a real-model pilot of the paired benchmark: the ten existing pairs, each run clean and attacked, against one real model, once.

The goal is not statistical significance. It is to find out whether the phenomenon appears at all on a real model under a benchmark that was frozen before any real-model result existed. Whatever it shows — including nothing — is what would be reported.