AI Security Research
Evidence Integrity Research
A research effort studying prompt injection against evidence-grounded LLM applications: systems that are expected to tie each claim they make to supporting source material. So far it has a threat model, a deterministic evidence-grounded testbed, and a clean-versus-attack paired benchmark, all on synthetic data. The harness has only been exercised with a deterministic test double, so there are no empirical findings yet.
- Research question
- Can adversarial instructions embedded in source documents corrupt claim-level evidence integrity in evidence-grounded LLM applications?
- Status
- Methodology and benchmark harness built; no real-model results yet
- Topics
- Evidence-grounded AI systems
- Prompt / evidence integrity
- Adversarial evaluation
- Reproducible experiments
- Repository
- Not public yet
- Paper
- No paper yet
- Experiments
Evidence-grounded testbed Complete
Synthetic cases, deterministic segmentation, structured assessments, and structural reference validation.
Paired attack benchmark Harness complete (test double only)
Ten clean-versus-attack pairs across four fixed attack classes, with deterministic paired metrics.
Real-model pilot Not run
The same frozen benchmark run once against a real model.
- Results
- No results yet