team research — afrl rome research site
AFRL_AIAI
Predicted vs. realized prompt-injection severity: pitting OWASP's predicted-risk score (AIVSS v0.8) against a DoD outcome-based measurement (CNSSI 1253 / FIPS 199) for adaptive indirect prompt injection on open-weight agentic LLMs.
Context. Summer 2026 internship at the Air Force Research Laboratory's Rome Research Site, as an AI Security Research Intern on a three-person team, which I led. The findings summary was reviewed and approved for public release (AFRL-2026-3601, PA-cleared 11 August 2026). The underlying codebase is not yet public, so this page describes the research rather than linking a repo.
The problem
Two families of AI-risk severity instruments get used in practice: predicted-severity scoring systems purpose-built for agentic AI (like OWASP's new AIVSS), and outcome-based categorization instruments used for real authorization-to-operate decisions (like the DoD's CNSSI 1253/FIPS 199). Nobody had pointed both at the same real attacks against real, deployed agentic models to see whether they actually agree, or even can agree in principle. This project builds a realistic HR-records agent, attacks it with adaptive indirect prompt injection across 13 distinct injection goals and multiple open-weight LLMs, and measures both instruments against the same real transcripts.
The approach
An 18-tool HR-records AgentDojo agent (five tool categories: read/PII access, modify, destructive, communication/export, workflow), eight legitimate user tasks × three specificity variants, and 13 injection goals — nine single-objective (confidentiality, integrity, availability) plus four compound goals — form a 312-scenario grid per model. Attacks escalate from a static pass to five adaptive passes, then to a frozen, SHA-pinned corpus of 16 real-world indirect-injection techniques drawn from garak, InjecAgent, published proof-of-concepts, and named CVE incidents.
Two independent scoring pipelines then run over the same evidence: a
from-spec AIVSS v0.8 calculator built on a tested port of FIRST.org's
own CVSS v4.0 reference implementation, scored once per injection goal;
and per-objective CNSSI/FIPS-199 checkers that grade each attack
transcript into a full confidentiality/
Tech stack
Key finding
AIVSS's predicted score turned out to be structurally blind to model and task: it varies only by injection goal, landing every model at "Critical" identically by construction. CNSSI's realized trigraph, by contrast, varied by model, task, and goal. That's a measurable, load-bearing granularity mismatch between what a predicted-risk score claims and what actually happened when the attacks were run for real, which matters directly to anyone deciding whether a predicted score can stand in for real red-team evidence in an AI system's authorization package. The project frames this explicitly as a demonstration, not a statistical validation.
Measurement integrity, treated as a first-class result. A hand-labeled checker-validation pass caught a real grading bug: a model's string-literal "null" argument was silently zeroing a tool's result set, so genuine attack attempts were being logged as never-reached rather than resisted. The fix recovered those scenarios and re-scored them properly — a self-audit of the measurement instrument itself, on top of the audit of the models under test.
Team & credit
I led the three-person team through the summer and owned the scenario and attack harness — the 312-scenario grid, the run orchestrator, and the model roster. Tyler owned severity measurement; Alex owned the presentable view and the assurance case. Credit to both for the parts of this that are theirs.