team research — afrl rome research site

AFRL_AIAI

Predicted vs. realized prompt-injection severity: pitting OWASP's predicted-risk score (AIVSS v0.8) against a DoD outcome-based measurement (CNSSI 1253 / FIPS 199) for adaptive indirect prompt injection on open-weight agentic LLMs.

Case no.
01
Status
PA CLEARED
Filed
MMXXVI

Context. Summer 2026 internship at the Air Force Research Laboratory's Rome Research Site, as an AI Security Research Intern on a three-person team, which I led. The findings summary was reviewed and approved for public release (AFRL-2026-3601, PA-cleared 11 August 2026). The underlying codebase is not yet public, so this page describes the research rather than linking a repo.

7–8
open-weight LLM agents attacked for real, on rented GPU infrastructure
2,000+
real attack transcripts collected (not simulated)
312
scenario grid per model (tasks × specificity × injection goals)
7.5–39.3%
range in realized attack success across the 7 models — one predicted score couldn't tell them apart

The problem

Two families of AI-risk severity instruments get used in practice: predicted-severity scoring systems purpose-built for agentic AI (like OWASP's new AIVSS), and outcome-based categorization instruments used for real authorization-to-operate decisions (like the DoD's CNSSI 1253/FIPS 199). Nobody had pointed both at the same real attacks against real, deployed agentic models to see whether they actually agree, or even can agree in principle. This project builds a realistic HR-records agent, attacks it with adaptive indirect prompt injection across 13 distinct injection goals and multiple open-weight LLMs, and measures both instruments against the same real transcripts.

The approach

An 18-tool HR-records AgentDojo agent (five tool categories: read/PII access, modify, destructive, communication/export, workflow), eight legitimate user tasks × three specificity variants, and 13 injection goals — nine single-objective (confidentiality, integrity, availability) plus four compound goals — form a 312-scenario grid per model. Attacks escalate from a static pass to five adaptive passes, then to a frozen, SHA-pinned corpus of 16 real-world indirect-injection techniques drawn from garak, InjecAgent, published proof-of-concepts, and named CVE incidents.

Two independent scoring pipelines then run over the same evidence: a from-spec AIVSS v0.8 calculator built on a tested port of FIRST.org's own CVSS v4.0 reference implementation, scored once per injection goal; and per-objective CNSSI/FIPS-199 checkers that grade each attack transcript into a full confidentiality/integrity/availability trigraph — deliberately never collapsed into one scalar, since CNSSI itself rejects that simplification. A hand-labeled 21-scenario validation pass checked the checkers themselves against ground truth before any result was trusted.

Tech stack

Python 3.13 AgentDojo (vendored fork) vLLM Rented GPU inference OWASP AIVSS v0.8 CVSS v4.0 CNSSI 1253 / FIPS 199 pytest

Key finding

AIVSS's predicted score turned out to be structurally blind to model and task: it varies only by injection goal, landing every model at "Critical" identically by construction. CNSSI's realized trigraph, by contrast, varied by model, task, and goal. That's a measurable, load-bearing granularity mismatch between what a predicted-risk score claims and what actually happened when the attacks were run for real, which matters directly to anyone deciding whether a predicted score can stand in for real red-team evidence in an AI system's authorization package. The project frames this explicitly as a demonstration, not a statistical validation.

Measurement integrity, treated as a first-class result. A hand-labeled checker-validation pass caught a real grading bug: a model's string-literal "null" argument was silently zeroing a tool's result set, so genuine attack attempts were being logged as never-reached rather than resisted. The fix recovered those scenarios and re-scored them properly — a self-audit of the measurement instrument itself, on top of the audit of the models under test.

Team & credit

I led the three-person team through the summer and owned the scenario and attack harness — the 312-scenario grid, the run orchestrator, and the model roster. Tyler owned severity measurement; Alex owned the presentable view and the assurance case. Credit to both for the parts of this that are theirs.