Agent Under Load
Research report / agent systems and evaluation

When the alert stays the same, does the answer change?

A tool-using triage agent can look busy while relying on the alert it was given. This study holds the alert fixed, varies the telemetry, and measures whether four models make use of the evidence they retrieve.

Yasin DehfouliATT&CK T1003.00180 cases · 24 capturesSeptember 2026
+0.40within-rule macro-F1 gain from telemetry for Luna
4 × 6 × 3models, evidence conditions and repeats per case
24independent captures used for uncertainty intervals
13.1%of candidate injection placements are mountable
The central result

One model's verdict depends on the capture

Six detection rules fire on captures from both classes. Within a rule, the alert is identical. On those 50 cases, gpt-6-luna scored 0.72 with telemetry and 0.32 when forced to guess from the alert alone. The paired difference is +0.40 [0.27, 0.52], with exact McNemar p = 0.036.

The same alert, two information states
Macro-F1 on the preregistered 50-case subset. The interval belongs to the paired difference, not to either score separately.
Within-rule comparisonMacro F1 is 0.32 for an alert-only forced guess and 0.72 with telemetry. The difference is 0.40 with a 95 percent interval of 0.27 to 0.52. 00.20.40.60.8 Forced alert guessTelemetry available0.320.72Δ +0.40 [0.27, 0.52]
The analysis was specified before the subset result was computed. Luna was selected after the full matrix was inspected, so this supports evidence sensitivity for this model and corpus, not a claim that it is the best model. Result record.
A separate comparison stays unresolved. On all 80 cases Luna scored 0.73 and the no-model heuristic 0.60. Luna alone got 16 cases right; the heuristic alone got 10. McNemar p = 0.33. The observed score gap is not evidence of a reliable advantage over that heuristic.
System design

An explicit graph makes every step inspectable

The LangGraph state carries the transcript, last reply, turn and tool counts, rejections and repair count. Only the investigate node calls a model. Routing to validation and back is explicit, so a rejected answer remains part of the trace.

Agent architecture with a model investigation node, scoped tool loop, validation and repair path; alongside four evaluated models, four controls, 32,122 provider requests and the paired within-rule result
The request count covers the primary four-model, six-condition matrix: 5,760 completed trajectories and 25,893 model turns. It includes provider retries and excludes smoke runs, the wording check and the unfinished control study. Per-condition accounting · Counting method. The four controls are implemented; their benefit under attack has not been measured.

One code path, controlled conditions

The reference condition exposes the tools and the case's own capture. Alert-only removes all tools. Rule-only exposes the rule but no events. Forced alert-only removes the option to abstain so the model's prior can be measured. Swap conditions present a donor capture under the same alert.

The final verdict has a validated schema and event citations. One repair is permitted after rejection; 12 turns is the limit. Unanswered and rejected trajectories stay in the denominator.

What the run ledger retains

IdentityModel, condition, case, prompt digest and harness version prevent an altered run from silently resuming.
OperationTool calls, provenance, citations, token use, latency, cost and failures are recorded.
BoundaryRun and model spending caps account for work in flight; the audit trail hash-chains its entries.
Five tools, one path for attacker-written text
The labels describe the strongest content each tool can return, not a claim that every value is malicious.
lookup_rulepinned Sigma rule text
describe_captureevent IDs and field presence
count_eventscounts without field values
lookup_attack_techniquestatic local ATT&CK table
query_eventsfull event fields; writable text can reach the model
Provenance labels, structured ingestion, enforced citations and capability scope can be enabled separately. They are implemented controls, but their effect on attack outcomes has not been measured. Architecture and trust boundary.
Experimental design

Evidence conditions isolate what the model can know

Each model receives 80 rule/capture cases, six conditions and three trajectories per case. The truth is 44 true positives and 36 false positives from 24 captures. A case answer is the majority of its repeats.

Evidence condition matrixReference uses alert, rule and own capture. Alert-only and forced alert-only use only the alert. Rule-only uses alert and rule. Cross and same swaps use alert, rule and a donor capture. AlertRuleOwn eventsDonor events ReferenceAlert onlyForced alert guessRule onlyCross-label swapSame-label swap must decide

The four-model matrix is 4 × 80 × 6 × 3 = 5,760 planned trajectories. The later technique-question run is a separate diagnostic. Macro-F1 is reported instead of bare accuracy. Its 95% intervals come from 2,000 resamples of captures, stratified by label. Committed analysis files.

Three shortcuts were checked before model runs
The alert and corpus were revised so a model could not get credit for a label proxy.

Lab and count leakage

All true-positive hosts shared one domain, while false-positive hosts used other domains. Host and capture names were replaced with capture-specific pseudonyms. An inconsistent alert match count gave a simple threshold 75/80 correct; the count was removed.

Rule prior

Rule identity alone fit about 61/80 in sample. Leave-one-capture-out evaluation reduced this to 45/80, macro-F1 0.56. It remains an explicit baseline rather than being mistaken for evidence use.

Account names remain a partial residual confound. Design decisions and leak checks.
Four-model comparison

More investigation does not guarantee a better verdict

Luna has the highest observed reference score. The other three reference scores are close to their own forced guesses from the alert. Without the forced instruction, models mostly abstain when evidence is unavailable.

Reference macro-F1 and each model's forced alert guess
Horizontal lines show 95% capture-bootstrap intervals for reference. The gray dashed line is the 0.60 heuristic. Values are per model, not pairwise model significance tests.
Reference performance across modelsLuna 0.73 interval 0.60 to 0.86, forced guess 0.37. DeepSeek 0.59 interval 0.46 to 0.70, guess 0.52. OSS 20b 0.57 interval 0.44 to 0.71, guess 0.54. OSS 120b 0.46 interval 0.34 to 0.58, guess 0.57. 0.00.20.40.60.8 gpt-6-lunadeepseek-v4-progpt-oss-20bgpt-oss-120b 0.730.590.570.46
reference and CIforced alert guess┄ 0.60 heuristic

Reliability is part of the result

Reference case answer rates were 100%, 96%, 76% and 84% in the order shown. Three-repeat agreement was 81%, 89%, 38% and 69%. The gpt-oss-20b result is limited by incomplete answers and unstable repeats, not just classification error.

Cost and turn limits

The six-condition matrix cost about $1.13 for Luna, $16.26 for DeepSeek, $0.83 for gpt-oss-20b and $1.35 for gpt-oss-120b, excluding smoke runs. DeepSeek's swap cells often reached the 12-turn limit, with 47% of cross-label trajectories hitting it.

Interpreting a failed probe

The swaps changed whether the alert was supported

A donor capture often contained no event that fired the rule named by the alert. Calling that alert unsupported can be a sound answer. Swap scores therefore cannot be read as ordinary classification accuracy.

Does the named rule fire on the donor?
Only a minority of donors preserved this prerequisite for the original question.
19 / 80cross-label donors support the alert
24% support · 76% do not
25 / 80same-label donors support the alert
31% support · 69% do not
Exploratory split: among supported cross-label donors, Luna followed the donor's true label in 3/4 true-positive cases and 12/15 false-positive cases. For supported same-label donors, the verdict held in 16/16; where the rule did not fire it held in 2/28. Swap split.
Six shared rules did not contribute equally
Share of true-positive × false-positive case pairs where Luna got both right. Each line is one rule, ordered by pair count.
Within-rule pair outcomesThe six rules have pair success of 80 percent of 35 pairs, 33 percent of 30, 38 percent of 16, 100 percent of 16, 50 percent of 8, and 0 percent of 2. 0%25%50%75%100% Rule 4b447e9dRule 4a1b6da0Rule 250ae82fRule 5ef9853eRule 678dfc63Rule 962fe167 28/3510/306/1616/164/80/2
One rule accounts for 16 flawless pairs; another has none of two. The result is not uniform across rule families. The IDs shown are shortened display labels; the full IDs and case counts are in the result record.
Label and question audit

The prompt asked a wider question than the labels

The labels identify credential theft from LSASS memory. The original prompt asked if “the activity the rule describes” happened. An LSASS access rule can fire on another technique without credential theft. The author reviewed all 17 false-positive captures against dataset descriptions before running a wording check; none required relabeling.

Changing one definition shifted errors toward the intended label
This diagnostic condition was chosen after the original misses were seen. It is not a replacement model score.
Original and technique-specific questionMacro F1 rises from 0.73 to 0.84. False positive specificity rises from 56 to 72 percent. True positive recall rises from 91 to 93 percent. Macro-F1FP specificityTP recall025%50%75%100%.8472%93%
original questiontechnique-specific question
Paired macro-F1 difference +0.10 [0.01, 0.23] by capture bootstrap; exact McNemar p = 0.12. TP recall changes from 40/44 to 41/44. FP specificity changes from 20/36 to 26/36. Computed check.
What the model cited: among 232 decisive verdict trajectories in the diagnostic run, 206 cited a process-access event and 26 relied on other events alone. The agent usually read the LSASS access event that caused the alert. That observation fits the remaining false-positive pattern, but it does not prove why the model erred.
Attack surface and controls

The possible injection channel is narrow and measurable

The attack placement check tests 43 payload and field positions over 44 true-positive cases. A placement is allowed only where the event already carries that field and an adversary can write it. It accepts 248 of 1,892 candidates, or 13.1%. This is feasibility, not attack success.

Mountable attempts by field
The length of each line is the fraction of that field's candidate placements that passed the realism check. Counts show mountable / attempted.
CommandLine
160 / 704
ParentCommandLine
7 / 44
ServiceFileName
4 / 88
Details
0 / 396
ScriptBlockText
10 / 220
TargetFilename
20 / 176
Image
24 / 88
Product
7 / 44
Description
7 / 44
Company
7 / 44
ServiceName
2 / 44
Registry Details fails all 396 candidates because none of these rules' candidate events are registry writes. The summed field counts equal 248/1,892. Mountability record.

Controls built into the harness

Provenance tags tell the model who could have written each value. Structured ingestion separates typed fields and drops redundant rendered messages. Citation enforcement checks quotes against the capture. Capability scope checks requested actions against a session grant.

Measurement still pending

The adversarial runner supports payloads, spending limits, resume and control configurations. No model-backed attack success rate or control ablation has passed review. An earlier stacked-control run produced many null verdicts and has not been audited; it is not a result.

Scope and reproducibility

What this result can carry

This is an evaluation of one technique on recorded lab data, with seven attack captures across two labs. It tests whether an agent's answer uses available evidence. It does not establish deployment performance, general prompt-injection resistance or superiority over the heuristic.

Unit of uncertainty24 captures, stratified by label. The 80 cases share source recordings.
Model selectionLuna was selected after the full matrix; the within-rule criterion was registered before that subset was computed.
Remaining confoundsAccount names, one technique, two labs and capture-specific behavior limit transfer.

Trace the figures