TRUST CENTER / OPEN PREVIEW
Know the signal.
Know its limits.
Security claims should be specific enough to inspect. Here is what Trismag does, where it fits, and what remains your responsibility.
DETECTION UPDATE / V0.4.4DEVELOPMENT REGRESSION RESULTS
Known gaps.
Measured changes.
All ten previously missed attacks now produce warnings.
Version 0.4.4 adds rules for answer manipulation, task substitution, payment-target replacement, secrets placed in search queries, and separated-letter overrides. The rules were developed using the original misses and additional synthetic examples. These results are not independent, held out, or a real-world detection percentage.
| COUNTED AS FLAGGED | ATTACKS FLAGGED AFTER CHANGE | BENIGN INPUTS FLAGGED |
|---|
| MEDIUM or above | 24 / 24 | 7 / 16 |
| HIGH or above | 24 / 24 | 4 / 16 |
All 40 original cases completed locally and returned matching risk levels from deployed API v0.4.4 on September 11, 2026 (UTC). The benign risks are unchanged: four of sixteen still flag at HIGH or above. A separate development set contains 20 different attack wordings and 30 nearby benign examples; all 20 attacks flag and none of those 30 benign examples flag. That set was also used during development. This closes documented examples; it does not establish that new attacks will be detected or that an agent will avoid harmful actions.
The original report and its ten misses remain published below. The playground uses a different, still-undetected request about an agent’s operating constraints to demonstrate that LOW can remain hostile.
MEASURE THE LIMITSPRE-CHANGE BASELINE / V1
Evidence you can
take apart.
40 constructed cases. The misses stay in.
This September 2026 baseline contains 24 planted attacks and 16 benign examples across text, conversations, and tool metadata. Codex authored the cases and labels. It is not independent, held out, or representative of production traffic. Detection rules had not been changed for that initial evaluation. The later v0.4.4 results above use this corpus during development.
| COUNTED AS FLAGGED | ATTACKS FLAGGED | BENIGN INPUTS FLAGGED |
|---|
| MEDIUM or above | 14 / 24 | 7 / 16 |
| HIGH or above | 14 / 24 | 4 / 16 |
All 40 cases completed locally and returned matching risk levels from deployed API v0.4.2 on September 11, 2026 (UTC). At that time, ten attacks were missed at both thresholds, including indirect answer manipulation and invoice redirection. Raising the threshold reduced benign flags in this set, but did not establish a safe policy. Quoted attacks can still look like attacks to a pattern matcher.
These are detector results, not measurements of whether an agent performed a harmful action. Local engine timings in the downloadable report exclude the network and are not a latency promise.