# Evaluate Trismag in one agent workflow

A two-week evaluation guide, not an offer of dedicated capacity or a paid plan.

Choose one working agent, one owner, and one source of untrusted input. Document
which tools it can use, which actions require approval, and what happens if a
scan fails. A LOW verdict must never grant permissions on its own.

Start with synthetic or deidentified content. Do not paste credentials or customer
secrets into the demo. Reference examples do not make network requests; the live
buttons do. The homepage no-key path provides risk only. Your free key adds
findings; current allowances are published at https://trismag-api.fly.dev/health.

First week: connect the HTTP API in observation mode. Handle 400/401/402/413/422/429,
timeouts, unknown risks and incomplete scans. Respect Retry-After. Record aggregate
counts in your own environment; you do not need to share raw payloads with us.

Second week: review both warnings and a sample of LOW results. Separate useful
warnings, benign flags, and missed planted attacks. Track integration effort,
review time and whether legitimate agent tasks still complete successfully.

At the end: decide whether to continue, change the workflow, or stop. Specify
which input categories were evaluated and the sample size. Do not extrapolate
from a few successes to universal protection.

Trismag is deterministic pattern matching, without an LLM in its detection loop.
Its public synthetic baseline includes misses and false alarms. It is an early
warning layer, alongside permissions, egress controls and independent approvals.


Current September 11, 2026: https://trismag.dev/privacy/ and https://trismag.dev/terms/ apply. Operator Jonathan Hill: Jonhillnj@gmail.com. Record request duration, input byte size, endpoint and whether the request was cold, using aggregate measurements in your environment. Compare scan overhead with legitimate task completion; do not send raw production content as telemetry.

The original pre-change baseline had ten misses. v0.4.4 now detects those known examples; it still has blind spots. The updated 24/24 result uses development cases, with 4/16 benign examples HIGH+. Read the versioned reports rather than treating this as a real-world percentage.
