How it works
- Pattern matching: A fast rule-based scan for known prompt injection signatures and risky field patterns, with negligible latency.
- ML classification: A local ML model (MiniLM) scores the content for novel or subtle attacks that pattern matching would miss. It scans the SFE-filtered payload (or the
tier2Fieldssubset when configured), and reports its score astier2Scorein the response metadata.
@stackone/defender package. The Scan Mode setting decides which responses get a Deep Scan (see Scan mode).
Risk level and scan metadata are returned alongside every response in every mode, so you can see what Defender detected even when it changed nothing.
When to use Defender
- You are building AI agents or MCP-based workflows that process third-party API responses
- Your integrations handle sensitive data such as emails, files, calendar events, or CRM records
- You want to observe risk signals on tool call responses without necessarily blocking them
Configure from the dashboard
Navigate to your project in the StackOne dashboard, then open the Defender tab in project settings. This is the baseline for every account in the project.
Defender settings apply project-wide. Per-account and per-request overrides take precedence where supported.
Core settings
Protection mode
Scan mode
Light is the default. The Light Scan is the same open source engine as the
@stackone/defender package; the Deep Scan runs only within StackOne’s hosted service, so Hybrid and Deep are available here but not in the package.
Advanced settings
Implement it
Tool Defense
Override Defender per toolset in the Agent SDK.
FAQ
Do I need to enable Defender?
Do I need to enable Defender?
It depends on when your project was created. New projects have Defender on. Older projects start with it off: turn it on from the Defender tab in project settings when your agents consume third-party data. Once on, it runs in Monitor, so tool results are unchanged until you choose Sanitize or Block.
Does Defender add latency?
Does Defender add latency?
The Light Scan adds little: pattern matching is negligible, and the ML classifier runs locally, so there is no external API call. The Semantic Field Extractor trims metadata and identifier fields before classification to keep latency low; for typical responses the added latency is under 100ms. A Deep Scan adds an LLM review to the responses it covers, which costs more time than the local checks.
What happens when a response is blocked?
What happens when a response is blocked?
The tool call returns an error to your agent indicating the response was blocked. The agent can handle this like any other tool error: retry, skip, or surface it to the user.
Can I see what Defender flagged without blocking?
Can I see what Defender flagged without blocking?
Yes. That is what Monitor mode does, and it is the default. Defender still scans and returns
riskLevel, tier2Score, and detections in the response metadata, which you can inspect in your logs, but the tool result reaches your agent unchanged.Will my data be used to train the AI model?
Will my data be used to train the AI model?
No. The Light Scan’s classification model runs locally within StackOne’s infrastructure, and the Deep Scan’s LLM is a model StackOne deploys and operates on its own inference endpoints. Your responses are not sent to a third-party AI service, and neither model is trained on your data.