Skip to main content
Defender protects your AI agents from prompt injection attacks by scanning API tool call responses before they reach your LLM. When an MCP tool returns data from a third-party provider (emails, CRM records, documents), that data could contain instructions designed to hijack your agent’s behavior. Responses are intercepted and classified by Defender, and depending on the Protection Mode you choose, Defender logs the verdict, removes high-risk content, or blocks the response before it reaches your agent.

How it works

Flow diagram. A tool call response enters the Scan Mode split: Light or Hybrid run the Light Scan of pattern matching and an ML classifier, Deep goes straight to the Deep Scan LLM review, and an uncertain Light Scan result escalates to the Deep Scan in Hybrid mode. Both scans feed Protection Mode, which acts on the verdict: Monitor passes the response unchanged with the verdict logged, Sanitize removes high-risk content, and Block refuses high or critical risk responses and sanitizes the rest
Defender’s standard check is the Light Scan, two checks run together on every response:
  • Pattern matching: A fast rule-based scan for known prompt injection signatures and risky field patterns, with negligible latency.
  • ML classification: A local ML model (MiniLM) scores the content for novel or subtle attacks that pattern matching would miss. It scans the SFE-filtered payload (or the tier2Fields subset when configured), and reports its score as tier2Score in the response metadata.
Responses that exceed the configured size limits can skip scanning entirely (see Advanced settings). The Deep Scan is a third check on top: an LLM reviews the response for risks the Light Scan misses. It runs only in StackOne’s hosted service, not in the open source @stackone/defender package. The Scan Mode setting decides which responses get a Deep Scan (see Scan mode). Risk level and scan metadata are returned alongside every response in every mode, so you can see what Defender detected even when it changed nothing.

When to use Defender

  • You are building AI agents or MCP-based workflows that process third-party API responses
  • Your integrations handle sensitive data such as emails, files, calendar events, or CRM records
  • You want to observe risk signals on tool call responses without necessarily blocking them

Configure from the dashboard

Navigate to your project in the StackOne dashboard, then open the Defender tab in project settings. This is the baseline for every account in the project.
Defender Settings
Defender settings apply project-wide. Per-account and per-request overrides take precedence where supported.

Core settings

Protection mode

Start in Monitor to see what Defender would act on in your traffic by reviewing logs. Move to Sanitize or Block when you are ready to enforce.

Scan mode

Light is the default. The Light Scan is the same open source engine as the @stackone/defender package; the Deep Scan runs only within StackOne’s hosted service, so Hybrid and Deep are available here but not in the package.

Advanced settings

Implement it

Tool Defense

Override Defender per toolset in the Agent SDK.

FAQ

It depends on when your project was created. New projects have Defender on. Older projects start with it off: turn it on from the Defender tab in project settings when your agents consume third-party data. Once on, it runs in Monitor, so tool results are unchanged until you choose Sanitize or Block.
The Light Scan adds little: pattern matching is negligible, and the ML classifier runs locally, so there is no external API call. The Semantic Field Extractor trims metadata and identifier fields before classification to keep latency low; for typical responses the added latency is under 100ms. A Deep Scan adds an LLM review to the responses it covers, which costs more time than the local checks.
The tool call returns an error to your agent indicating the response was blocked. The agent can handle this like any other tool error: retry, skip, or surface it to the user.
Yes. That is what Monitor mode does, and it is the default. Defender still scans and returns riskLevel, tier2Score, and detections in the response metadata, which you can inspect in your logs, but the tool result reaches your agent unchanged.
No. The Light Scan’s classification model runs locally within StackOne’s infrastructure, and the Deep Scan’s LLM is a model StackOne deploys and operates on its own inference endpoints. Your responses are not sent to a third-party AI service, and neither model is trained on your data.