What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
- Current maintenance status and activity of the project: confirm
- Required external dependencies for each detection layer: verify against project docs
Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.
What it does
Rebuff inspects user input before it reaches your prompt and estimates whether it is an injection attempt. It is built as four independent layers because no single technique holds up alone. The first is heuristic: pattern matching for the recognizable shapes of injection, instructions to disregard prior context, attempts to elicit the system prompt, delimiter breakouts. The second sends the input to a language model with a classifier prompt asking whether the text is trying to subvert instructions. The third embeds the input and compares it against a vector store of past attacks, so one that worked before is recognized in variant form. Each layer returns a score and you set the rejection threshold.
The fourth layer works on the output side. Rebuff inserts a canary token into the prompt, a unique marker the model is told to keep private, then checks responses for it. If the token appears in output, the system prompt leaked and you have direct evidence of extraction rather than a guess. Detected attacks can be written back to the vector store, so the deployment learns from what it sees.
Where it fits
In the request path of an LLM application, called before the prompt is assembled and again on the response before it is returned. It is a library, so the application team owns it. Two of the four layers need external services, a model endpoint for the classifier and a vector store for similarity, which adds a dependency to every request you screen.
Strengths
- Layering independent methods is the right architecture for a problem where any single filter is bypassable.
- Canary tokens give a high confidence signal of system prompt leakage, which most detectors cannot produce at all.
- Confirmed attacks feed the similarity store, improving detection against variants of what you have already seen.
- Small, readable codebase, useful as a reference implementation even if you do not deploy it.
Limitations
- Heuristic and similarity layers detect known shapes. Novel phrasing, obfuscation and non English payloads get through, and the classifier is itself a model that can be talked out of its judgment.
- The classifier and embedding calls add latency and inference overhead to every screened request, and the project reads as a reference implementation more than a hardened control.
- Input screening is one control among several. It does nothing about what an agent may do once an injection succeeds.
Who it suits
Teams that want to understand how injection detection works, or that need a defensible first filter in front of a low risk LLM feature. Not the right foundation for a high risk agentic system, where constraining tool permissions matters more than filtering the input.
Used Rebuff? Recommend it under your own name and title.
Recommend this tool