AppSecNews
AI Security Open source Growing

NVIDIA NeMo Guardrails

by NVIDIA

Toolkit for defining rails on an LLM conversation using a dedicated modeling language, controlling input, output, topic, retrieval and tool execution.

Visit github.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run NVIDIA NeMo Guardrails in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 1 point in this profile is not yet confirmed against vendor documentation.
  • Current rail types and bundled third-party detector integrations, confirm against project docs

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

NeMo Guardrails puts a state machine between the user and the model. You describe conversational behavior in Colang, a purpose-built modeling language, by defining canonical forms for what users say and flows for how the application should respond. At runtime the toolkit embeds the incoming message, matches it by vector similarity against the canonical forms you defined, and executes the matching flow. That is a different mechanism from a classifier: instead of scoring a message for maliciousness, you are declaring which conversations exist and what happens in each.

Rails apply at different points. Input rails inspect or reject what the user sent. Dialogue rails steer the conversation itself, which is how you keep a support assistant from discussing anything but support. Retrieval rails filter chunks coming back from a knowledge base, the control point for injection hidden in indexed documents. Execution rails constrain tool and action calls. Output rails check the response before it reaches the user, and can delegate to fact-checking, hallucination scoring or a third-party safety model.

Where it fits

In the application path in production, self-hosted. In practice it is a joint effort: application teams write the flows, security reviews the input, retrieval and execution rails. It assumes your application has a bounded purpose you can enumerate. An open-ended general assistant is hard to express as dialogue rails, and that mismatch is the most common reason adoption stalls.

Strengths

  • Rails at five distinct points let you place control where the risk is, rather than filtering everything through one input check.
  • Retrieval rails address indirect prompt injection through documents, which input filtering alone cannot reach.
  • Dialogue rails give topic control that is auditable: the allowed conversations are written down, not inferred from a classifier's scores.
  • Pluggable output checks let you bring your own safety model instead of relying on whatever ships in the box.

Limitations

  • Colang is a new language with its own semantics. The learning curve is the main adoption cost and it is not small.
  • Rails add work per turn: embedding lookups plus, for some rails, extra model calls. Latency compounds as you add them.
  • Canonical forms need curation and do not generalize as well to unseen phrasing as the examples suggest, so coverage is ongoing maintenance.

Who it suits

Well matched to teams building task-scoped assistants with a definable conversational scope, who want control expressed as explicit policy they can review. A poor match for an open-ended assistant, for a team that wants a drop-in filter with no configuration work, or for anyone unwilling to take on a domain-specific language as part of their security stack.

Used NVIDIA NeMo Guardrails? Recommend it under your own name and title.

Recommend this tool