AppSecNews
AI Security Open source and commercial Established

Arize AI

by Arize AI

LLM and ML observability platform that captures application traces and runs evaluations over them, with an open source tracing and evaluation component.

Visit arize.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run Arize AI in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
  • Feature boundary between the open source component and the commercial platform: confirm
  • Which safety and security evaluators ship built in: verify against vendor docs

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

Arize instruments AI applications and collects what they do. Instrumentation is built on OpenTelemetry with a semantic convention for LLM workloads, so a trace captures a request end to end: the prompt as sent including retrieved context, the model and parameters, tool calls with their arguments and results, token counts, latency, and the final response. For agentic applications the trace is a tree, and the nested spans show which step produced a bad answer.

On top of the traces sits an evaluation layer. Evaluators run over recorded spans and score them, typically using a language model as judge against a rubric, covering answer relevance, retrieval quality, hallucination, and whether a response contains content it should not. That mechanism is what gives the platform a security role: evaluate inputs for injection attempts and outputs for leaked sensitive data, then alert on rates rather than individual events. Datasets and experiments let you pin inputs, replay them against a changed prompt or model, and compare scores. The open source component provides tracing and evaluation you can run locally or self host, with the commercial platform adding managed storage, monitoring and collaboration.

Where it fits

Development and production both. Engineers use the local tracing view while building, and the same instrumentation feeds a hosted or self hosted backend once the application ships. Ownership usually sits with the AI engineering team, not security: security gets value here by defining evaluators and consuming the metrics, not by operating the tool. You need to instrument the application, and you need to decide what a bad response looks like before an evaluator can find one.

Strengths

  • OpenTelemetry based instrumentation avoids a proprietary agent and slots into existing observability pipelines.
  • Trace trees make agent and retrieval failures diagnosable at the step that caused them.
  • Dataset and experiment workflow turns prompt changes into measured comparisons rather than vibes.
  • A usable open source path exists for teams that cannot send prompt data to a vendor.

Limitations

  • This is an observability platform first. Security detection is something you configure with evaluators, not a hardened control, and it does not block anything in line by itself.
  • LLM as judge evaluation costs money per span and carries its own error rate, so evaluating full production volume is usually impractical and sampling hides rare events.
  • Traces contain prompts and retrieved context, which means they contain whatever sensitive data your application handles. Retention and access control become a real obligation.

Who it suits

Teams operating LLM features at enough volume that behavior must be measured rather than assumed, with engineering capacity to instrument and to write evaluators. Not a substitute for a runtime guardrail if your requirement is blocking bad input or output rather than observing it.

Used Arize AI? Recommend it under your own name and title.

Recommend this tool