What we still need to verify : 3 points in this profile are not yet confirmed against vendor documentation.
- Current product names for the commercial platform and guardrail components: verify against vendor docs
- Which LLM safety metrics ship built in versus require configuration: confirm
- Self hosted deployment options for the commercial platform: confirm
Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.
What it does
WhyLabs is built around a distinctive idea: instead of shipping raw data to a monitoring backend, you compute a statistical profile locally and ship that. The open source logging library summarizes a batch or stream into sketches, distributions, cardinality estimates, null rates and type counts, producing a compact object that describes the data without containing it. Profiles merge across time windows and distributed workers, so monitoring a large pipeline needs no proportionally large data path, and the monitoring system never holds your records.
The platform consumes those profiles and detects change: input drift, prediction shift, data quality regressions, segment level anomalies an aggregate would hide. For language model workloads the same machinery is applied to text, with metrics over prompts and responses covering similarity to known injection patterns, sentiment and toxicity, readability, refusal detection and pattern matched sensitive data. Those become monitored series like any other, so you alert on the rate of injection-like inputs rather than triaging individual requests. A control oriented layer evaluates the same metrics in line, so a request can be blocked or a response rewritten rather than only recorded.
Where it fits
Production, alongside the serving path, owned by the ML platform or AI engineering team. Security consumes the output rather than operating the tool, which is the workable arrangement in this category. You need to instrument the points where data enters and predictions leave, and a reference window representing normal behavior before drift means anything. It does not evaluate a model you have not deployed.
Strengths
- Profile based telemetry keeps raw data inside your boundary, removing the biggest objection to sending production traffic to a monitoring vendor.
- Mergeable profiles scale to high volume and distributed pipelines without sampling away the tail.
- Covers classical ML and LLM workloads with one monitoring model, useful for teams running both.
- The logging library is genuinely open source and usable on its own.
Limitations
- Statistical monitoring detects aggregate change. A single targeted attack that does not shift a distribution is invisible to it.
- The text safety metrics are heuristic and similarity based, so they approximate injection and toxicity rather than reliably detecting them, and they are weaker outside English.
- Profiles are summaries by design, so an alert often cannot be traced to the specific records that caused it, and baselines need tending: seasonal traffic produces drift alerts that are correct and useless.
Who it suits
Teams running models in production at volume, especially where data governance rules out sending raw inputs to a vendor. Not the right tool for pre release adversarial testing, or for a team whose only AI usage is a low volume feature calling a hosted API.
Used WhyLabs? Recommend it under your own name and title.
Recommend this tool