AppSecNews
Secret Scanning Open source Established

detect-secrets

by Yelp

Python secret scanner built around a baseline file, so teams can block new credentials from entering a repository without first cleaning up every historical finding.

Visit github.com (leaves AppSecNews, opens in a new tab) Leaves AppSecNews for the vendor's own site.

No endorsements yet

Run detect-secrets in production? A named recommendation helps the next team shortlisting it.

Recommend this tool

Endorsers verify their identity through LinkedIn. Titles and companies are self declared, shown as they were when each person signed, and reviewed by an editor before anything is published. Endorsements are never paid for.

What we still need to verify : 2 points in this profile are not yet confirmed against vendor documentation.
  • Which bundled plugins support live credential verification: confirm against current plugin list
  • Exact detector set changes between releases: confirm plugin inventory

Treat these points as unconfirmed. They are open items in the catalog's verification queue, and this note stays until each is checked against the vendor's documentation.

What it does

detect-secrets scans source files through a set of pluggable detectors. Some are keyword driven, looking for assignments to variables named like password or secret_key. Some are provider specific expressions for recognizable formats such as AWS access key IDs, Slack tokens, private key headers and JSON Web Tokens. The rest are entropy based, computing Shannon entropy over base64 and hex character sets and flagging strings random enough to be a credential rather than prose or an identifier.

The design decision that sets it apart is the baseline. The tool writes a .secrets.baseline file recording a hashed fingerprint of every finding at a point in time, and later scans surface only what is new. An interactive audit command walks a human through each entry to mark it a true or false positive, and that verdict persists. This turns an unbounded cleanup project into a line you can hold from today forward.

Where it fits

Most teams wire it in as a pre-commit hook, so a credential is caught before it reaches a remote, plus a CI job that re-runs the scan and fails if the baseline has drifted. Developers operate the hook, security owns the baseline policy and the audit. For that to work the baseline must be committed and treated as a reviewed artifact, and someone has to run the audit rather than regenerating the baseline whenever it complains.

Strengths

  • The baseline mechanism makes adoption possible on a legacy repository that already contains findings you cannot fix this quarter.
  • Plugins are ordinary Python classes, so adding a detector for your own internal token format is a short, testable change.
  • Filters let you exclude paths, file types and known allow-listed values with reasonable precision.
  • Fingerprints are hashed, so the baseline records that a secret exists at a location without storing the secret itself.

Limitations

  • Entropy heuristics generate real noise. Minified JavaScript, lockfiles, test fixtures and encoded binary blobs all trip the high entropy plugins, and the audit burden falls on a human.
  • It inspects the working tree rather than deep git history by default, so a secret removed from the current files but still reachable in older commits needs a different approach.
  • Regenerating the baseline is easier than auditing it, and undisciplined teams use that escape hatch until the tool stops meaning anything.

Who it suits

A Python-friendly engineering organization that wants a self-contained commit-time control and will own the baseline as a real artifact. It suits large legacy codebases better than most alternatives for that reason. Teams wanting validated findings, historical repository sweeps or centralized incident tracking should choose a managed platform instead.

Used detect-secrets? Recommend it under your own name and title.

Recommend this tool