Skip to content

Repository files navigation

dspm

DSPM scanner: discovers sensitive data (PII, credentials & secrets, financial, healthcare, regional-compliance identifiers) in cloud data stores and posts findings to the CSPM backend.

Supported connectors: S3, Azure Blob Storage / ADLS Gen2, PostgreSQL, MySQL, MariaDB, MSSQL, MongoDB/DocumentDB, DynamoDB, RDS/Aurora, Azure Database for PostgreSQL / MySQL, Azure SQL, Cosmos DB (NoSQL and MongoDB APIs), Google Workspace (Drive), Salesforce.

Setup

python -m venv venv && ./venv/bin/pip install -r requirements.txt

Configuration is read from environment variables (a .env file in the project root is loaded automatically by settings.py). .env.example lists every variable with the recommended production values — cp .env.example .env and fill in the targets and credentials.

Running

There are two entry points:

  1. Worker (src/dspm_scanner_worker_handler.py) — scans whole resources (buckets and/or databases) configured via environment variables; several targets are scanned two at a time. This is what the Docker image runs.

    python -m src.dspm_scanner_worker_handler

    Findings are written to output/findings/<OBJECT_NAME>-<YYYY-MM-DD>.json (one file per target) and uploaded as a zip archive to CSPM_URL if configured. The JSON has the same layout for buckets and databases: findings holds one entry per scanned object key, schema.table or collection (an empty list when it is clean), and files_scanned counts them. With KEEP_SCANNED_FILES=true the files the scanner downloaded or exported (S3 objects, Azure blobs, Drive files, Salesforce attachments) are kept under output/scanned/ as well, to check what the parsers actually saw; leave it off anywhere but a test run.

  2. Master (src/dspm_scanner_master_handler.py) — AWS Lambda handler that scans one target per invocation payload (also accepts SQS-wrapped payloads, S3 event notifications, and DynamoDB Stream batches).


Worker mode — environment variables per connector

Common (all connectors)

Variable Required Description
OBJECT_TYPE yes* Selects the connector, see sections below
OBJECT_NAME yes* S3 bucket name, DynamoDB table name, Azure Blob container, or database name for the DB connectors
OBJECTS_TO_SCAN no Several targets at once: a JSON object {"name": "type", ...} (e.g. {"bucket-a": "s3", "appdb": "postgres"}) or a JSON list of names that all use OBJECT_TYPE. Overrides OBJECT_NAME/OBJECT_TYPE (* not needed when set)
CSPM_URL no CSPM backend base URL; findings upload is skipped when unset
ARTIFACT_TOKEN with CSPM_URL Bearer token for the findings upload (api/v1/artifact/)
LABEL_ID no Label the uploaded findings are filed under in the CSPM backend, default test
OBJECT_REGION no AWS region for the S3 and DynamoDB clients (applies to every S3 and DynamoDB target)
ENABLED_REGIONS no Comma-separated regional compliance packs, default US,IN,GB (valid: US, CA, GB, DE, SE, FI, PL, ES, IT, TR, IN, SG, AU, KR, TH, ZA, NG, PH; UK is accepted as an alias for GB). Also used as the regions for national-format phone numbers
REPORT_TOKEN_LIKE_VALUES no false (default): random-looking tokens with no supporting evidence (credential-named field, key=/token: keyword, known format) are dropped; true keeps them as possible candidates reported as Secret.TokenLikeValue (Medium) when a whole column is made of them
MIN_CONFIDENCE no Lowest confidence tier reported: possible, likely (default) or very_likely (see Classification below). The legacy SCORE_THRESHOLD float is still accepted (0.9 → very_likely, 0.8 → likely, lower → possible)
ADAPTIVE_SAMPLING no false (default). true stops reading a table/collection once its column verdicts have settled (see Sampling)
SAMPLE_STRATEGY no head (default): the first SAMPLE_LIMIT rows/documents. random: TABLESAMPLE on PostgreSQL/MSSQL tables larger than twice the limit, $sample on MongoDB, head elsewhere (see Sampling)
SAMPLE_LIMIT no Rows / documents read per table or collection, default 10000
NER_ENABLED no true (default): person names in free text through spaCy; false skips the model
NER_MODEL no en_core_web_trf (transformer, best accuracy — the default when installed) or en_core_web_sm (12 MB, ~40× faster per cell); both ship in requirements.txt
REPORT_PRIVATE_IPS no false (default): RFC 1918 / loopback / link-local addresses are infrastructure, not PII.IPAddress
DISABLED_DETECTORS no Detector names never reported (comma-separated or JSON list), e.g. PII.IPAddress,MAC_ADDRESS
ALLOW_LIST / ALLOW_REGEX no Macie-style exceptions: exact values (comma-separated or JSON list) / JSON list of regexes over values that are never findings
COLUMN_RATIO / MIN_COUNT no Override every detector's column-classification share (policy default 0.5) / distinct-possible-hits promotion count (policy default 10); empty keeps the per-detector policies
AGGREGATION_THRESHOLD no Hits per (detector, column) that collapse into one column-level finding, default 25; 0 disables
OUTPUT_DIR no Findings/work directory. Default <repo>/output; the container image sets /app/output — point it at a mounted volume to persist findings
KEEP_SCANNED_FILES no true keeps every file a connector downloaded or exported under <OUTPUT_DIR>/scanned/, mirroring the source (s3/<bucket>/<key>, gdrive/<user or drive>/<file id>/<name>, salesforce/<host>/<object>/<id>/<name>), instead of deleting it after classification. Local testing only: it duplicates the sensitive data on disk

S3

Variable Required Description
OBJECT_TYPE yes S3
OBJECT_NAME yes Bucket name
AWS_ACCOUNT_ID yes Account that owns the bucket(s); recorded in the findings and required by the CSPM backend
AWS_ACCESS_KEY_ID no Static IAM credentials with s3:ListBucket + s3:GetObject. Leave both unset to use the instance profile / IRSA, or AWS_PROFILE + AWS_CONFIG_FILE for an assume-role profile into another account (see deployments/vm/README.md)
AWS_SECRET_ACCESS_KEY no with AWS_ACCESS_KEY_ID

Objects larger than 100 MB are skipped. Archives (.zip/.tar/.gz/.bz2) are unpacked and scanned recursively; CSV/TSV, Parquet, Excel, JSON, XML, PDF, DOCX and images (OCR) have dedicated parsers, everything else falls back to plain-text scanning.

Azure Blob Storage / ADLS Gen2

Variable Required Description
OBJECT_TYPE yes AZURE_BLOB (or ADLS, AZURE_STORAGE)
OBJECT_NAME yes Container name; account/container or the container URL when the identity reads several accounts
AZURE_STORAGE_ACCOUNT yes* Storage account of the container(s): the account name or its https:// URL (* unless every name carries its account)
AZURE_SUBSCRIPTION_ID yes Subscription that owns the account; recorded in the findings as account_id and required by the CSPM backend
AZURE_STORAGE_ENDPOINT_SUFFIX no core.windows.net (default), core.usgovcloudapi.net, core.chinacloudapi.cn
AZURE_CLIENT_ID, AZURE_TENANT_ID, AZURE_CLIENT_SECRET no A service principal (app registration). Leave unset to use the VM's managed identity or AKS workload identity; a bare AZURE_CLIENT_ID selects a user-assigned managed identity. Creating the registration and the roles it needs: deployments/vm/azure/README.md
AZURE_STORAGE_SAS_TOKEN / AZURE_STORAGE_CONNECTION_STRING / AZURE_STORAGE_ACCOUNT_KEY no Credential fallbacks when no identity is available, tried in that order before the identity chain; a read + list SAS is the safe one

The identity needs the role Storage Blob Data Reader on the account (or its resource group) and a network path: the storage firewall allowing the scanner's subnet, or a private endpoint (deployments/vm/azure/). The S3 rules apply: blobs over 100 MB are skipped, archives are unpacked and scanned recursively, workbooks yield one entry per sheet. Directory placeholders of hierarchical (ADLS Gen2) accounts, archive-tier blobs (they need rehydration first), page blobs and soft-deleted blobs are skipped and counted in the scanner stats. Resource ids are blob URLs (https://<account>.blob.core.windows.net/<container>/<blob>) and never carry a SAS token; a blob that could not be downloaded is reported in errors, not recorded as clean.

PostgreSQL / MySQL / MariaDB / MSSQL

Variable Required Description
OBJECT_TYPE yes POSTGRES (or POSTGRESQL), MYSQL, MARIADB, MSSQL (or SQLSERVER)
OBJECT_NAME yes Database name to scan
DB_HOST yes* Database host
DB_PORT yes* Typical defaults: 5432 (postgres), 3306 (mysql/mariadb), 1433 (mssql)
DB_USERNAME yes* A read-only account is sufficient and recommended
DB_PASSWORD yes*
DB_URI no Full SQLAlchemy connection string, e.g. postgresql+psycopg2://user:pass@host:5432/db. Overrides all DB_* fields above (* not needed when DB_URI is set)

Drivers used: psycopg2 (postgres), PyMySQL (mysql/mariadb), pymssql (mssql). All non-system schemas of the database are discovered and scanned, up to 10 000 rows per table.

MongoDB / DocumentDB

Variable Required Description
OBJECT_TYPE yes MONGODB (or MONGO, DOCUMENTDB)
OBJECT_NAME yes Database name to scan
DB_HOST yes* MongoDB host
DB_PORT yes* Typically 27017
DB_USERNAME no* Omit for unauthenticated instances
DB_PASSWORD no*
DB_URI no Full MongoDB URI, e.g. mongodb://user:pass@host:27017/?authSource=admin. Overrides all DB_* fields above (* not needed when DB_URI is set)

All non-system.* collections of the database are discovered and scanned, up to 10 000 documents per collection. Documents are walked recursively; nested fields are reported with dotted paths.

Reaching a replica set through kubectl port-forward / an SSH tunnel: the members advertise cluster-internal hostnames (…rs0-0.…svc.cluster.local) that do not resolve locally, so topology discovery fails with Could not reach any servers. The scanner adds directConnection=true automatically when DB_HOST is localhost/127.0.0.1; with DB_URI, append ?directConnection=true yourself.

DynamoDB

Variable Required Description
OBJECT_TYPE yes DYNAMODB (or DDB)
OBJECT_NAME yes Table name; OBJECTS_TO_SCAN lists several tables of the same account and region
OBJECT_REGION yes* Region of the tables; the DynamoDB client is regional (* or AWS_DEFAULT_REGION / the profile's region)
AWS_ACCOUNT_ID yes Account that owns the tables; recorded in the findings and required by the CSPM backend
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY no Static keys, exactly as for S3; leave unset for the instance profile / IRSA, or AWS_PROFILE for an assume-role profile

One table is one unit, like a relation or a collection: a Scan reads up to SAMPLE_LIMIT items (10 000), which counts against the table's read capacity, and every item is classified like a MongoDB document (attribute paths as context, nested maps and lists walked). The identity needs dynamodb:Scan on the tables (dynamodb:DescribeTable for a preflight); deployments/vm/aws/member-role-policy.json grants both for one region. Change-data capture through DynamoDB Streams remains a master-mode feature.

Azure Database for PostgreSQL / MySQL, Azure SQL, Cosmos DB for MongoDB

Ordinary wire-protocol databases to the scanner: the SQL and MongoDB connectors above behind an Azure OBJECT_TYPE alias, so the CSPM can tell the platform apart. Same DB_* variables, one instance per server.

Service OBJECT_TYPE Connection
Azure Database for PostgreSQL Flexible Server AZURE_POSTGRES DB_HOST=<server>.postgres.database.azure.com, port 5432. psycopg2 negotiates the TLS the server requires; pin the CA with DB_URI=postgresql+psycopg2://…?sslmode=verify-full&sslrootcert=/etc/ssl/certs/ca-certificates.crt
Azure Database for MySQL Flexible Server AZURE_MYSQL Port 3306. PyMySQL negotiates TLS by itself; for *.mysql.database.azure.com hosts the scanner also verifies the server certificate against the system CA bundle (a DB_URI with its own ssl_ca= is left alone)
Azure SQL Database / Managed Instance AZURE_SQL DB_HOST=<server>.database.windows.net, port 1433, a contained user in db_datareader. SQL logins only: pymssql cannot present Entra tokens
Cosmos DB for MongoDB, RU or vCore COSMOS_MONGO DB_URI = the account's read-only connection string (RU, port 10255) or the vCore mongodb+srv:// URI. Request-rate throttling (error 16500) is retried from the last document read, honouring the server's RetryAfterMs

Without passwords. DB_AUTH=azure_entra makes the scanner identity's Microsoft Entra access token the password of every connection to a Flexible Server (PostgreSQL and MySQL). DB_USERNAME is the identity's name on the server, created by the server's Entra administrator, on PostgreSQL: SELECT * FROM pgaadauth_create_principal('dspm-scanner-vm', false, false); GRANT pg_read_all_data TO "dspm-scanner-vm";. Tokens last about an hour and are fetched per connection, so long scans keep working.

Cosmos DB for NoSQL

Variable Required Description
OBJECT_TYPE yes COSMOS_NOSQL (or COSMOS, AZURE_COSMOS)
OBJECT_NAME yes Database name to scan
AZURE_COSMOS_ENDPOINT yes https://<account>.documents.azure.com:443/, or just the account name
AZURE_COSMOS_KEY no The account's read-only key. Leave unset to use the identity chain with the data-plane role Cosmos DB Built-in Data Reader, assigned with az cosmosdb sql role assignment create (the portal's IAM blade cannot grant data-plane roles)
AZURE_SUBSCRIPTION_ID no Recorded in the findings as account_id

Every container of the database is scanned, up to SAMPLE_LIMIT items each (SELECT TOP n * FROM c); the Cosmos system properties _rid, _self, _etag, _attachments, _ts are dropped before classification and items are walked recursively with dotted paths like MongoDB documents. The NoSQL API has no server-side random sample, so SAMPLE_STRATEGY=random reads the head. Request-rate limits (429) are retried by the SDK. Resource ids are https://<account>.documents.azure.com/<database>/<container>.

Google Workspace (Drive)

Variable Required Description
OBJECT_TYPE yes GOOGLE_WORKSPACE (or GDRIVE, GOOGLE_DRIVE, GOOGLEWORKSPACE)
OBJECT_NAME yes What to scan when the GOOGLE_* variables below are unset: a user email (their My Drive) or a shared-drive id
GOOGLE_SA_KEY_FILE yes* Path to the service-account key JSON (* or set GOOGLE_APPLICATION_CREDENTIALS; Application Default Credentials when neither is set)
GOOGLE_IMPERSONATE_USER no Workspace user whose My Drive is scanned (requires domain-wide delegation)
GOOGLE_DRIVE_ID no One shared drive to scan instead

Setup (the model the DSPM vendors document): create a GCP service account, enable the Drive API, and authorize the account in the Google Admin console for domain-wide delegation with the single read-only scope https://www.googleapis.com/auth/drive.readonly - the scanner cannot modify data by construction. Google-native files are exported (Docs → .docx, Sheets → .xlsx so every sheet is scanned - the CSV export is first-sheet-only - Slides → text; the Drive API caps exports at 10 MB), everything else is downloaded as-is; both go through the same file parsers as S3 objects, one Drive file per findings entry. Files over 100 MB are skipped.

Salesforce

Variable Required Description
OBJECT_TYPE yes SALESFORCE (or SFDC)
OBJECT_NAME yes My Domain name (acme for acme.my.salesforce.com) or the full instance URL; SF_DOMAIN overrides it
SF_CONSUMER_KEY yes Connected App credentials (OAuth 2.0 client-credentials flow)
SF_CONSUMER_SECRET yes
SF_OBJECTS no Pin the sObjects to scan (comma-separated or JSON list); default: every queryable business object that holds records
SF_INCLUDE_FILES no true (default): also scan the files attached to records - the latest ContentVersion of every File plus classic Attachment bodies - through the file parsers
SF_API_VERSION no Default v62.0

Setup: a Connected App with the client-credentials flow enabled, run as an integration user holding a permission set with "API Enabled", "View All Data" and "Query All Files". One sObject is one unit, scanned like a table: text-typed fields only (string / textarea / email / phone / url / picklist), up to SAMPLE_LIMIT records per object. *Share / *History / *Feed / *ChangeEvent, Apex* / Datacloud* and other system objects are excluded, and empty objects are skipped via limits/recordCount. Incremental scans filter on SystemModstamp.

Example .env (PostgreSQL)

OBJECT_TYPE=POSTGRES
OBJECT_NAME=appdb
DB_HOST=127.0.0.1
DB_PORT=5432
DB_USERNAME=scanner
DB_PASSWORD=secret
CSPM_URL=https://cspm.example.com/
ARTIFACT_TOKEN=eyJ...

Master mode — payload fields per connector

Invocation payload shape:

{
  "scan_type": "...",
  "target": { ... },
  "config": { "enabled_regions": ["US", "IN"] }
}

ARTIFACT_TOKEN + CSPM_URL environment variables control the findings upload in this mode.

scan_type: "s3"

Target field Required Description
bucket yes Bucket name
key yes Object key
version_id no Specific object version
last_modified no With config.last_scan_time, enables skip-if-unchanged

scan_type: "postgres" | "postgresql" | "mysql" | "mariadb" | "mssql" | "sqlserver" | "azure_postgres" | "azure_mysql" | "azure_sql"

Target field Required Description
host, port yes* Database endpoint
username, password yes* Credentials
database yes* Database name
connection_string no Full SQLAlchemy DSN; overrides the fields above (*)
password_secret no AWS Secrets Manager ARN/name, or an Azure Key Vault secret URI (https://<vault>.vault.azure.net/secrets/<name>, read with the scanner identity); fills any missing username/password/host/port/database/uri
auth no azure_entra: the scanner identity's token is the password (Azure Database for PostgreSQL / MySQL Flexible Server)
schema no Restrict to one schema (default: all non-system schemas)
tables no Restrict to specific tables, plain or schema-qualified: ["users", "sales.orders"]
include_views no Also scan views (default false)
incremental_column no Timestamp column for incremental scans
last_scan_time no Only rows where incremental_column > last_scan_time are scanned
sample_limit no Max rows per table (default 10000)

scan_type: "rds" | "aurora"

Same fields as the SQL engines above (set engine to one of postgres/mysql/mariadb/mssql), plus:

Target field Required Description
engine yes Database engine of the RDS instance
reader_endpoint no Aurora reader endpoint
use_reader no Route the scan to reader_endpoint

scan_type: "mongo" | "mongodb" | "documentdb" | "cosmos_mongo"

Target field Required Description
host, port yes* MongoDB endpoint (port defaults to 27017)
username, password no* Omit for unauthenticated instances
uri no Full MongoDB URI; overrides the fields above (*)
password_secret no AWS Secrets Manager ARN/name or Key Vault secret URI, as for SQL
database no Restrict to one database (default: all non-system databases)
collection no Restrict to one collection (default: all non-system.* collections)
incremental_field no Field for incremental scans, with last_scan_time
last_scan_time no Only documents where incremental_field > last_scan_time
sample_limit no Max documents per collection (default 10000)

scan_type: "dynamodb"

Target field Required Description
table_name yes DynamoDB table name
region no AWS region (default from environment)
sample_limit no Max items (default 10000)

Uses ambient AWS credentials (Lambda role / environment). DynamoDB Stream CDC batches are handled automatically when the Lambda is wired to a stream.

scan_type: "azure_blob" | "adls"

Target field Required Description
account yes* Storage account name or URL (* default AZURE_STORAGE_ACCOUNT)
container yes Container name
blob no One blob; omit to scan the container (prefix restricts it to a virtual directory)
version_id, snapshot no With blob
last_scan_time no Skip blobs not modified since (ISO 8601)
max_blobs no Cap on blobs listed per run
connection_string, sas_token, account_key, endpoint_suffix no Credential / endpoint overrides; password_secret may supply them from Key Vault or Secrets Manager

Without overrides the identity chain is used (service principal from AZURE_CLIENT_ID / AZURE_TENANT_ID / AZURE_CLIENT_SECRET, workload identity, managed identity), as in worker mode.

scan_type: "cosmos" | "cosmos_nosql"

Target field Required Description
endpoint or account yes* Account URL or name (* default AZURE_COSMOS_ENDPOINT)
key no Read-only key; omit for the identity chain (default AZURE_COSMOS_KEY)
database, container no Restrict to one database / container (default: all)
last_scan_time no Only items whose _ts is after it
sample_limit no Max items per container (default 10000)
password_secret no Key Vault secret URI or Secrets Manager ARN supplying key / endpoint

scan_type: "google_workspace" | "gdrive" | "google_drive"

{
  "scan_type": "google_workspace",
  "target": {
    "sa_key_file": "/path/key.json",
    "impersonate_user": "user@example.com",
    "drive_id": "0AbCd...",
    "folder_id": "1XyZ...",
    "max_files": 500,
    "last_scan_time": "2026-08-01T00:00:00Z"
  }
}

impersonate_user (My Drive via domain-wide delegation) or drive_id (a shared drive) selects the corpus; folder_id restricts to one folder; omit sa_key_file to use Application Default Credentials.

scan_type: "salesforce" | "sfdc"

{
  "scan_type": "salesforce",
  "target": {
    "domain": "acme",
    "consumer_key": "3MVG9...",
    "consumer_secret": "...",
    "objects": ["Contact", "Lead"],
    "include_files": true,
    "sample_limit": 10000,
    "last_scan_time": "2026-08-01T00:00:00Z"
  }
}

Like the database targets, password_secret (an AWS Secrets Manager ARN or an Azure Key Vault secret URI) can supply consumer_key / consumer_secret / domain - or a ready access_token + instance_url - instead of inline values.

config keys (all scan types)

Key Default Description
enabled_regions [] Regional compliance packs (ISO alpha-2), 62 packs: AE, AR, AT, AU, BE, BG, BR, CA, CH, CL, CN, CZ, DE, DK, EE, EG, ES, FI, FR, GB, GH, GR, HK, HR, HU, ID, IE, IL, IN, IS, IT, JP, KR, LK, LT, LU, LV, MX, MY, NG, NL, NO, NZ, PH, PK, PL, PT, RO, RS, RU, SA, SE, SG, SI, SK, TH, TR, TW, UA, US, VN, ZA — national ids, tax numbers, passports, driver licences, health identifiers with their public check-digit algorithms
phone_regions enabled_regions Regions used to parse national-format phone numbers (5678942315 in a mobile field); international +… numbers are always detected
chunk_size 5000 (SQL) / 1000 (Mongo) Rows/documents fetched per batch
connect_timeout 10 Connection timeout in seconds (SQL and Mongo)
log_queries false (worker mode: always on) Log every query issued during DB scans (dialect-compiled SQL with bound values, Mongo filters, DynamoDB scans). Note: emits table/column names into logs
last_scan_time – S3 only: skip objects not modified since this timestamp
keep_files_dir – Keep every downloaded/exported file under this directory, mirroring the source path, instead of deleting it after classification (the worker sets <OUTPUT_DIR>/scanned when KEEP_SCANNED_FILES=true). Local testing only
aggregation_threshold 25 A (detector, column) pair firing on at least this many rows/documents/cells collapses into one column-level finding with an occurrences count. 0 disables
min_confidence likely Lowest confidence tier reported: possible / likely / very_likely (a legacy score_threshold float is accepted)
column_ratio per detector (0.5) Share of a column's sampled non-empty values that must match before the column is classified (Sentra's 50 % rule); overrides every detector policy
column_min_matches per detector (3) Distinct matching values needed before a column verdict is drawn
min_count per detector (10) Distinct possible hits in one unit (file, table column) that promote them to likely
allow_list / allow_regex [] Macie-style allow lists: exact values / regexes never reported (public phone numbers, sample data)
adaptive_sampling false Stop reading a unit once its column verdicts settle; settle_min_records (2000), settle_window (1000), settle_margin (0.05) tune the stop rule
sample_strategy head random draws rows with TABLESAMPLE (PostgreSQL, MSSQL) / $sample (MongoDB) instead of reading the head; per-target sample_strategy overrides it
ner true Person names in prose via the spaCy model
field_suppression true Structural field-name rules: token detectors never fire in *_id/hash/etag/path/… fields, digit-run detectors never fire in counter/timestamp fields. Corroborated findings are exempt
decode_base64 true Decode base64 blobs (Authorization: Basic …, base64 JSON, PEM) and scan the plaintext
entropy_report_uncorroborated false Report random-looking tokens that have no supporting evidence as Secret.TokenLikeValue (Medium) instead of dropping them
disabled_detectors [] Detector names never reported
report_private_ips false Report RFC 1918 / loopback / link-local addresses as PII.IPAddress; by default only public IPs count
direct_connection auto MongoDB: add directConnection=true to the built URI. Automatic for localhost/127.0.0.1, i.e. port-forwarded replica sets whose members advertise cluster-internal hostnames
column_suppression id/hash rule for entropy Scanner-level escape hatch: per-detector regexes of column/field names to skip on top of the engine's own field rules. Pass {} to disable, or your own {detector: regex} map
entropy_min_length 24 Minimum token length for the entropy detector
entropy_min_entropy 4.5 Shannon-entropy threshold for base64-shaped tokens (hex tokens use 3.0 and need 32+ chars)

Classification

connector  ──►  Record / TextBlob stream  ──►  src/pipeline (per unit)  ──►  findings
                                                    │
                                              src/engine (per value)

The code is split the way the vendor engines we studied are (Wiz, Cyera, Orca, Sentra, Varonis, Amazon Macie, Google Sensitive Data Protection, Microsoft Purview, Nightfall — see How the vendors do it below):

  • Connectors (src/scanners/) know a data source. They enumerate its units (tables, collections, objects, sheets) and turn each unit into a stream of Records (rows / documents made of Cells that carry the value, its column name or field path, and a rendered location) or TextBlobs (pages, paragraphs, text blocks). They never call the detection engine.
  • The engine (src/engine/) judges one value at a time: pattern + validator + context words + field name → a score and a confidence tier. It stays the place where detectors are defined.
  • The pipeline (src/pipeline/) does everything a DSPM product does after pattern matching, once, for every connector: context policy, record-level corroboration, column density verdicts, minimum counts, adaptive sampling and aggregation.

Confidence tiers

Every finding carries confidence and evidence:

Confidence tiers express heuristic evidence strength, not calibrated probabilities of correctness.

Tier Meaning Examples
very_likely validated shape and corroboration checksum-valid national id next to its keyword or in a column named for it, JWT with a decodable header, vendor-prefixed token, opaque value in a credential-named column, a classified column of likely cells
likely one strong signal valid e-mail (IANA TLD, no demo/automated sender), Luhn + issuer prefix with separators, mod-97 IBAN, structured street address, a plausible shape backed by its column name or a context word, a classified column of possible cells
possible plausible shape only SSN-shaped digit groups, a bare mod-10/mod-11-valid number, a random-looking token, a BIC-shaped word. Never reported on its own

MIN_CONFIDENCE picks the lowest tier reported (likely by default — Google's default is possible, Nightfall recommends likely). Checksums are not enough on their own: a mod-10/mod-11 check passes ~10 % of random numbers, so a checksum-valid number with no context word, field hint or column evidence is possible (a US phone number is a valid NHS number one time in ten). Epoch timestamps are never identifiers; private/loopback IPs are infrastructure (report_private_ips); documented examples (AKIAIOSFODNN7EXAMPLE, test cards, hunter2 in advisory text) are never real.

evidence lists what backs the finding: checksum, format (self-validating shape), context:<word>, field (the column/path names the entity), key:<name> (assigned to a credential keyword), column:<ratio> (column density), record:identity (same record holds two identity signals), count:<n> (many in one file), shape / needs_context / uncorroborated (why something stayed possible).

Detector policies

Each detector has a reporting policy (src/engine/policy.py, looked up by name with category defaults, so a new recognizer needs no entry unless it deviates):

Field Meaning Vendor precedent
context required: a hit with neither validation nor context is capped at possible (national ids, bank accounts, cards, SWIFT, opaque tokens); boost: context raises the tier; none: self-identifying (e-mail, JWT, IBAN, vendor tokens) Macie keyword requirements per data type
column_ratio, column_min_matches share of a column's sampled values (and distinct matches) that classify the column Sentra 50 % rule, Google column profiles
min_count, count_promotion distinct possible hits in one unit that become likely; off for word-shaped patterns (SWIFT/BIC) and random strings Orca statistical scan, Purview "low-confidence patterns with 20+ instances", Nightfall minimum findings
identity, identity_corroboration identity signals (name, e-mail, phone, address, DOB) promote a possible national id / card in the same record Purview supporting elements, Cyera identifiability
negative_fields column names that veto the detector (txn, hash, invoice, port, amount …) DLP negative keywords, Google exclude-by-hotword

Field-name rules

Structured data carries its strongest signal in the field name (DetectionEngine.scan_text(text, field_name=...); the name is never mixed into the scanned text). Macie counts a keyword in the column name or any element of the JSON path as proximity, and so does the engine: src/engine/context.py classifies credential-named fields (token, secret_key, authorization, cookie, webhook_url, …) whose values are credentials whatever their entropy, identifier fields (*_id, hash, etag, path, …) that never yield token findings, counter/timestamp fields that never yield card/id numbers, full_name/first_name fields that yield PII.PersonName, and mobile/phone fields that enable national-format phone parsing (a phone-named column whose numbers do not parse for the enabled regions still yields possible phones that column density can promote). Overlapping matches keep the most specific detector (a JWT is not also a bearer token, a card number is not also a bank account).

Names in prose. A name in a name-labelled column is a field rule; a name inside a support note, a PDF page or a comment field needs a model (Macie NAME, Purview named entities, Cyera NER). With spaCy (requirements.txt ships en_core_web_trf, a RoBERTa transformer with OntoNotes NER F1 0.90, and en_core_web_sm, F1 0.84, as the fallback; NER_MODEL chooses, NER_ENABLED=false skips), src/engine/ner.py accepts PERSON entities that look like written names — two to four title-case tokens, no digits or acronyms, no company suffix — and entities the small model mislabels when their first token is in the shipped given-name lexicon (src/engine/data/given_names.txt). A name next to a context word or honorific (patient, customer, employee, regards, dear, Mr/Dr …) is likely; otherwise it is possible and the pipeline decides: a document with ten distinct names is likely, a lone name in a sentence stays hidden.

Column, record and file verdicts (src/pipeline)

  • Column density — a column whose sampled non-empty cells match a detector at column_ratio (default 50 %) with at least column_min_matches matching cells and distinct values is classified as that detector one tier above its cells. Each cell contributes once, even if it contains several matches; validation density and the majority confidence tier also count cells. Forty SSN-shaped values under a meaningless header can classify a column, while three shapes in one of six cells represent 1/6 density, not 3/6.
  • Unit-name context — the table, collection, sheet or object name is part of the path to a value (Macie counts a keyword "in the name of an element in the path"): a possible card number in credit_cards.number or an SSN-shaped value in ssn_export.csv is likely (unit:<name>). Only possible candidates are lifted; very weak patterns still need their column name.
  • Sibling columns — companion names can raise an established column verdict one more tier: expiry/cvv next to a card column, for example (siblings:<columns>). They must share the same parent field path. A sibling name elsewhere in the unit cannot promote an isolated weak match on its own.
  • Record corroboration — eligible possible identifiers next to two identity signals become likely (record:identity). For documents, those signals must belong to the same containing object; separate array members remain separate contexts. Bare-number and weak-checksum candidates marked needs_context require stronger support, such as a detector-specific field/keyword or sufficient column density; a nearby name and email alone do not establish their type.
  • Column exclusivity — the dominant detector suppresses weak alternative interpretations, such as a phone number that happens to pass an NHS checksum. Independently supported findings in separate spans or cells remain reportable: a notes field can contain both email addresses and phone numbers, and an email-heavy column can contain an explicitly labelled SSN. Record or column promotion alone does not establish independent support.
  • Minimum counts — in a document or a mixed column, min_count distinct possible hits of one detector become likely (count:<n>): a file with 30 SSN-shaped numbers is not a coincidence, one is.
  • Aggregation — a (detector, column) pair with aggregation_threshold or more hits collapses into one column-level finding. occurrences counts findings; column_matches and column_validated_matches count matching cells and cells with validation evidence. column_sampled is the non-empty cell denominator for column_ratio and column_validated_ratio. Validation evidence includes checksums and recognized formats; it does not prove an identity or credential is real. These details are available in pipeline findings; the worker's grouped output currently retains counts and confidence but omits the column statistics.
  • Allow lists — allow_list / allow_regex suppress known values (public contact numbers, sample data), like Macie allow lists.

Sampling

Connectors read up to sample_limit rows/documents per unit (10 000). By default that is the head of the table; SAMPLE_STRATEGY=random (or sample_strategy per target) draws a random sample instead — TABLESAMPLE SYSTEM (p) on PostgreSQL and MSSQL when the planner estimate says the table holds more than twice the limit (p is sized to about three times the limit before LIMIT cuts it), $sample on MongoDB, the head on MySQL/MariaDB where no cheap random read exists. With adaptive_sampling the pipeline stops reading a unit once settle_min_records records have been seen, no new (column, detector) pair appeared for settle_window records and no column sits within settle_margin of its classification ratio. Rows already read count in the stats.

The current adaptive stop rule is a stability heuristic, not a statistical confidence bound. A stable sample, especially a head or page sample, does not guarantee discovery of rare sensitive rows or prove the unread portion is clean. For detectors that require validation density, the stop rule uses the validated cell ratio, as column classification does.

Output

Every finding carries resource_id, detector, category, severity, value (capped at 200 characters), location, confidence, evidence and a value_hash (sha256 prefix for correlating the same value across scans). Column-level findings add the aggregation fields above. The worker's clubbed JSON entries carry the highest confidence among their findings.

Recognizer packs. src/engine/recognizers/ holds 159 country-specific and generic recognizers across 62 region packs as native Rule objects (src/engine/rules.py: pattern scores, context words, validators/invalidators; test vectors in tests/test_recognizers_*.py). Rules are grouped by region pack; generic ones (IBAN, crypto wallets, IP, MAC, IMEI, ICCID, VIN, passport MRZ, coordinates, ICD-10 / NDC codes, medical record numbers) always run, URL/UUID are shipped disabled. Validators return True (checksum holds), False (dropped) or None where the algorithm is not authoritative for every number (Danish CPR after 2007, Latvian 32-prefixed codes, UK UTR, Mexican RFC). Every detector name must exist in fixtures/findings-mapping.json (tests/test_detector_names.py).

Regression corpus. tests/fixtures/detection_corpus.json is an anonymised corpus built from real Postgres/Mongo scans: 370+ reviewed false positives that must stay silent and 160+ true positives that must stay detected (tests/test_regression_corpus.py). tests/test_pipeline.py covers the column/record/file rules and the connector contract.

The original corpus checks selected expected/forbidden detectors; its reported precision and recall are targeted regression measures. The additional fully labelled synthetic corpus, tests/fixtures/accuracy_cases.json, also counts unexpected detections on positive cases and checks exact values and cell/blob locations. It covers mixed content, sparse weak matches and encoded payloads. Run it with:

python -m scripts.evaluate_accuracy --check --output /tmp/dspm-accuracy.json

The evaluator disables aggregation, reports per-detector precision/recall/F1, and writes only case identifiers and counts, not detected values. Undefined metrics are null. These development cases are not a held-out production benchmark. See the competitive assessment and accuracy review for sources, measured changes and the next engineering priorities.

The current tuning priority is reducing false alarms. The additional tests/fixtures/false_alarm_cases.json corpus covers weak numeric references beside identity fields, context crossing between nested objects or array members, and positive controls. Run it with python -m scripts.evaluate_accuracy --corpus tests/fixtures/false_alarm_cases.json --check. The evaluator accepts flat rows, nested documents, or text cases. See false-alarm reduction for measured results and the recall tradeoff.

Sample dataset. sample_data/all_detectors.jsonl / .txt hold one synthetic example per detector (built by python -m tests.sample_dataset_builder; --check scans them and lists anything not detected) — scan them to see every finding type the engine can produce, and tests/test_sample_dataset.py keeps them in sync with the catalogue.

Adding a connector

A connector is a BaseScanner subclass (src/scanners/base.py) that implements how to reach and read a source; classification is inherited:

from src.pipeline import Cell, Record, TextBlob          # what a connector emits
from src.scanners.base import BaseScanner

class SnowflakeScanner(BaseScanner):
    def __init__(self, engine, config=None, client=None):
        super().__init__(engine, config, client)
        self.stats = {"tables_scanned": 0, "rows_scanned": 0, "errors": 0}

    def iter_scan(self, target):                          # one unit at a time, so callers can checkpoint
        for schema, table in self._list_tables(target):
            resource_id = f"snowflake://{target['account']}/{schema}.{table}"
            findings = self.classify(                     # the whole pipeline: context policy, column verdicts,
                resource_id,                              # record corroboration, min counts, aggregation
                self._rows(schema, table, target),
                location_fn=lambda column, n: f"Table '{table}', Column '{column}' ({n} matches)",
            )
            self.stats["tables_scanned"] += 1
            yield resource_id, f"{schema}.{table}", self.dedup_findings(findings)

    def scan(self, target):
        return self.collect(target)

    def _rows(self, schema, table, target):               # a generator of Records; count rows in stats
        for idx, row in enumerate(self._query(schema, table, target.get("sample_limit", 10000))):
            self.stats["rows_scanned"] += 1
            yield Record([Cell(str(v), column, f"Table '{table}', Row {idx}, Column '{column}'") for column, v in row.items() if v])

Rules of the contract:

  1. Emit Records for rows/documents (Cell.field is the column name or dotted path the engine uses as context; Cell.key the aggregation key when array indices should collapse; src.pipeline.document_record builds one from any nested document) and TextBlobs for free text (locate(start, end) renders span locations).
  2. Keep location strings connector-specific and stable; pass location_fn(column, n) for column-level findings.
  3. Count what you read in self.stats inside the generator (it keeps counting when adaptive sampling stops early) and let classify() record read errors (self.stats["errors"], error_details).
  4. Files of any origin go through src/scanners/files.iter_units(path, resource_id, config) — an object-store connector only downloads, into self.workdir(resource_id), released with self.discard_workdir() afterwards, so KEEP_SCANNED_FILES works for every connector.
  5. New detectors are Rules in src/engine/recognizers/ plus a fixtures/findings-mapping.json entry, and, only if they deviate from their category, a DetectorPolicy in src/engine/policy.py.

TLS to databases

  • SQL engines: pass DBAPI options via the target's connect_args (master mode), e.g. {"sslmode": "verify-full", "sslrootcert": "/certs/rds-ca.pem"} for PostgreSQL — or put them on the DSN: DB_URI=postgresql+psycopg2://user:pass@host:5432/db?sslmode=require. # pragma: allowlist secret
  • MongoDB / DocumentDB: use URI parameters: DB_URI=mongodb://user:pass@host:27017/?tls=true&tlsCAFile=/certs/global-bundle.pem (DocumentDB requires TLS with the Amazon CA bundle). # pragma: allowlist secret
  • Mount the CA bundle into the container (e.g. via a ConfigMap/Secret volume) and reference it by path.

Tests

python run_tests.py

About

No description, website, or topics provided.

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages