DSPM scanner: discovers sensitive data (PII, credentials & secrets, financial, healthcare, regional-compliance identifiers) in cloud data stores and posts findings to the CSPM backend.
Supported connectors: S3, Azure Blob Storage / ADLS Gen2, PostgreSQL, MySQL, MariaDB, MSSQL, MongoDB/DocumentDB, DynamoDB, RDS/Aurora, Azure Database for PostgreSQL / MySQL, Azure SQL, Cosmos DB (NoSQL and MongoDB APIs), Google Workspace (Drive), Salesforce.
python -m venv venv && ./venv/bin/pip install -r requirements.txtConfiguration is read from environment variables (a .env file in the project root is loaded automatically by settings.py). .env.example lists every variable with the recommended production values — cp .env.example .env and fill in the targets and credentials.
There are two entry points:
-
Worker (
src/dspm_scanner_worker_handler.py) — scans whole resources (buckets and/or databases) configured via environment variables; several targets are scanned two at a time. This is what the Docker image runs.python -m src.dspm_scanner_worker_handler
Findings are written to
output/findings/<OBJECT_NAME>-<YYYY-MM-DD>.json(one file per target) and uploaded as a zip archive toCSPM_URLif configured. The JSON has the same layout for buckets and databases:findingsholds one entry per scanned object key,schema.tableor collection (an empty list when it is clean), andfiles_scannedcounts them. WithKEEP_SCANNED_FILES=truethe files the scanner downloaded or exported (S3 objects, Azure blobs, Drive files, Salesforce attachments) are kept underoutput/scanned/as well, to check what the parsers actually saw; leave it off anywhere but a test run. -
Master (
src/dspm_scanner_master_handler.py) — AWS Lambda handler that scans one target per invocation payload (also accepts SQS-wrapped payloads, S3 event notifications, and DynamoDB Stream batches).
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes* | Selects the connector, see sections below |
OBJECT_NAME |
yes* | S3 bucket name, DynamoDB table name, Azure Blob container, or database name for the DB connectors |
OBJECTS_TO_SCAN |
no | Several targets at once: a JSON object {"name": "type", ...} (e.g. {"bucket-a": "s3", "appdb": "postgres"}) or a JSON list of names that all use OBJECT_TYPE. Overrides OBJECT_NAME/OBJECT_TYPE (* not needed when set) |
CSPM_URL |
no | CSPM backend base URL; findings upload is skipped when unset |
ARTIFACT_TOKEN |
with CSPM_URL |
Bearer token for the findings upload (api/v1/artifact/) |
LABEL_ID |
no | Label the uploaded findings are filed under in the CSPM backend, default test |
OBJECT_REGION |
no | AWS region for the S3 and DynamoDB clients (applies to every S3 and DynamoDB target) |
ENABLED_REGIONS |
no | Comma-separated regional compliance packs, default US,IN,GB (valid: US, CA, GB, DE, SE, FI, PL, ES, IT, TR, IN, SG, AU, KR, TH, ZA, NG, PH; UK is accepted as an alias for GB). Also used as the regions for national-format phone numbers |
REPORT_TOKEN_LIKE_VALUES |
no | false (default): random-looking tokens with no supporting evidence (credential-named field, key=/token: keyword, known format) are dropped; true keeps them as possible candidates reported as Secret.TokenLikeValue (Medium) when a whole column is made of them |
MIN_CONFIDENCE |
no | Lowest confidence tier reported: possible, likely (default) or very_likely (see Classification below). The legacy SCORE_THRESHOLD float is still accepted (0.9 → very_likely, 0.8 → likely, lower → possible) |
ADAPTIVE_SAMPLING |
no | false (default). true stops reading a table/collection once its column verdicts have settled (see Sampling) |
SAMPLE_STRATEGY |
no | head (default): the first SAMPLE_LIMIT rows/documents. random: TABLESAMPLE on PostgreSQL/MSSQL tables larger than twice the limit, $sample on MongoDB, head elsewhere (see Sampling) |
SAMPLE_LIMIT |
no | Rows / documents read per table or collection, default 10000 |
NER_ENABLED |
no | true (default): person names in free text through spaCy; false skips the model |
NER_MODEL |
no | en_core_web_trf (transformer, best accuracy — the default when installed) or en_core_web_sm (12 MB, ~40× faster per cell); both ship in requirements.txt |
REPORT_PRIVATE_IPS |
no | false (default): RFC 1918 / loopback / link-local addresses are infrastructure, not PII.IPAddress |
DISABLED_DETECTORS |
no | Detector names never reported (comma-separated or JSON list), e.g. PII.IPAddress,MAC_ADDRESS |
ALLOW_LIST / ALLOW_REGEX |
no | Macie-style exceptions: exact values (comma-separated or JSON list) / JSON list of regexes over values that are never findings |
COLUMN_RATIO / MIN_COUNT |
no | Override every detector's column-classification share (policy default 0.5) / distinct-possible-hits promotion count (policy default 10); empty keeps the per-detector policies |
AGGREGATION_THRESHOLD |
no | Hits per (detector, column) that collapse into one column-level finding, default 25; 0 disables |
OUTPUT_DIR |
no | Findings/work directory. Default <repo>/output; the container image sets /app/output — point it at a mounted volume to persist findings |
KEEP_SCANNED_FILES |
no | true keeps every file a connector downloaded or exported under <OUTPUT_DIR>/scanned/, mirroring the source (s3/<bucket>/<key>, gdrive/<user or drive>/<file id>/<name>, salesforce/<host>/<object>/<id>/<name>), instead of deleting it after classification. Local testing only: it duplicates the sensitive data on disk |
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | S3 |
OBJECT_NAME |
yes | Bucket name |
AWS_ACCOUNT_ID |
yes | Account that owns the bucket(s); recorded in the findings and required by the CSPM backend |
AWS_ACCESS_KEY_ID |
no | Static IAM credentials with s3:ListBucket + s3:GetObject. Leave both unset to use the instance profile / IRSA, or AWS_PROFILE + AWS_CONFIG_FILE for an assume-role profile into another account (see deployments/vm/README.md) |
AWS_SECRET_ACCESS_KEY |
no | with AWS_ACCESS_KEY_ID |
Objects larger than 100 MB are skipped. Archives (.zip/.tar/.gz/.bz2) are unpacked and scanned recursively; CSV/TSV, Parquet, Excel, JSON, XML, PDF, DOCX and images (OCR) have dedicated parsers, everything else falls back to plain-text scanning.
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | AZURE_BLOB (or ADLS, AZURE_STORAGE) |
OBJECT_NAME |
yes | Container name; account/container or the container URL when the identity reads several accounts |
AZURE_STORAGE_ACCOUNT |
yes* | Storage account of the container(s): the account name or its https:// URL (* unless every name carries its account) |
AZURE_SUBSCRIPTION_ID |
yes | Subscription that owns the account; recorded in the findings as account_id and required by the CSPM backend |
AZURE_STORAGE_ENDPOINT_SUFFIX |
no | core.windows.net (default), core.usgovcloudapi.net, core.chinacloudapi.cn |
AZURE_CLIENT_ID, AZURE_TENANT_ID, AZURE_CLIENT_SECRET |
no | A service principal (app registration). Leave unset to use the VM's managed identity or AKS workload identity; a bare AZURE_CLIENT_ID selects a user-assigned managed identity. Creating the registration and the roles it needs: deployments/vm/azure/README.md |
AZURE_STORAGE_SAS_TOKEN / AZURE_STORAGE_CONNECTION_STRING / AZURE_STORAGE_ACCOUNT_KEY |
no | Credential fallbacks when no identity is available, tried in that order before the identity chain; a read + list SAS is the safe one |
The identity needs the role Storage Blob Data Reader on the account (or its resource group) and a network path: the storage firewall allowing the scanner's subnet, or a private endpoint (deployments/vm/azure/). The S3 rules apply: blobs over 100 MB are skipped, archives are unpacked and scanned recursively, workbooks yield one entry per sheet. Directory placeholders of hierarchical (ADLS Gen2) accounts, archive-tier blobs (they need rehydration first), page blobs and soft-deleted blobs are skipped and counted in the scanner stats. Resource ids are blob URLs (https://<account>.blob.core.windows.net/<container>/<blob>) and never carry a SAS token; a blob that could not be downloaded is reported in errors, not recorded as clean.
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | POSTGRES (or POSTGRESQL), MYSQL, MARIADB, MSSQL (or SQLSERVER) |
OBJECT_NAME |
yes | Database name to scan |
DB_HOST |
yes* | Database host |
DB_PORT |
yes* | Typical defaults: 5432 (postgres), 3306 (mysql/mariadb), 1433 (mssql) |
DB_USERNAME |
yes* | A read-only account is sufficient and recommended |
DB_PASSWORD |
yes* | |
DB_URI |
no | Full SQLAlchemy connection string, e.g. postgresql+psycopg2://user:pass@host:5432/db. Overrides all DB_* fields above (* not needed when DB_URI is set) |
Drivers used: psycopg2 (postgres), PyMySQL (mysql/mariadb), pymssql (mssql). All non-system schemas of the database are discovered and scanned, up to 10 000 rows per table.
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | MONGODB (or MONGO, DOCUMENTDB) |
OBJECT_NAME |
yes | Database name to scan |
DB_HOST |
yes* | MongoDB host |
DB_PORT |
yes* | Typically 27017 |
DB_USERNAME |
no* | Omit for unauthenticated instances |
DB_PASSWORD |
no* | |
DB_URI |
no | Full MongoDB URI, e.g. mongodb://user:pass@host:27017/?authSource=admin. Overrides all DB_* fields above (* not needed when DB_URI is set) |
All non-system.* collections of the database are discovered and scanned, up to 10 000 documents per collection. Documents are walked recursively; nested fields are reported with dotted paths.
Reaching a replica set through
kubectl port-forward/ an SSH tunnel: the members advertise cluster-internal hostnames (…rs0-0.…svc.cluster.local) that do not resolve locally, so topology discovery fails with Could not reach any servers. The scanner addsdirectConnection=trueautomatically whenDB_HOSTislocalhost/127.0.0.1; withDB_URI, append?directConnection=trueyourself.
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | DYNAMODB (or DDB) |
OBJECT_NAME |
yes | Table name; OBJECTS_TO_SCAN lists several tables of the same account and region |
OBJECT_REGION |
yes* | Region of the tables; the DynamoDB client is regional (* or AWS_DEFAULT_REGION / the profile's region) |
AWS_ACCOUNT_ID |
yes | Account that owns the tables; recorded in the findings and required by the CSPM backend |
AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY |
no | Static keys, exactly as for S3; leave unset for the instance profile / IRSA, or AWS_PROFILE for an assume-role profile |
One table is one unit, like a relation or a collection: a Scan reads up to SAMPLE_LIMIT items (10 000), which counts against the table's read capacity, and every item is classified like a MongoDB document (attribute paths as context, nested maps and lists walked). The identity needs dynamodb:Scan on the tables (dynamodb:DescribeTable for a preflight); deployments/vm/aws/member-role-policy.json grants both for one region. Change-data capture through DynamoDB Streams remains a master-mode feature.
Ordinary wire-protocol databases to the scanner: the SQL and MongoDB connectors above behind an Azure OBJECT_TYPE alias, so the CSPM can tell the platform apart. Same DB_* variables, one instance per server.
| Service | OBJECT_TYPE |
Connection |
|---|---|---|
| Azure Database for PostgreSQL Flexible Server | AZURE_POSTGRES |
DB_HOST=<server>.postgres.database.azure.com, port 5432. psycopg2 negotiates the TLS the server requires; pin the CA with DB_URI=postgresql+psycopg2://…?sslmode=verify-full&sslrootcert=/etc/ssl/certs/ca-certificates.crt |
| Azure Database for MySQL Flexible Server | AZURE_MYSQL |
Port 3306. PyMySQL negotiates TLS by itself; for *.mysql.database.azure.com hosts the scanner also verifies the server certificate against the system CA bundle (a DB_URI with its own ssl_ca= is left alone) |
| Azure SQL Database / Managed Instance | AZURE_SQL |
DB_HOST=<server>.database.windows.net, port 1433, a contained user in db_datareader. SQL logins only: pymssql cannot present Entra tokens |
| Cosmos DB for MongoDB, RU or vCore | COSMOS_MONGO |
DB_URI = the account's read-only connection string (RU, port 10255) or the vCore mongodb+srv:// URI. Request-rate throttling (error 16500) is retried from the last document read, honouring the server's RetryAfterMs |
Without passwords. DB_AUTH=azure_entra makes the scanner identity's Microsoft Entra access token the password of every connection to a Flexible Server (PostgreSQL and MySQL). DB_USERNAME is the identity's name on the server, created by the server's Entra administrator, on PostgreSQL: SELECT * FROM pgaadauth_create_principal('dspm-scanner-vm', false, false); GRANT pg_read_all_data TO "dspm-scanner-vm";. Tokens last about an hour and are fetched per connection, so long scans keep working.
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | COSMOS_NOSQL (or COSMOS, AZURE_COSMOS) |
OBJECT_NAME |
yes | Database name to scan |
AZURE_COSMOS_ENDPOINT |
yes | https://<account>.documents.azure.com:443/, or just the account name |
AZURE_COSMOS_KEY |
no | The account's read-only key. Leave unset to use the identity chain with the data-plane role Cosmos DB Built-in Data Reader, assigned with az cosmosdb sql role assignment create (the portal's IAM blade cannot grant data-plane roles) |
AZURE_SUBSCRIPTION_ID |
no | Recorded in the findings as account_id |
Every container of the database is scanned, up to SAMPLE_LIMIT items each (SELECT TOP n * FROM c); the Cosmos system properties _rid, _self, _etag, _attachments, _ts are dropped before classification and items are walked recursively with dotted paths like MongoDB documents. The NoSQL API has no server-side random sample, so SAMPLE_STRATEGY=random reads the head. Request-rate limits (429) are retried by the SDK. Resource ids are https://<account>.documents.azure.com/<database>/<container>.
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | GOOGLE_WORKSPACE (or GDRIVE, GOOGLE_DRIVE, GOOGLEWORKSPACE) |
OBJECT_NAME |
yes | What to scan when the GOOGLE_* variables below are unset: a user email (their My Drive) or a shared-drive id |
GOOGLE_SA_KEY_FILE |
yes* | Path to the service-account key JSON (* or set GOOGLE_APPLICATION_CREDENTIALS; Application Default Credentials when neither is set) |
GOOGLE_IMPERSONATE_USER |
no | Workspace user whose My Drive is scanned (requires domain-wide delegation) |
GOOGLE_DRIVE_ID |
no | One shared drive to scan instead |
Setup (the model the DSPM vendors document): create a GCP service account, enable the Drive API, and authorize the account in the Google Admin console for domain-wide delegation with the single read-only scope https://www.googleapis.com/auth/drive.readonly - the scanner cannot modify data by construction. Google-native files are exported (Docs → .docx, Sheets → .xlsx so every sheet is scanned - the CSV export is first-sheet-only - Slides → text; the Drive API caps exports at 10 MB), everything else is downloaded as-is; both go through the same file parsers as S3 objects, one Drive file per findings entry. Files over 100 MB are skipped.
| Variable | Required | Description |
|---|---|---|
OBJECT_TYPE |
yes | SALESFORCE (or SFDC) |
OBJECT_NAME |
yes | My Domain name (acme for acme.my.salesforce.com) or the full instance URL; SF_DOMAIN overrides it |
SF_CONSUMER_KEY |
yes | Connected App credentials (OAuth 2.0 client-credentials flow) |
SF_CONSUMER_SECRET |
yes | |
SF_OBJECTS |
no | Pin the sObjects to scan (comma-separated or JSON list); default: every queryable business object that holds records |
SF_INCLUDE_FILES |
no | true (default): also scan the files attached to records - the latest ContentVersion of every File plus classic Attachment bodies - through the file parsers |
SF_API_VERSION |
no | Default v62.0 |
Setup: a Connected App with the client-credentials flow enabled, run as an integration user holding a permission set with "API Enabled", "View All Data" and "Query All Files". One sObject is one unit, scanned like a table: text-typed fields only (string / textarea / email / phone / url / picklist), up to SAMPLE_LIMIT records per object. *Share / *History / *Feed / *ChangeEvent, Apex* / Datacloud* and other system objects are excluded, and empty objects are skipped via limits/recordCount. Incremental scans filter on SystemModstamp.
OBJECT_TYPE=POSTGRES
OBJECT_NAME=appdb
DB_HOST=127.0.0.1
DB_PORT=5432
DB_USERNAME=scanner
DB_PASSWORD=secret
CSPM_URL=https://cspm.example.com/
ARTIFACT_TOKEN=eyJ...Invocation payload shape:
{
"scan_type": "...",
"target": { ... },
"config": { "enabled_regions": ["US", "IN"] }
}ARTIFACT_TOKEN + CSPM_URL environment variables control the findings upload in this mode.
| Target field | Required | Description |
|---|---|---|
bucket |
yes | Bucket name |
key |
yes | Object key |
version_id |
no | Specific object version |
last_modified |
no | With config.last_scan_time, enables skip-if-unchanged |
scan_type: "postgres" | "postgresql" | "mysql" | "mariadb" | "mssql" | "sqlserver" | "azure_postgres" | "azure_mysql" | "azure_sql"
| Target field | Required | Description |
|---|---|---|
host, port |
yes* | Database endpoint |
username, password |
yes* | Credentials |
database |
yes* | Database name |
connection_string |
no | Full SQLAlchemy DSN; overrides the fields above (*) |
password_secret |
no | AWS Secrets Manager ARN/name, or an Azure Key Vault secret URI (https://<vault>.vault.azure.net/secrets/<name>, read with the scanner identity); fills any missing username/password/host/port/database/uri |
auth |
no | azure_entra: the scanner identity's token is the password (Azure Database for PostgreSQL / MySQL Flexible Server) |
schema |
no | Restrict to one schema (default: all non-system schemas) |
tables |
no | Restrict to specific tables, plain or schema-qualified: ["users", "sales.orders"] |
include_views |
no | Also scan views (default false) |
incremental_column |
no | Timestamp column for incremental scans |
last_scan_time |
no | Only rows where incremental_column > last_scan_time are scanned |
sample_limit |
no | Max rows per table (default 10000) |
Same fields as the SQL engines above (set engine to one of postgres/mysql/mariadb/mssql), plus:
| Target field | Required | Description |
|---|---|---|
engine |
yes | Database engine of the RDS instance |
reader_endpoint |
no | Aurora reader endpoint |
use_reader |
no | Route the scan to reader_endpoint |
| Target field | Required | Description |
|---|---|---|
host, port |
yes* | MongoDB endpoint (port defaults to 27017) |
username, password |
no* | Omit for unauthenticated instances |
uri |
no | Full MongoDB URI; overrides the fields above (*) |
password_secret |
no | AWS Secrets Manager ARN/name or Key Vault secret URI, as for SQL |
database |
no | Restrict to one database (default: all non-system databases) |
collection |
no | Restrict to one collection (default: all non-system.* collections) |
incremental_field |
no | Field for incremental scans, with last_scan_time |
last_scan_time |
no | Only documents where incremental_field > last_scan_time |
sample_limit |
no | Max documents per collection (default 10000) |
| Target field | Required | Description |
|---|---|---|
table_name |
yes | DynamoDB table name |
region |
no | AWS region (default from environment) |
sample_limit |
no | Max items (default 10000) |
Uses ambient AWS credentials (Lambda role / environment). DynamoDB Stream CDC batches are handled automatically when the Lambda is wired to a stream.
| Target field | Required | Description |
|---|---|---|
account |
yes* | Storage account name or URL (* default AZURE_STORAGE_ACCOUNT) |
container |
yes | Container name |
blob |
no | One blob; omit to scan the container (prefix restricts it to a virtual directory) |
version_id, snapshot |
no | With blob |
last_scan_time |
no | Skip blobs not modified since (ISO 8601) |
max_blobs |
no | Cap on blobs listed per run |
connection_string, sas_token, account_key, endpoint_suffix |
no | Credential / endpoint overrides; password_secret may supply them from Key Vault or Secrets Manager |
Without overrides the identity chain is used (service principal from AZURE_CLIENT_ID / AZURE_TENANT_ID / AZURE_CLIENT_SECRET, workload identity, managed identity), as in worker mode.
| Target field | Required | Description |
|---|---|---|
endpoint or account |
yes* | Account URL or name (* default AZURE_COSMOS_ENDPOINT) |
key |
no | Read-only key; omit for the identity chain (default AZURE_COSMOS_KEY) |
database, container |
no | Restrict to one database / container (default: all) |
last_scan_time |
no | Only items whose _ts is after it |
sample_limit |
no | Max items per container (default 10000) |
password_secret |
no | Key Vault secret URI or Secrets Manager ARN supplying key / endpoint |
{
"scan_type": "google_workspace",
"target": {
"sa_key_file": "/path/key.json",
"impersonate_user": "user@example.com",
"drive_id": "0AbCd...",
"folder_id": "1XyZ...",
"max_files": 500,
"last_scan_time": "2026-08-01T00:00:00Z"
}
}impersonate_user (My Drive via domain-wide delegation) or drive_id (a shared drive) selects the corpus; folder_id restricts to one folder; omit sa_key_file to use Application Default Credentials.
{
"scan_type": "salesforce",
"target": {
"domain": "acme",
"consumer_key": "3MVG9...",
"consumer_secret": "...",
"objects": ["Contact", "Lead"],
"include_files": true,
"sample_limit": 10000,
"last_scan_time": "2026-08-01T00:00:00Z"
}
}Like the database targets, password_secret (an AWS Secrets Manager ARN or an Azure Key Vault secret URI) can supply consumer_key / consumer_secret / domain - or a ready access_token + instance_url - instead of inline values.
| Key | Default | Description |
|---|---|---|
enabled_regions |
[] |
Regional compliance packs (ISO alpha-2), 62 packs: AE, AR, AT, AU, BE, BG, BR, CA, CH, CL, CN, CZ, DE, DK, EE, EG, ES, FI, FR, GB, GH, GR, HK, HR, HU, ID, IE, IL, IN, IS, IT, JP, KR, LK, LT, LU, LV, MX, MY, NG, NL, NO, NZ, PH, PK, PL, PT, RO, RS, RU, SA, SE, SG, SI, SK, TH, TR, TW, UA, US, VN, ZA — national ids, tax numbers, passports, driver licences, health identifiers with their public check-digit algorithms |
phone_regions |
enabled_regions |
Regions used to parse national-format phone numbers (5678942315 in a mobile field); international +… numbers are always detected |
chunk_size |
5000 (SQL) / 1000 (Mongo) | Rows/documents fetched per batch |
connect_timeout |
10 | Connection timeout in seconds (SQL and Mongo) |
log_queries |
false (worker mode: always on) |
Log every query issued during DB scans (dialect-compiled SQL with bound values, Mongo filters, DynamoDB scans). Note: emits table/column names into logs |
last_scan_time |
– | S3 only: skip objects not modified since this timestamp |
keep_files_dir |
– | Keep every downloaded/exported file under this directory, mirroring the source path, instead of deleting it after classification (the worker sets <OUTPUT_DIR>/scanned when KEEP_SCANNED_FILES=true). Local testing only |
aggregation_threshold |
25 |
A (detector, column) pair firing on at least this many rows/documents/cells collapses into one column-level finding with an occurrences count. 0 disables |
min_confidence |
likely |
Lowest confidence tier reported: possible / likely / very_likely (a legacy score_threshold float is accepted) |
column_ratio |
per detector (0.5) |
Share of a column's sampled non-empty values that must match before the column is classified (Sentra's 50 % rule); overrides every detector policy |
column_min_matches |
per detector (3) |
Distinct matching values needed before a column verdict is drawn |
min_count |
per detector (10) |
Distinct possible hits in one unit (file, table column) that promote them to likely |
allow_list / allow_regex |
[] |
Macie-style allow lists: exact values / regexes never reported (public phone numbers, sample data) |
adaptive_sampling |
false |
Stop reading a unit once its column verdicts settle; settle_min_records (2000), settle_window (1000), settle_margin (0.05) tune the stop rule |
sample_strategy |
head |
random draws rows with TABLESAMPLE (PostgreSQL, MSSQL) / $sample (MongoDB) instead of reading the head; per-target sample_strategy overrides it |
ner |
true |
Person names in prose via the spaCy model |
field_suppression |
true |
Structural field-name rules: token detectors never fire in *_id/hash/etag/path/… fields, digit-run detectors never fire in counter/timestamp fields. Corroborated findings are exempt |
decode_base64 |
true |
Decode base64 blobs (Authorization: Basic …, base64 JSON, PEM) and scan the plaintext |
entropy_report_uncorroborated |
false |
Report random-looking tokens that have no supporting evidence as Secret.TokenLikeValue (Medium) instead of dropping them |
disabled_detectors |
[] |
Detector names never reported |
report_private_ips |
false |
Report RFC 1918 / loopback / link-local addresses as PII.IPAddress; by default only public IPs count |
direct_connection |
auto | MongoDB: add directConnection=true to the built URI. Automatic for localhost/127.0.0.1, i.e. port-forwarded replica sets whose members advertise cluster-internal hostnames |
column_suppression |
id/hash rule for entropy | Scanner-level escape hatch: per-detector regexes of column/field names to skip on top of the engine's own field rules. Pass {} to disable, or your own {detector: regex} map |
entropy_min_length |
24 |
Minimum token length for the entropy detector |
entropy_min_entropy |
4.5 |
Shannon-entropy threshold for base64-shaped tokens (hex tokens use 3.0 and need 32+ chars) |
connector ──► Record / TextBlob stream ──► src/pipeline (per unit) ──► findings
│
src/engine (per value)
The code is split the way the vendor engines we studied are (Wiz, Cyera, Orca, Sentra, Varonis, Amazon Macie, Google Sensitive Data Protection, Microsoft Purview, Nightfall — see How the vendors do it below):
- Connectors (
src/scanners/) know a data source. They enumerate its units (tables, collections, objects, sheets) and turn each unit into a stream ofRecords (rows / documents made ofCells that carry the value, its column name or field path, and a rendered location) orTextBlobs (pages, paragraphs, text blocks). They never call the detection engine. - The engine (
src/engine/) judges one value at a time: pattern + validator + context words + field name → a score and a confidence tier. It stays the place where detectors are defined. - The pipeline (
src/pipeline/) does everything a DSPM product does after pattern matching, once, for every connector: context policy, record-level corroboration, column density verdicts, minimum counts, adaptive sampling and aggregation.
Every finding carries confidence and evidence:
Confidence tiers express heuristic evidence strength, not calibrated probabilities of correctness.
| Tier | Meaning | Examples |
|---|---|---|
very_likely |
validated shape and corroboration | checksum-valid national id next to its keyword or in a column named for it, JWT with a decodable header, vendor-prefixed token, opaque value in a credential-named column, a classified column of likely cells |
likely |
one strong signal | valid e-mail (IANA TLD, no demo/automated sender), Luhn + issuer prefix with separators, mod-97 IBAN, structured street address, a plausible shape backed by its column name or a context word, a classified column of possible cells |
possible |
plausible shape only | SSN-shaped digit groups, a bare mod-10/mod-11-valid number, a random-looking token, a BIC-shaped word. Never reported on its own |
MIN_CONFIDENCE picks the lowest tier reported (likely by default — Google's default is possible, Nightfall recommends likely). Checksums are not enough on their own: a mod-10/mod-11 check passes ~10 % of random numbers, so a checksum-valid number with no context word, field hint or column evidence is possible (a US phone number is a valid NHS number one time in ten). Epoch timestamps are never identifiers; private/loopback IPs are infrastructure (report_private_ips); documented examples (AKIAIOSFODNN7EXAMPLE, test cards, hunter2 in advisory text) are never real.
evidence lists what backs the finding: checksum, format (self-validating shape), context:<word>, field (the column/path names the entity), key:<name> (assigned to a credential keyword), column:<ratio> (column density), record:identity (same record holds two identity signals), count:<n> (many in one file), shape / needs_context / uncorroborated (why something stayed possible).
Each detector has a reporting policy (src/engine/policy.py, looked up by name with category defaults, so a new recognizer needs no entry unless it deviates):
| Field | Meaning | Vendor precedent |
|---|---|---|
context |
required: a hit with neither validation nor context is capped at possible (national ids, bank accounts, cards, SWIFT, opaque tokens); boost: context raises the tier; none: self-identifying (e-mail, JWT, IBAN, vendor tokens) |
Macie keyword requirements per data type |
column_ratio, column_min_matches |
share of a column's sampled values (and distinct matches) that classify the column | Sentra 50 % rule, Google column profiles |
min_count, count_promotion |
distinct possible hits in one unit that become likely; off for word-shaped patterns (SWIFT/BIC) and random strings |
Orca statistical scan, Purview "low-confidence patterns with 20+ instances", Nightfall minimum findings |
identity, identity_corroboration |
identity signals (name, e-mail, phone, address, DOB) promote a possible national id / card in the same record |
Purview supporting elements, Cyera identifiability |
negative_fields |
column names that veto the detector (txn, hash, invoice, port, amount …) |
DLP negative keywords, Google exclude-by-hotword |
Structured data carries its strongest signal in the field name (DetectionEngine.scan_text(text, field_name=...); the name is never mixed into the scanned text). Macie counts a keyword in the column name or any element of the JSON path as proximity, and so does the engine: src/engine/context.py classifies credential-named fields (token, secret_key, authorization, cookie, webhook_url, …) whose values are credentials whatever their entropy, identifier fields (*_id, hash, etag, path, …) that never yield token findings, counter/timestamp fields that never yield card/id numbers, full_name/first_name fields that yield PII.PersonName, and mobile/phone fields that enable national-format phone parsing (a phone-named column whose numbers do not parse for the enabled regions still yields possible phones that column density can promote). Overlapping matches keep the most specific detector (a JWT is not also a bearer token, a card number is not also a bank account).
Names in prose. A name in a name-labelled column is a field rule; a name inside a support note, a PDF page or a comment field needs a model (Macie NAME, Purview named entities, Cyera NER). With spaCy (requirements.txt ships en_core_web_trf, a RoBERTa transformer with OntoNotes NER F1 0.90, and en_core_web_sm, F1 0.84, as the fallback; NER_MODEL chooses, NER_ENABLED=false skips), src/engine/ner.py accepts PERSON entities that look like written names — two to four title-case tokens, no digits or acronyms, no company suffix — and entities the small model mislabels when their first token is in the shipped given-name lexicon (src/engine/data/given_names.txt). A name next to a context word or honorific (patient, customer, employee, regards, dear, Mr/Dr …) is likely; otherwise it is possible and the pipeline decides: a document with ten distinct names is likely, a lone name in a sentence stays hidden.
- Column density — a column whose sampled non-empty cells match a detector at
column_ratio(default 50 %) with at leastcolumn_min_matchesmatching cells and distinct values is classified as that detector one tier above its cells. Each cell contributes once, even if it contains several matches; validation density and the majority confidence tier also count cells. Forty SSN-shaped values under a meaningless header can classify a column, while three shapes in one of six cells represent 1/6 density, not 3/6. - Unit-name context — the table, collection, sheet or object name is part of the path to a value (Macie counts a keyword "in the name of an element in the path"): a
possiblecard number incredit_cards.numberor an SSN-shaped value inssn_export.csvislikely(unit:<name>). Onlypossiblecandidates are lifted; very weak patterns still need their column name. - Sibling columns — companion names can raise an established column verdict one more tier:
expiry/cvvnext to a card column, for example (siblings:<columns>). They must share the same parent field path. A sibling name elsewhere in the unit cannot promote an isolated weak match on its own. - Record corroboration — eligible
possibleidentifiers next to two identity signals becomelikely(record:identity). For documents, those signals must belong to the same containing object; separate array members remain separate contexts. Bare-number and weak-checksum candidates markedneeds_contextrequire stronger support, such as a detector-specific field/keyword or sufficient column density; a nearby name and email alone do not establish their type. - Column exclusivity — the dominant detector suppresses weak alternative interpretations, such as a phone number that happens to pass an NHS checksum. Independently supported findings in separate spans or cells remain reportable: a notes field can contain both email addresses and phone numbers, and an email-heavy column can contain an explicitly labelled SSN. Record or column promotion alone does not establish independent support.
- Minimum counts — in a document or a mixed column,
min_countdistinctpossiblehits of one detector becomelikely(count:<n>): a file with 30 SSN-shaped numbers is not a coincidence, one is. - Aggregation — a (detector, column) pair with
aggregation_thresholdor more hits collapses into one column-level finding.occurrencescounts findings;column_matchesandcolumn_validated_matchescount matching cells and cells with validation evidence.column_sampledis the non-empty cell denominator forcolumn_ratioandcolumn_validated_ratio. Validation evidence includes checksums and recognized formats; it does not prove an identity or credential is real. These details are available in pipeline findings; the worker's grouped output currently retains counts and confidence but omits the column statistics. - Allow lists —
allow_list/allow_regexsuppress known values (public contact numbers, sample data), like Macie allow lists.
Connectors read up to sample_limit rows/documents per unit (10 000). By default that is the head of the table; SAMPLE_STRATEGY=random (or sample_strategy per target) draws a random sample instead — TABLESAMPLE SYSTEM (p) on PostgreSQL and MSSQL when the planner estimate says the table holds more than twice the limit (p is sized to about three times the limit before LIMIT cuts it), $sample on MongoDB, the head on MySQL/MariaDB where no cheap random read exists. With adaptive_sampling the pipeline stops reading a unit once settle_min_records records have been seen, no new (column, detector) pair appeared for settle_window records and no column sits within settle_margin of its classification ratio. Rows already read count in the stats.
The current adaptive stop rule is a stability heuristic, not a statistical confidence bound. A stable sample, especially a head or page sample, does not guarantee discovery of rare sensitive rows or prove the unread portion is clean. For detectors that require validation density, the stop rule uses the validated cell ratio, as column classification does.
Every finding carries resource_id, detector, category, severity, value (capped at 200 characters), location, confidence, evidence and a value_hash (sha256 prefix for correlating the same value across scans). Column-level findings add the aggregation fields above. The worker's clubbed JSON entries carry the highest confidence among their findings.
Recognizer packs. src/engine/recognizers/ holds 159 country-specific and generic recognizers across 62 region packs as native Rule objects (src/engine/rules.py: pattern scores, context words, validators/invalidators; test vectors in tests/test_recognizers_*.py). Rules are grouped by region pack; generic ones (IBAN, crypto wallets, IP, MAC, IMEI, ICCID, VIN, passport MRZ, coordinates, ICD-10 / NDC codes, medical record numbers) always run, URL/UUID are shipped disabled. Validators return True (checksum holds), False (dropped) or None where the algorithm is not authoritative for every number (Danish CPR after 2007, Latvian 32-prefixed codes, UK UTR, Mexican RFC). Every detector name must exist in fixtures/findings-mapping.json (tests/test_detector_names.py).
Regression corpus. tests/fixtures/detection_corpus.json is an anonymised corpus built from real Postgres/Mongo scans: 370+ reviewed false positives that must stay silent and 160+ true positives that must stay detected (tests/test_regression_corpus.py). tests/test_pipeline.py covers the column/record/file rules and the connector contract.
The original corpus checks selected expected/forbidden detectors; its reported precision and recall are targeted regression measures. The additional fully labelled synthetic corpus, tests/fixtures/accuracy_cases.json, also counts unexpected detections on positive cases and checks exact values and cell/blob locations. It covers mixed content, sparse weak matches and encoded payloads. Run it with:
python -m scripts.evaluate_accuracy --check --output /tmp/dspm-accuracy.jsonThe evaluator disables aggregation, reports per-detector precision/recall/F1, and writes only case identifiers and counts, not detected values. Undefined metrics are null. These development cases are not a held-out production benchmark. See the competitive assessment and accuracy review for sources, measured changes and the next engineering priorities.
The current tuning priority is reducing false alarms. The additional tests/fixtures/false_alarm_cases.json corpus covers weak numeric references beside identity fields, context crossing between nested objects or array members, and positive controls. Run it with python -m scripts.evaluate_accuracy --corpus tests/fixtures/false_alarm_cases.json --check. The evaluator accepts flat rows, nested documents, or text cases. See false-alarm reduction for measured results and the recall tradeoff.
Sample dataset. sample_data/all_detectors.jsonl / .txt hold one synthetic example per detector (built by python -m tests.sample_dataset_builder; --check scans them and lists anything not detected) — scan them to see every finding type the engine can produce, and tests/test_sample_dataset.py keeps them in sync with the catalogue.
A connector is a BaseScanner subclass (src/scanners/base.py) that implements how to reach and read a source; classification is inherited:
from src.pipeline import Cell, Record, TextBlob # what a connector emits
from src.scanners.base import BaseScanner
class SnowflakeScanner(BaseScanner):
def __init__(self, engine, config=None, client=None):
super().__init__(engine, config, client)
self.stats = {"tables_scanned": 0, "rows_scanned": 0, "errors": 0}
def iter_scan(self, target): # one unit at a time, so callers can checkpoint
for schema, table in self._list_tables(target):
resource_id = f"snowflake://{target['account']}/{schema}.{table}"
findings = self.classify( # the whole pipeline: context policy, column verdicts,
resource_id, # record corroboration, min counts, aggregation
self._rows(schema, table, target),
location_fn=lambda column, n: f"Table '{table}', Column '{column}' ({n} matches)",
)
self.stats["tables_scanned"] += 1
yield resource_id, f"{schema}.{table}", self.dedup_findings(findings)
def scan(self, target):
return self.collect(target)
def _rows(self, schema, table, target): # a generator of Records; count rows in stats
for idx, row in enumerate(self._query(schema, table, target.get("sample_limit", 10000))):
self.stats["rows_scanned"] += 1
yield Record([Cell(str(v), column, f"Table '{table}', Row {idx}, Column '{column}'") for column, v in row.items() if v])Rules of the contract:
- Emit
Records for rows/documents (Cell.fieldis the column name or dotted path the engine uses as context;Cell.keythe aggregation key when array indices should collapse;src.pipeline.document_recordbuilds one from any nested document) andTextBlobs for free text (locate(start, end)renders span locations). - Keep
locationstrings connector-specific and stable; passlocation_fn(column, n)for column-level findings. - Count what you read in
self.statsinside the generator (it keeps counting when adaptive sampling stops early) and letclassify()record read errors (self.stats["errors"],error_details). - Files of any origin go through
src/scanners/files.iter_units(path, resource_id, config)— an object-store connector only downloads, intoself.workdir(resource_id), released withself.discard_workdir()afterwards, soKEEP_SCANNED_FILESworks for every connector. - New detectors are
Rules insrc/engine/recognizers/plus afixtures/findings-mapping.jsonentry, and, only if they deviate from their category, aDetectorPolicyinsrc/engine/policy.py.
- SQL engines: pass DBAPI options via the target's
connect_args(master mode), e.g.{"sslmode": "verify-full", "sslrootcert": "/certs/rds-ca.pem"}for PostgreSQL — or put them on the DSN:DB_URI=postgresql+psycopg2://user:pass@host:5432/db?sslmode=require. # pragma: allowlist secret - MongoDB / DocumentDB: use URI parameters:
DB_URI=mongodb://user:pass@host:27017/?tls=true&tlsCAFile=/certs/global-bundle.pem(DocumentDB requires TLS with the Amazon CA bundle). # pragma: allowlist secret - Mount the CA bundle into the container (e.g. via a ConfigMap/Secret volume) and reference it by path.
python run_tests.py