Sources (bucket sync)
A Source connects an object-storage bucket to a memory bank. Hindsight lists the bucket on a schedule you choose, ingests every new or changed object as a document, and removes documents whose objects have been deleted — no upload pipeline to run on your side.
Amazon S3 and S3-compatible stores (MinIO, Cloudflare R2, …) are supported today. Google Cloud Storage and Azure Blob Storage appear in the provider list as "coming soon".
Sources is available on Enterprise plans. If the Sources tab on your bank's Profile page shows a locked panel, book a call or contact your Hindsight representative to enable it for your organization.
Before you start
Create a read-only IAM identity for Hindsight. It needs exactly two permissions: list the bucket, and read objects under the prefix you will sync.
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": "s3:ListBucket",
"Resource": "arn:aws:s3:::my-company-docs",
"Condition": { "StringLike": { "s3:prefix": ["reports/*"] } }
},
{
"Effect": "Allow",
"Action": "s3:GetObject",
"Resource": "arn:aws:s3:::my-company-docs/reports/*"
}
]
}
Drop the Condition and use arn:aws:s3:::my-company-docs/* if you want to
sync the whole bucket. Hindsight never writes to your bucket.
Adding a source
- Open the bank, go to Profile → Sources and click Add source.
- Fill in:
- Name — how the source is listed.
- Bucket and Region (required for Amazon S3 — a wrong region is
reported by Test connection as "wrong region", with the right one when
AWS tells us). For S3-compatible stores only, set a Custom endpoint
(for example
https://<account>.r2.cloudflarestorage.com); the region may then be left blank. The endpoint must be a publichttps://address — private, internal and plain-http endpoints are refused — and changing it later requires re-entering the access key and secret. - Prefix — only objects under this key prefix are synced. Leave it blank for the entire bucket.
- Access key ID / Secret access key — the read-only credentials above. The secret is encrypted at rest and never shown again; to rotate it, edit the source and enter a new pair.
- Extraction method — the same choice as Add Document. Standard is free and handles most documents; image files are skipped (listed in the run with a reason). Enhanced (Iris) carries an additional charge and extracts scanned documents, complex layouts and images. Changing it only affects objects synced afterwards.
- Schedule — pick how often to sync: every 15 minutes, every hour (the default), every 6 hours, daily at a time, or weekly on a day at a time. Times are entered in your local time zone and stored in UTC, so a daily schedule shifts by an hour when daylight saving changes. Custom accepts a five-field cron expression evaluated in UTC. Whatever you pick, the preview shows the next three runs in UTC and in your local time. Runs must be at least 5 minutes apart.
- Click Test connection. Hindsight lists the first few objects under the prefix so you can confirm the credentials and prefix are right.
- Click Create source. Creation also validates the credentials; a source with a bad key is rejected rather than saved.
Only organization owners and admins can add, edit, sync or delete sources, or test a connection. Members can see the list and the run history.
What happens on each run
Each run:
- Lists the bucket under the prefix.
- Compares what it finds with the documents already in the bank that came from this source, using the object's ETag. Unchanged objects cost nothing.
- Ingests new and changed objects. Text-like files (
.md,.txt,.json,.csv,.html, …) are retained directly; everything else (PDF, Office documents, images, …) is converted to text with the source's extraction method first, then retained. - Deletes documents whose objects no longer exist in the bucket.
Objects are ingested as documents with the id s3://<bucket>/<key> and the tag
source:<source id>, so you can filter them on the Documents tab and in
recall.
A run submits the objects; conversion and memory extraction finish asynchronously
in the background, exactly as if you had called
/files/retain yourself. The run's status reflects submission:
| Status | Meaning |
|---|---|
| Succeeded | Every object that needed work was submitted. |
| Partial | Some objects could not be read or submitted. Expand the run to see the per-object errors; they are retried on the next run. |
| Failed | The run could not start or was stopped. A connection problem or running out of credits is retried on schedule (credits: as soon as the balance allows). Bad credentials or a deleted bank pause the source: it shows Needs attention, the organization owner and billing contact receive an email, and your organization webhooks receive a source.paused event. Fix the cause, then re-enable the source. |
| Interrupted | The run was cut off (for example during a deployment). The next run resumes where it stopped. |
Very large buckets are swept in slices: a run stops after its object limit (5,000 by default) and the next run continues from the same point. A partial sweep never deletes documents, so nothing is removed until Hindsight has seen the whole bucket.
Objects larger than 25 MB are skipped and reported in the run.
Protection against accidental mass deletes
A full sweep that would remove more than 20% of the source's documents (and more than 10), or that finds no objects at all while the bank still holds documents from the source, is treated as a mistake — a prefix typo, a narrowed IAM policy, a bucket rename — rather than applied. The run ends Failed with the message "refused to delete N of M documents", the source pauses and you are notified as for any pause. New and changed objects found in that sweep are still ingested.
If the removal is intended, re-enable the source: the next run applies the deletes once, then normal protection resumes. Editing a source's bucket or prefix restarts the sweep from scratch, so the first run after such an edit often triggers this guard on purpose.
Notifications
When a source is paused, Hindsight:
- emails the organization owner and billing contact with the source, bucket and reason, and a link back to the Sources tab;
- sends a
source.pausedevent to every organization-level webhook subscribed to it. These are the org-wide subscriptions used for SIEM delivery; register one withPOST /ext/organizations/{org_id}/webhooksand"event_types": ["source.paused"](ask your Hindsight representative if you need help). The payload carriesorg_id,bank_id,source_id,source_name,provider,bucket,prefix,run_id,reasonandtimestamp, signed like every other delivery.
Runs that fail for transient reasons do not notify; check the run history on the Sources tab.
Sync now
Sync now queues a run outside the schedule; it starts within about a minute. Expand the row to follow it. Only one run per source is ever in progress.
Billing
Synced content is billed the same way as content you retain through the API: by the memory extracted from it. Listing the bucket, comparing ETags and converting files to text are not metered. If your organization runs out of credits mid-run the run stops with Failed and resumes once credits are available.
Deleting a source
Deleting a source stops the schedule and removes the source's configuration and
credentials. Documents it already ingested stay in the bank. Delete them
from the Documents tab (filter by the source: tag) if you no longer want
them.
Notes and limits
- Documents you delete by hand from the bank are re-ingested on the next run if the object still exists in the bucket. Remove or move the object instead, or narrow the prefix.
- Memory Defense scans synced content like any other retain, so buckets that contain secrets or personal data will raise security events.
- Under Standard extraction, scanned (image-only) PDFs produce empty documents and image files are skipped; choose Enhanced (Iris) for those.
- Original files are not stored in Hindsight; the bucket remains the source of truth.