Skip to content
SmiKar Software

How Forage Works

4 min read · Last updated · Page version 3

Forage is in development. The whole product is, not just part of it. What these pages describe as working is built and in beta testing with customers; the retention and disposal half is not built yet. Every page says which is which.

The sequence

1. Forage reads Squirrel's record of the estate. It builds its inventory of SharePoint Online files and Squirrel-archived files from the records Squirrel already keeps, so SharePoint is not crawled a second time.

Because the inventory comes from Squirrel, Forage sees what Squirrel records and nothing else. It reads the sites on Squirrel's list of sites.

2. It reads each document, and records what happened to it. PDF, Word, Excel and PowerPoint files including templates and macro-enabled files, older Office formats, Outlook and standard email messages, and rich text, plain text, CSV, HTML and XML, without needing Microsoft Office installed.

Scanned PDFs and image files are read with optical character recognition, so scanned contracts and forms become searchable. A page that cannot be read with enough confidence is left out and counted rather than indexed, and a page that passes can still contain misread words.

For every document it reads, Forage records how it was read, page by page, and keeps a cryptographic fingerprint of the file.

3. It keeps a searchable index inside your own deployment. One full-text search across live and archived content, with highlighted extracts and filters by source and document type.

4. Where content classification is enabled, it applies your content rules. Forage recognises what documents contain using transparent rules built from phrases, patterns, checksum validation and proximity. Every finding keeps the words that produced it, and the same document always gives the same result.

5. The text stays in your deployment. Forage keeps the text it reads, and its search index of that text, inside your own deployment, and sends no document content to any external AI service.

Reading the archive

Forage reads the content Squirrel has archived, opening the encrypted, compressed archive copies inside your own deployment with your own archive key. It reads the current version of each file.

Your archive key has to be provided to the deployment for this to work. It is not automatic.

Keeping current

Forage keeps its picture of the estate current by asking Squirrel what was added, changed or deleted.

When a file is reported as changed, Forage fetches it and compares its fingerprint with the one it recorded, and reads it again only if the content really changed. A rename or a metadata edit moves a file's modified date as readily as new content does, so a change is treated as a trigger to look, never as a conclusion that the content differs.

Forage picks up changes on a regular sweep rather than instantly, and it can be no more current than Squirrel's own record.

Accounting for everything

Every file in the estate is accounted for, whether read, excluded by type or by policy, or failed with a recorded reason. Coverage is reported against the whole estate rather than against what was convenient to read.

That is the point of the Coverage and Exceptions pages: what was not read is visible, with the reason, rather than quietly missing.

The half in development

Designed, being built, and not available today:

Retention rules would start a clock on each document. Candidates for disposal would go to a person for review, with documents on legal hold filtered out first. Only approved documents would be destroyed, each after a fresh check immediately beforehand, and every destruction would produce a certificate.

See safety principles for the design, and limitations for the full list of what is not yet available.


Questions go to sales@smikar.com.

More in Squirrel

See all pages →