Forage Limitations
8 min read · Last updated · Page version 2
This page states plainly what Forage does not do, how today's features behave at their edges, and what is designed but not yet available.
Part 1: scope, by design
Sources. Forage covers SharePoint Online and the content Squirrel has archived from it. OneDrive, Exchange mailboxes, Microsoft Teams messages and other systems are not in scope.
Squirrel is required. Forage builds its picture of the estate from Squirrel's record, so a site or file that Squirrel does not record is not seen by Forage, and Forage is only as current as that record. Forage reads the sites on Squirrel's list of sites. Files Squirrel holds for a site that is not on that list are outside the estate and are not read, and the Coverage page shows how many there are.
Documents only. Forage reads documents. Video and audio are not read.
The current version only. Forage reads the current version of each document. Earlier versions, whether in SharePoint's version history or kept inside an archived file, are not read.
Nothing is left out silently. Files Forage does not read are still counted, with the reason: a type it does not read, a type an administrator has switched off, or a file that could not be read. The one exception is a switched-off type whose records an administrator has chosen to delete.
Part 2: how today's features behave
Reading documents
Protected files. Forage never removes or bypasses a password or encryption. A password-protected or encrypted PDF or modern Office file is recorded as protected and is not read; this includes Office files encrypted by a sensitivity label or by rights management. An older-format Office file with a password may instead be recorded as unreadable or partly read; its content is not searchable either way. A PDF that opens without a password but restricts printing or copying is read normally.
Formats read in part. Excel workbooks are read for their words, but cells that hold only a number or a date are not yet searchable, so a rule looking for a number or a date stored that way will not find it there; every such workbook is marked as partly read. Older Word and PowerPoint files are read for their text but not their layout, and are marked as partly read. Attachments inside email messages are listed by name; their contents are not read as part of the message.
Formats not read today. Some formats are not read today, including binary Excel workbooks, compressed (zip) folders and design files. They are counted, with the reason.
Pictures inside documents. Pictures embedded in Word, Excel and PowerPoint files are not read. A PDF page is read with optical character recognition only when it carries no text at all, or when the document's text is unreadable, so pictures on a page with readable text, including a scanned page with a line of added text, are not read.
Very long documents. Very long documents are read up to a set amount of text, long scanned documents up to a set number of pages, and any document up to a time limit. What was read is searchable, and the document is marked as partly read.
Scanned pages
Pictures with no text. Photographs, drawings and other pictures with no text have nothing to search. Forage does not describe what a picture shows.
Language today. Scanned pages are read with an English reader today. In other Latin-script languages, accented letters are often dropped or misread, and text in other scripts, such as Cyrillic, is largely lost.
Handwriting and poor scans. Handwriting is read poorly or not at all. Faint, skewed or rotated pages are often read poorly. A page that cannot be read with enough confidence is left out and counted rather than indexed, and a page that passes can still contain misread words.
Pages that mix scripts. A page that mixes scripts, such as Latin and Cyrillic text on one page, is read poorly: part of it is lost.
A stronger reader is being added. Forage is being extended to read scanned pages with a stronger optical character recognition model on deployments with a GPU, including rotated pages and languages beyond English.
Search
Search languages. Search matches forms of the same English word. In other languages it matches the word as written, ignoring accents. Text in Chinese or Japanese characters is not searchable today.
Results are not trimmed to permissions. Search results are not limited to what each person may open in SharePoint.
What has been read so far. Search and findings cover only what Forage has read so far. The Coverage page shows how much of the estate that is, and its estate figures are a periodic count that says when it was taken, not a live figure.
Content rules and labels
Rules match exactly. Content rules match exactly the words and patterns written into them, ignoring capital letters, with checksum checks and nearby-word conditions where a rule asks for them. They make no judgement about meaning. Unlike search, a rule does not match other forms of a word: each form it should find has to be listed.
Identifier coverage and accuracy. Built-in checksum checks cover payment card numbers, international bank account numbers, Australian business and tax file numbers, and UK NHS numbers. Identifiers from other countries need rules written for them, without a built-in checksum. No accuracy figure is measured for any rule; a dry run's checked sample is the evidence.
Findings cover what has been examined. Findings cover only documents Forage has read and examined against the rules active now. When the rules change, Forage re-examines everything in the background and the counts refill. Forage keeps the current finding for each document and rule; when a new version of a rule examines a document, the earlier version's finding is replaced, not kept as history.
Microsoft labels. Search results show the retention label and the sensitivity label that Squirrel recorded for each file. Forage never applies, changes or removes a label, does not act on retention labels, and does not use sensitivity labels in rules or as a search filter.
Keeping current and reporting
Changes arrive on a cycle. Forage picks up changes on a regular sweep, not instantly, and can be no more current than Squirrel's own record, which Squirrel updates on its own schedule. The sweep keeps up to date the sites Forage has already read; sites it has not read yet, including new ones, are read as part of reading the estate.
Files moved between sites. A file moved into a site without being changed may not be noticed by the regular sweep. A full re-check of the estate finds it, and is run deliberately rather than as part of the sweep.
Scanned-page failures. Scanned pages that could not be read are not yet listed on the Exceptions page; their documents are counted as partly read.
Archived content, access and audit
Azure's offline archive tier. Where Squirrel has placed an archived file in Azure's offline archive tier, Forage asks Azure to bring it back online before reading it. Azure charges for this, it can take up to a day, and Forage does not return the file to the offline tier afterwards.
Access is all or nothing. There are no narrower roles today. Every Full Administrator who can use the Records section can use every page in it: search across all content, the rules, and the settings.
What the audit log records. The audit log records settings changes, content rule changes made in the portal, and searches, meaning the words used and who searched. It does not record search results, by design, because that would copy sensitive content into the log. Which documents an administrator opens or reads is not yet recorded.
Part 3: not yet available
- Retention rules and retention clocks. A retention rule cannot be saved today.
- Retention clocks started by an event, such as an employee leaving, taken from Entra ID.
- Disposition review.
- Legal hold.
- Destruction. Forage cannot delete a document from SharePoint or the archive.
- Disposal certificates.
- Roles narrower than Full Administrator, such as records manager, reviewer and search operator.
- Limiting a search to a case.
- An audit log protected against change, and an export of the audit log for regulators.
- Recording which documents an administrator opens or reads.
- Exporting search results, including for eDiscovery.
- Creating, renaming and retiring labels in the portal.
- A report of documents holding sensitive content that do not carry a matching sensitivity label.
- Suggested terms for a rule, and a measured precision for each rule.
- Describing pictures that contain no text.
- Recognising identity documents in images.
Questions about how a deployment is secured go to SmiKar (sales@smikar.com).