Forage: What Is Built and What Is Not
5 min read · Last updated · Page version 3
Forage is in development as a whole, not only the parts listed further down as not built. What appears under Built and in testing is working and is being tested with customers; it is not generally available.
There are no dates on this page, and no delivery order.
Built and in testing
- Estate inventory. Forage builds its inventory of SharePoint Online files and Squirrel-archived files from the records Squirrel already keeps, so SharePoint is not crawled a second time.
- Reading the archive. Forage reads the content Squirrel has archived, opening the encrypted, compressed archive copies inside your own deployment with your own archive key, and reads the current version of each file.
- Document types. PDF; Word, Excel and PowerPoint files, including templates and macro-enabled files; older Office formats; Outlook and standard email messages; and rich text, plain text, CSV, HTML and XML, without needing Microsoft Office.
- Scanned documents. Scanned PDFs and image files are read with optical character recognition, so scanned contracts and forms become searchable.
- A record of how each document was read. For every document it reads, Forage records how it was read, page by page, and keeps a cryptographic fingerprint of the file.
- Coverage accounting. Every file in the estate is accounted for, whether read, excluded by type or by policy, or failed with a recorded reason.
- Search. One full-text search across live and archived content, with highlighted extracts, filters by source and document type, English word-stem matching, and each file's name and location on every result.
- Keeping current. Forage keeps its picture of the estate current by asking Squirrel what was added, changed or deleted, and reads a document again only if its content really changed.
- Content classification, where it is enabled on a deployment. Forage recognises what documents contain, such as bank details, payment card numbers, tax and company identifiers, passport numbers, health information and contract terms, using transparent rules. Every finding keeps the words that produced it.
- Writing and testing rules, where content classification is enabled. Administrators write, version, test and switch on content rules in the portal, and a dry run before a rule goes live shows the most it could match, a sample checked with the real rule, and a projected range across the documents read so far.
- Labels and the documents behind them, where content classification is enabled. Administrators see, label by label, how many documents contain each kind of content, how serious it is and in which sites, and can open the documents behind any label or rule and read the matched text in context.
- Labels Squirrel recorded, on results. Search results show the retention label and the sensitivity label that Squirrel recorded for each file.
- What could not be read. Forage shows what could not be read, grouped by cause and ranked by the documents affected, and says whether the cause lies with the file itself, the source, the infrastructure or Forage, and whether it will be retried or needs someone to act.
- Choosing what is read. Administrators choose live content, archived content or both, and which document types are read. Every change is recorded with who made it and when.
- Progress reporting. The Dashboard and Coverage pages show how much of the estate is searchable, which sites have been read, whether read sites are being kept current, the processing rate, and an estimated time to finish at the rate actually achieved.
- Audit log. Forage records settings changes, content rule changes made in the portal, and searches, and records how each document was read.
- Per-customer deployment. Each customer has its own Forage deployment, with its own database and its own search index. There is no shared index and no path for one customer's query to reach another customer's data.
- The source is left alone. Forage only reads SharePoint, the archive and Squirrel's records. It does not change, move or delete any source document, and it writes nothing to Squirrel.
- No external AI service. Forage sends no document content to any external AI service. Its content classification uses no AI model. Scanned pages are read by optical character recognition models that run inside the deployment.
- Nothing is deleted. Forage does not delete anything.
- Recoverable index. Forage's search index can be rebuilt from the text Forage already holds, without reading SharePoint or the archive again.
- Forage keeps a copy of the text. Forage keeps the text it reads, and a search index of that text, inside your deployment.
Content classification is not on every deployment, and a new deployment does not necessarily start with it or with starter labels. Confirm with SmiKar what a given deployment includes.
Being extended
- Stronger reading of scans on deployments with a GPU. Forage is being extended to read scanned pages with a stronger optical character recognition model on deployments with a GPU, including rotated pages and languages beyond English.
In development, not available yet
- Retention rules and retention clocks, including clocks that start from an event such as an employee leaving, taken from Entra ID.
- Disposition review, where a person decides on every candidate for disposal.
- Legal hold.
- Destruction, with a fresh check of each document immediately before anything is removed.
- Disposal certificates.
- An audit log protected against change, and its export for regulators.
- Roles narrower than Full Administrator, and limiting a search to a case.
- eDiscovery export.
- Creating and editing labels in the portal, and a report of documents found to hold sensitive content that do not carry a matching sensitivity label.
- Help with tuning rules: suggested terms, and a measured precision for each rule.
- Describing images that contain no text.
- Recognising identity documents in images.
See limitations for how today's features behave at their edges.
Questions go to sales@smikar.com.