Nutshell AI Supported File Types
6 min read · Last updated · Page version 5
Nutshell AI summarises the formats below. Files outside this list are still archived by Squirrel; they just do not receive a Nutshell summary. Per-extension processing can be toggled on or off from the File Processing & Security page in the Squirrel admin portal.
Supported formats
| Category | Extensions |
|---|---|
.pdf | |
| Word | .docx, .dotx, .dotm |
| Excel | .xlsx, .xlsm, .xltx, .xltm |
| PowerPoint | .pptx, .pptm, .potx, .potm, .ppsx, .ppsm |
| Plain text | .txt, .text |
| Legacy Office | .doc, .xls, .ppt, .csv, .rtf |
Modern formats (the Open XML .*x / .*m family) are the fastest to process - Nutshell reads their text directly. Legacy Office formats (.doc, .xls, .ppt, .csv, .rtf) are fully supported but slower because they require an additional parsing step before summarisation.
How each file type is read
- Word - full text body. Fastest and most accurate to summarise.
- Excel - read as a table rather than as prose. See Spreadsheets below.
- PowerPoint - visible text on slides, slide titles, and speaker notes. The summary captures the flow and themes of the deck.
- PDF - text content where the PDF has a text layer, and OCR where it does not. See Scanned PDFs below.
- Text and legacy Office - straightforward reads. Legacy formats add conversion overhead but produce the same quality of summary as their modern equivalents.
Spreadsheets
A spreadsheet is not prose, and summarising it as though it were produces a paragraph that says less than the file does. Nutshell reads a workbook and profiles it instead.
For each sheet it reports the real sheet name, how many rows and columns it holds, what each column contains, which values dominate a column and how often, and the range the numbers cover. Those figures are counted from the file, not generated, so they are right regardless of how the AI writes them up. One AI pass then turns the profile into readable prose and the counted facts sit underneath it.
In Extract mode, where no AI runs at all, the profile is the summary.
This is also why a spreadsheet summary is quick. Reading a 1,700-row permissions export used to mean dozens of AI passes producing near-identical paragraphs, because every slice of a table looks like every other slice. It is now a single pass, and the result is specific where the file is specific - which principal types, how many distinct users, which groups - rather than "a list of users and their permissions".
Formulas are passed over rather than read as values, and a row keeps its shape, so a header row still means something and a value stays under the right column.
Very large spreadsheets
A workbook can hold far more data than any summary could describe - a system export running to tens of millions of cells is not unusual. Nutshell reads such a workbook up to a fixed budget rather than end to end, so the stub describes the opening portion of its largest sheets rather than every row.
When that happens the summary says so. A sample is never presented as though it were the whole file.
This bounds the summary only. The workbook itself is archived intact and comes back whole on restore.
Scanned PDFs
A scanned or image-only PDF has no text layer, so Nutshell reads it with OCR. This is automatic - there is nothing to configure, and no separate OCR licence to buy. It costs roughly a second per document on top of the normal read.
Scans are read in their own language. Nutshell detects the writing system a scan uses - Latin, Cyrillic, Greek, Japanese, Chinese, Korean and the rest - and reads it with the matching model, rather than assuming English. This matters more than it sounds: read with the wrong model, a Russian contract came back as Tpunoxenne 2.9 / Koutpaxt Ne 74-E-38-1 and a Japanese patent as {Oevt4 Sid, ARMS OAM. Those were not bad scans. Latin-script documents benefit too, because the Latin model covers accented characters that the English one truncates.
A scan too poor to read says so. OCR reports how confident it was, per word. Where that confidence is low across a document, the stub states plainly that the document is a scan whose text could not be read reliably enough to summarise, and lists the lines that did read clearly - typically the letterhead, dates and reference numbers, which is usually enough to recognise the file. That is deliberate: the alternative is a stub full of noise, or worse, fluent AI prose invented on top of noise and indistinguishable from a real summary.
A PDF that does carry a text layer, but a corrupted one, is detected and sent to OCR as well rather than being summarised from the garbled text.
Files outside the supported list
Files that are not in the supported list - images, video, audio, executables, custom binary formats - are still archived by Squirrel and can still be restored on demand. They just do not get a summary written back to their SharePoint stub. Search and Copilot fall back to filename and site metadata for those, which is the same behaviour you get without Nutshell.
If a file type your organisation relies on is not on the list, contact sales@smikar.com - the format catalog is extended based on customer demand.
Toggling per-extension processing
The File Processing & Security page shows every supported extension with a per-extension toggle. Turning an extension off tells Nutshell to skip files with that extension - useful when a particular format in your tenant is high-volume and low-signal (for example, some organisations skip .csv because summaries of raw data extracts add limited value).
Toggling an extension does not affect Squirrel's archive behaviour - files with that extension continue to archive normally. Only summarisation is skipped.
See also
- Summarisation modes - coverage settings that interact with file size and type.
- How Nutshell works - where the extraction step reads each format.
- Redaction - how sensitive fields are stripped before the summary is written.
Need help? support@smikar.com.