Skip to content
SmiKar Software

Nutshell AI Summarisation Modes

10 min read · Last updated · Page version 11

Nutshell AI offers five summary types. The type controls two things: how much of a document is read, and how the summary is written - including whether the AI writes it at all. It is set per-tenant on the AI Processing Settings page in the Squirrel admin portal.

The five types

Listed most thorough to fastest.

TypeUses AIReadsWhat the stub containsBest for
DetailedYesEvery pageComprehensive write-up with the key points, context and supporting detailSmall, high-value repositories where reading depth matters more than speed
StandardYesThe opening, a middle section and the closing pagesBalanced summary of the main points and essential informationThe default for most tenants
BriefYesThe opening pages onlyConcise overview of only the most critical informationLarge repositories where a short read is enough
LiteYes, one short pass on most documentsEvery pageKey passages selected automatically from across the whole document, then rewritten by the AI - kept under the document's own headings where it has them, otherwise three or four paragraphs - followed by a line of the dates, amounts and reference codes found in itVolume archives: well under half of Brief's processing per document, so many more documents an hour, while still reading as written prose
ExtractNoEvery pageThe same automatically selected passages, shown word for word as they appear in the document and grouped under its own headings, plus the dates and amounts lineMaximum throughput, where policy says no AI, or where exact figures and reference codes matter more than readable prose

Detailed sends the whole document to the AI. Reading a document and giving all of it to the AI are two different steps. In Detailed they are the same: the relevance filter that sits between them runs in boilerplate-only form, stripping disclaimers, confidentiality notices and other standing text but dropping no content, so every sentence of the document itself reaches the model. The lighter types do narrow by relevance on the way in, which is part of how they stay fast.

A long Detailed summary reads as one piece. Detailed assembles its result from several passes over the document. Only the first of those introduces the document, a later part that reintroduces it has that framing removed while keeping its facts, and a point an earlier part already made is dropped before the parts are joined. The detail itself is never compressed - not compressing it is what the type is for.

Detailed, Lite and Extract read every page of a paged document, including scanned pages, which are read by OCR in whatever language they are written in - see supported file types. Where a scan is too poor for OCR to read reliably, every type says so in the stub rather than summarising noise. Brief reads the opening pages only, and Standard reads those plus a slice from the middle and the closing pages. Longer un-paged documents are narrowed by relevance instead, keeping proportionally more of the text in the more thorough types.

The difference between the two groups is worth understanding, because it is not the same axis as speed. Lite and Extract are fast despite reading everything: they pick the passages that matter and then do little or no generation, where Brief is fast because it reads less. So a long document can be better served by Lite than by Brief, even though Brief sits above it in the table.

Throughput figures are measured against a ~7,500-word, 1 MB Word document. Real numbers vary with file size, file type, and SharePoint API throttling.

The two that are not a normal summary

Lite and Extract work the same way up to the last step. Both read the whole document and pick out its key passages automatically. Where the document has real structure - numbered sections, chapters, appendices - both keep those headings and organise the result under them. Both finish with a line of the dates, amounts and reference codes found in the document, which is often the fastest way to recognise a file you are looking for.

Not every document gets an outline. Where the text is mostly short label-and-value lines rather than sentences - fill-in forms, bilingual invoices and the like - there is no meaningful heading structure to follow, so both types skip the outline and return the passages plainly.

A spreadsheet is handled differently again in every type: it is profiled rather than read as prose - row and column counts, what each column holds, the values that dominate it - and in Extract that profile is the whole summary. See spreadsheets.

The difference is the final step. Lite hands those passages to the AI to be rewritten as prose: one to three sentences under each heading, or three to four paragraphs where the document has no headings. Extract skips that step entirely and shows the passages word for word as they appear in the document.

On most documents Lite is a single AI call, which is where its speed comes from. A long document is too much to restate in one reply, so it is split on its own headings and polished in several passes, then joined - about one document in seven takes more than one call. The result is the same shape either way; only the processing cost differs.

Extract is the one to choose where AI is not permitted at all, or where throughput matters more than readability. Be aware that it reads as excerpts rather than a summary, and that scanned or table-heavy documents look rougher than they would under any of the AI types.

Where the exact figures matter, choose Extract over Lite. This is the one place the two types genuinely diverge. Extract reproduces passages word for word, so a part number, a torque figure or a standard reference arrives exactly as the document wrote it. Lite asks the model to rewrite the same passages, and a model paraphrasing technical text drops reference codes - measured across a set of engineering and contract documents, Extract carried about 71% of their distinct facts into the stub against Lite's 21%. Lite reads far better; Extract is more faithful. Pick on which of those the summary is for.

How much text an extract keeps scales with the document. There is no fixed cap. A longer document keeps proportionally more, up to a ceiling, and the space is spent on the lines a reader needs rather than simply the opening ones - a line carrying a date, an amount, a percentage, a reference code, or obligation language such as shall, must, warranty, penalty or deadline is kept ahead of generic prose, and the passages are still shown in document order. On a 940,000-character equipment manual that means about 3% of the text carrying roughly three quarters of its facts.

The temperature slider is greyed out while Extract is selected, because that type never calls the AI and there is nothing for temperature to affect.

Choosing a type

Changing the type applies from that point on. Documents already summarised are not redone, so switching type does not rewrite the stubs you already have. A common pattern is to switch to Lite or Extract while a very large site is being onboarded, then back to an AI type once the backlog has drained.

  • Brief is enough when the goal is discovery - someone searches, finds a candidate document via its summary, and decides whether to restore it. The opening slice of a document usually contains the executive summary, abstract, or introduction, so Brief mode captures the "what is this about" answer at the highest throughput.
  • Standard is the sensible default for most tenants. Reading three slices catches the closing recommendations, action items, or conclusions that a Brief scan would miss on longer documents.
  • Detailed is for content where every clause matters - contracts, policies, technical specifications, records with retention obligations. Detailed reads everything, so a summary produced by Detailed can be treated as an authoritative précis of the original.

Mode is per-tenant, but Nutshell will re-summarise on demand, so it is reasonable to start with Standard and move to Detailed selectively if you find summaries missing important tail content.

Per-worker throughput and scaling

A worker is a GPU-backed processing engine. Nutshell licensing scales by worker - capacity is added by adding workers, and throughput is reported per worker. Scaling is roughly linear: doubling the worker count roughly doubles files-per-minute.

For example, at Standard mode:

  • 1 worker → ~2,400 files / hour
  • 4 workers → ~9,600 files / hour
  • 10 workers → ~24,000 files / hour

Real throughput is bounded by SharePoint API rate limits and Azure Blob read throughput; adding workers beyond that ceiling produces diminishing returns.

Intelligent resource management

Each Nutshell worker monitors its own CPU and GPU load in real time and adjusts how many documents it processes concurrently:

  • Under heavier load - Nutshell reduces per-worker concurrency to stay stable and avoid thermal throttling.
  • When spare capacity is available - Nutshell scales concurrency back up automatically.

The effect is consistent throughput across large batches without manual tuning. Operators do not need to size worker concurrency by hand.

Large backlogs drain faster. When a backlog builds up, Nutshell adds extra upload workers to clear it. The ceiling on that surge capacity is set high enough to keep scaling on backlogs of several hundred files, so a large catch-up run - after an onboarding pass or an outage - clears at a sustained rate without anyone intervening.

What summaries leave out

Two rules apply in every mode:

  • Boilerplate is omitted, not paraphrased. Disclaimers, forward-looking-statement notices, and confidentiality, copyright or trademark notices are dropped from the summary entirely. They are not material facts about the content, and a summary that ends by noting a document "includes a disclaimer regarding forward-looking statements" has spent its closing sentence on nothing.
  • No self-references. Summaries state what the content says directly, rather than describing the file - so you get "the migration completes in March" rather than "this document explains that the migration completes in March".

Sample output per type

Side-by-side sample summaries show what each mode actually produces from the same source document. Reading them is the fastest way to decide which mode to configure.

See also


Need help? support@smikar.com.

More in Squirrel

See all pages →