PDF Unitization
Voir en françaisPDF Unitization
Use this template for repeatable, defensible AI review in Claira.
When you select this template, Claira displays the prompt configuration panel. Map the output to a review field in your case — typically a Memo field, since the output is a multi-line report ending in the unitization string.
What this prompt is for
A single PDF often holds many separate documents stitched together — a scan batch, a collection produced as one file, an email with its attachments flattened. This template identifies where each constituent document begins and ends and returns the page-range string the Nuix Discover image viewer uses to split the file back into individual records. Use it when the file is obviously a stack of distinct records and any competent reviewer would draw the same boundaries. When the split turns on something the pages alone will not tell you — by individual, by record type, by chapter — use PDF Unitization (Custom Split) instead.
Step-by-step in Claira
- Open Prompting > Template picker in your case.
- Select PDF Unitization and map the output to a review field (a Memo field works best).
- Set the Scan as dropdown to Image so Claira reads the actual PDF pages — extracted text does not preserve page boundaries.
- Run the compiled PDF, review the proposed boundaries and any UNCERTAIN entries, then paste the unitization string into the image viewer's split dialog.
The prompt
TASK: PDF UNITIZATION
This PDF is a compilation - several separate documents scanned or merged into one
file. Identify where each constituent document starts and ends and output the page
ranges needed to split them apart. You are drawing boundaries, not summarizing.
Use physical PDF page numbers starting at 1.
A NEW DOCUMENT STARTS WHERE YOU SEE:
- New letterhead, cover page, or title page
- A new email header block (From/To/Date/Subject) not quoted inside a thread
- A new date paired with a new author or addressee
- Pagination restarting, e.g. "Page 1 of 6" after "Page 4 of 4"
- A change in document type or format
- A slip sheet, exhibit divider, or fax cover sheet
DO NOT SPLIT ON:
- A page break inside a continuing document
- A section, article, or chapter heading within the same document
- A signature, notary, schedule, or exhibit page belonging to the document before it
- A continuation page of a table or list
- Repeated letterhead on later pages of the same letter
- An earlier message quoted inside an email thread - one thread is one document
RULES:
- Every page appears in exactly one range. No gaps, no overlaps, no page twice.
- Ranges ascend and cover page 1 through the last page.
- Assign a blank or separator page to the document it introduces, otherwise to the
document before it.
- If no signal on a page starts a new document, it continues the previous one. Never
split on a weak hint - keep the pages together and flag it.
- If the file is one unitary document, say so and return a single range.
OUTPUT:
TOTAL PAGES: [n]
DOCUMENTS:
1. [start-end] | [type] | [date, or "no date"] | [signal that started it]
2. [start-end] | ...
UNITIZATION STRING: [e.g. 1-4; 5-9; 10-22]
UNCERTAIN: [page and reason, or "none"]
Hyphen for a range, comma to join non-contiguous pages in the same document,
semicolon between documents. Account for every page exactly once.Recommended customizations
- If you know what the compilation holds (e.g. a medical chart, a batch of invoices), name the expected document types so boundary signals are easier to recognize.
- If you know roughly how many documents to expect, say so — a large mismatch is a red flag worth surfacing.
- Keep the output contract intact: the UNITIZATION STRING line is the deliverable, and downstream QC depends on the DOCUMENTS table and UNCERTAIN list.
Worked example
Input excerpt
A 22-page PDF: pages 1-4 are a signed engagement letter, pages 5-9 an email thread
printed with its header block, pages 10-22 an invoice batch on repeating letterhead.Expected output shape
A structured report listing total pages, one line per constituent document with its
page range, type, date, and starting signal, then the unitization string
(e.g. 1-4; 5-9; 10-22) and any uncertain pages.Troubleshooting
- If boundaries look wrong, confirm the scan ran in Image mode — text mode loses pagination and the model cannot see page breaks.
- If the string has gaps or repeats a page, re-run and remind the model every page must appear exactly once; Nuix Discover errors if the same page appears twice in one pass.
- If too many pages land in UNCERTAIN, the boundaries may not be self-evident — switch to PDF Unitization (Custom Split) and describe the split you want.
Related
Was this page helpful?
Continue reading
Need more help?
Contact our support team at support@claira.to — we are here to help.