ClairaClaira Help Desk

PDF Unitization

Voir en français

Use Claira to find the document boundaries inside a compiled PDF, split it apart in the Nuix Discover image viewer, then objectively code the new sub-documents.

PDF Unitization

A single PDF often holds many separate documents stitched together — a scan batch, a collection produced as one file, an email with its attachments flattened. Unitization is the process of identifying where each constituent document begins and ends so the file can be split back into individual records.

In Nuix Discover's image viewer, the split is driven by a page range string. This workflow uses Claira to draw the boundaries and produce that string, applies the split in the image viewer, and then runs objective coding over the newly created sub-documents — because they are born with almost no metadata of their own.

When to use this workflow

  • Scan batches. Paper collections scanned as one continuous PDF per box or folder.
  • Compiled productions. A party produced many records as a single file.
  • Flattened families. Emails printed to PDF with their attachments appended.
  • Oversized records. A chart, ledger, or file that reviewers need split by person, date, or record type to be reviewable.

The workflow

Step 1: Find the compiled PDFs

Identify the documents that need unitization. High page counts, generic titles (Scan_001.pdf), and documents whose first pages look nothing like their last pages are the usual tells. Tag them so you can work through the set methodically.

Step 2: Run a PDF Unitization template

Open the document in single review, then pick the template that fits:

  • PDF Unitization — the boundaries are self-evident: any competent reviewer would draw the same lines.
  • PDF Unitization (Custom Split) — the split turns on something the pages alone will not tell you. Replace the placeholder block with one or two plain sentences, e.g. "One document per invoice." or "One document per named individual, all of that person's records together."
Set the Scan as dropdown to Image so Claira reads the actual PDF pages. Extracted text does not preserve page boundaries, and unitization is entirely about page boundaries.

Claira returns a structured report: total pages, one line per constituent document (page range, type, date, starting signal), the unitization string, and a list of uncertain pages.

Step 3: QC the proposed boundaries

Before splitting anything, check Claira's work:

  • Resolve every UNCERTAIN entry. Open those pages in the viewer and decide where they belong.
  • Spot-check the boundary pages. For each proposed range, look at its first and last page — boundary errors live at the edges, not the middles.
  • Verify coverage. The ranges must account for every page exactly once — no gaps, no overlaps. Nuix Discover errors if the same page appears twice in one unitization pass.

Step 4: Split the file in the image viewer

Open the document in the Nuix Discover image viewer and start a unitization. Paste Claira's string, using the standard syntax:

  • Hyphen for a contiguous range: 5-9
  • Comma to join non-contiguous pages into the same output document: 1-3, 7
  • Semicolon between output documents: 1-4; 5-9; 10-22
  • A single-page document is written as a bare number
  • Page order within a set is honoured, so 7-5 collates in reverse
Example unitization string

1-4; 5-9; 10-22

Step 5: Objectively code the new sub-documents

The new documents do not inherit metadata from the source file — only Document Date is populated, set to the date of the unitization. Every other field starts empty, which is exactly what the Objective Coding workflow is for.

Collect the newly created documents and run objective coding over them:

  • Use Multi-Code to fill date, title, author, and document type in a single pass, or run the individual templates — Extract dates, Extract author, Identify document type.
  • Write to dedicated AI fields first (e.g. OC Date, OC Author) rather than your primary metadata fields, so you can QC before committing.
  • A sub-document that was image-only in the source is still image-only after the split — scan it as Image here too.

Step 6: QC and promote

Spot-check the coded values against the sub-documents, resolve fallback values, then copy the QC'd results from the AI fields into your primary metadata fields. From this point the sub-documents behave like any other records in the review.

Best practices

  • One compiled PDF at a time. Unitization is a per-file decision; run the template in single review rather than as a bulk scan, and keep the report next to the document in a Memo field.
  • Split first, code second. Coding the compiled file wastes effort — the values you extract belong to the sub-documents, not the container.
  • Keep the source document. Do not delete the compiled original; it is your audit trail for where every sub-document came from.
  • Escalate to Custom Split deliberately. If the self-evident version keeps flagging pages as uncertain, that is your cue that the split needs an instruction, not a re-run.

Limitations

  • Claira proposes; the image viewer disposes. Claira outputs the page-range string, but the split itself happens in Nuix Discover's image viewer — review the string before applying it.
  • Page numbers are physical. The ranges count PDF pages starting at 1, not Bates numbers or printed page numbers.
  • Metadata is not inherited. Only Document Date is populated on the new documents (set to the unitization date) — plan for the objective coding pass, not against it.
  • Very long files stretch any model. For files with hundreds of pages, consider splitting in stages: a coarse first pass, then unitize the larger chunks again.

Need help? Contact support@claira.to

Was this page helpful?

Need more help?

Contact our support team at support@claira.to — we are here to help.