Skip to main content
Table of Contents

Rearranging pages

Delete, move, duplicate and rotate the pages of an opened document, merge other documents into it, and split pages off into a new one.

Once a document is open (see Editing existing documents), its page list is yours to restructure. This is the toolkit behind the everyday PDF chores an application ends up doing for its users: dropping the blank page a scanner produced, turning a sideways page upright, stapling a terms-and-conditions appendix onto every generated quote, splitting one archive into per-customer files. Reach for these methods rather than regenerating the document whenever the content already exists — they carry the original page content across untouched, including fonts and images you have no source for. They are also the only way to grow an opened document: NewPage is a generator method and refuses to run on an opened PDF, so a page you need to create is generated into a separate document and brought in with MergeDocument.

Every one of these methods is a transaction of its own: it either completes or leaves the document exactly as it was. That is per call, not per sequence — if the third of five operations fails, the first two are still applied to the in-memory document. Nothing reaches disk until SaveDocument runs, so abandoning the session always leaves the source file untouched. See The transactional model.

Deleting, moving and duplicating

These three share one rule that causes most of the confusion: indices always refer to the current page order, which the previous operation may have changed. Every index in this guide is zero-based.

Delete, move and duplicate pagesPascal
procedure TForm1.ReorderDocument(const AFileName: string);
var
  p: TTMSFNCPDFLib;
begin
  p := TTMSFNCPDFLib.Create;
  try
    p.OpenDocument(AFileName);
    try
      { Page indices are zero based and always refer to the CURRENT page order,
        so plan a sequence of mutations against the result of the previous one.
        DeletePages takes the indices of the original order in one call, which
        is why it is safer than deleting one page at a time. }
      p.DeletePages([1, 3]);

      { Move what is now the last page to the front. A move destination is the
        FINAL position of the page, and must be in 0..GetPageCount - passing -1
        here raises rather than appending. }
      p.MovePage(p.GetPageCount - 1, 0);

      { Duplicate the cover. A duplicate destination is an INSERTION point, so
        the copy lands in front of the current page 1; pass -1 (the default) to
        append it instead. }
      p.DuplicatePage(0, 1);

      p.SaveDocument(AFileName);
    finally
      p.CloseDocument;
    end;
  finally
    p.Free;
  end;
end;
Method Behaviour
DeletePage(PageIndex) Removes one page. A document must retain at least one page, so deleting the last remaining page fails.
DeletePages(PageIndices) Removes several pages, named against the order before the call — safer than repeated single deletes.
MovePage(PageIndex, DestinationIndex) Moves one page so that it ends up at DestinationIndex in the resulting list.
MovePages(PageIndices, DestinationIndex) Moves several pages as one block to that final position, preserving the order you supplied them in.
DuplicatePage(PageIndex, DestinationIndex = -1) Copies a page and inserts the copy before DestinationIndex. Pass -1 (the default) to append.
DuplicatePages(PageIndices, DestinationIndex = -1) Copies several pages; -1 appends them.

Two things about the destination argument are easy to get backwards:

  • Move takes a final position; duplicate and merge take an insertion point. MovePage(5, 0) puts the page at index 0. DuplicatePage(5, 0) inserts the copy in front of the current page 0.
  • -1 means append for DuplicatePage/DuplicatePages and MergeDocument, but not for MovePage/MovePages. A move destination must be in 0..GetPageCount; anything else raises Invalid PDF page destination index. An out-of-range insertion point raises Invalid PDF insertion index.

The plural forms are not just convenience wrappers. Because they resolve every index against the same starting order, they avoid the classic bug where deleting page 1 makes the page you wanted to delete next shift down by one. They are also markedly cheaper: each call copies the whole in-memory document before mutating it, so deleting ten pages one at a time does ten full copies where DeletePages does one.

Bookmarks follow the pages

Deleting a page also deals with the outline (bookmark) tree that pointed at it, so the document does not keep entries that jump nowhere:

  • A bookmark whose destination was a deleted page is removed, and the sibling chain is rebuilt around what remains.
  • A bookmark that also has surviving children is kept, but loses its destination — it stays as a grouping entry instead of taking the reader to a page that no longer exists.
  • When nothing survives, the outline is dropped from the document catalog altogether.

Merging works the other way round: only pages and their form fields come across, so bookmarks defined in the merged document are not carried into the destination. Rebuild the outline yourself if the assembled document needs one.

Rotating pages

Rotation is stored as an attribute on the page rather than baked into its content, so it is cheap and lossless. RotatePage and RotatePages add to whatever angle the page already carries, clockwise, in multiples of 90 degrees; a negative multiple turns the other way. Anything that is not a multiple of 90 raises PDF page rotation must be a multiple of 90.

Rotate scanned pages uprightPascal
procedure TForm1.MakeAllPagesPortrait(const AFileName: string);
var
  p: TTMSFNCPDFLib;
  Info: TTMSFNCPDFPageInfo;
  Landscape: TArray<Integer>;
  I: Integer;
begin
  p := TTMSFNCPDFLib.Create;
  try
    p.OpenDocument(AFileName);
    try
      SetLength(Landscape, 0);
      for I := 0 to p.GetPageCount - 1 do
      begin
        Info := p.GetDocumentPageInfo(I);

        { Info.Width / Info.Height are the visible size in points, with the
          page's UserUnit scale and the rotation it already carries applied,
          so this test reflects what a reader displays. }
        if Info.Width > Info.Height then
        begin
          SetLength(Landscape, Length(Landscape) + 1);
          Landscape[High(Landscape)] := I;
        end;
      end;

      { One RotatePages call rather than one RotatePage per page: every
        structural method copies the whole in-memory document before mutating
        it, so the plural form is the cheaper route. Degrees must be a multiple
        of 90 and turns clockwise on top of the angle the page already has;
        pass -90 to turn the other way. }
      if Length(Landscape) > 0 then
        p.RotatePages(Landscape, 90);

      { Saving over the file that was opened is safe: it was fully parsed into
        memory and closed before any write happens. }
      p.SaveDocument(AFileName);
    finally
      p.CloseDocument;
    end;
  finally
    p.Free;
  end;
end;

GetDocumentPageInfo reports Width and Height as the visible size in points, with the page's UserUnit scale and its existing rotation already applied, so testing Width > Height tells you what a reader actually displays — which is what you want when straightening a scan. Rotation in the same record gives you the absolute angle, normalized to 0, 90, 180 or 270, if you would rather compute a correction than apply a relative turn. Collect the pages first and rotate them with one RotatePages call, as the snippet does, rather than calling RotatePage inside the loop.

Merging documents

MergeDocument copies pages out of another PDF into the open one — the method behind assembling a quote from a cover, a body and a standard appendix. Its eight overloads accept a file or stream and select all pages, an index array, a PageIndex/PageCount pair, or a TTMSFNCPDFPageRange. Every overload takes an InsertAt position and an optional password for the source document.

Merge documentsPascal
procedure TForm1.BuildPack(const ABaseFile, AAppendFile, ATermsFile,
  ATargetFile: string);
var
  p: TTMSFNCPDFLib;
begin
  p := TTMSFNCPDFLib.Create;
  try
    p.OpenDocument(ABaseFile);
    try
      { InsertAt = -1 appends every page of the other document. }
      p.MergeDocument(AAppendFile, -1);

      { Merge a contiguous zero-based source range. InsertAt is a position in
        the OPEN document and inserts before it, so these two pages become
        pages 1 and 2. }
      p.MergeDocument(ATermsFile, TTMSFNCPDFPageRange.Create(0, 2), 1);

      { Save under a new name so the originals stay untouched. }
      p.SaveDocument(ATargetFile);
    finally
      p.CloseDocument;
    end;
  finally
    p.Free;
  end;
end;

Four details worth internalising:

  • InsertAt = -1 appends; any other value inserts before that position in the open document. InsertAt beyond the current page count raises.
  • The page indices in the selective overloads address the source document being merged in, not the open one.
  • Use TTMSFNCPDFPageRange.Create(FirstPage, PageCount) when a contiguous range is clearer than an index array; both values use zero-based source-page indices.
  • The file overloads re-read the source from disk, so they merge what is saved there — not any unsaved edits you are holding in memory for that file.
  • Merging the same source repeatedly does not multiply the result. On save, objects that are byte-identical are written once and referenced from every page that uses them, so the fonts and images that a set of documents from one template all embed collapse into a single copy. A document assembled from 30 merges of the same source drops from 2,326,030 bytes to 109,066.

Form fields stay interactive

A cloned page brings its widget annotations with it, but a PDF reader only offers a field for input when the document catalog lists it as well — and the AcroForm that does the listing sits beside the page tree, out of reach of the cloned pages. MergeDocument closes that gap: the widgets that came in are collected and their root fields are registered in the receiving document's AcroForm, together with the source's default resources and default appearance string. The result is a merged form the reader still lets you fill in, instead of a picture of one.

Adding new fields to an opened document is a different matter and is not supported: FormFields raises Form fields cannot be added to an opened document, and names the supported route — generate the fields into a new document and merge that in. NewPage reports the same way. Both used to fail silently, so an application written against an older build may be relying on a no-op that now raises.

Merging part of a form

A field can own widgets on several pages — a radio group spread over two pages, for instance. When the merge takes only some of those pages, the field comes across trimmed to the widgets that are actually present, so nothing in the result points at a page that was left behind. Merge every page of a form when you need the whole field intact.

Extracting pages

ExtractPages is the inverse of merging: it returns a brand-new in-memory PDF holding just the pages you name, and leaves the open document untouched. Use it to split an archive into per-customer files, or to hand one chapter to another part of your application without writing the whole document out.

Extract pages into a new documentPascal
procedure TForm1.ExtractChapter(const ASourceFile, ATargetFile: string);
var
  p: TTMSFNCPDFLib;
  Extracted: TMemoryStream;
begin
  p := TTMSFNCPDFLib.Create;
  try
    p.OpenDocument(ASourceFile);
    try
      { ExtractPages returns a NEW in-memory PDF that the caller owns and must
        free. The opened document is left unchanged. }
      Extracted := p.ExtractPages([2, 3, 4]);
      try
        Extracted.SaveToFile(ATargetFile);
      finally
        Extracted.Free;
      end;
    finally
      p.CloseDocument;
    end;
  finally
    p.Free;
  end;
end;

Important

The returned TMemoryStream is created for you and owned by the caller. Free it when you are done, as the snippet does — this is not an internal buffer you are borrowing. It comes back positioned at 0, ready to save or re-open.

The extracted document is a genuinely new PDF, and it gets its own document information: the PDF library as producer and the moment of extraction as the creation date. It does not inherit the source's Title, Author or Subject — those describe the whole document, not the pages you took out of it — and it carries no outline. When the split-off file should carry the metadata anyway, re-open the extracted stream and set the fields before writing it:

Extract pages and carry the metadata overPascal
procedure TForm1.ExtractChapterWithMetadata(const ASourceFile, ATargetFile: string);
var
  p: TTMSFNCPDFLib;
  Extracted: TMemoryStream;
  SourceTitle, SourceAuthor: string;
begin
  p := TTMSFNCPDFLib.Create;
  try
    p.OpenDocument(ASourceFile);
    try
      SourceTitle := p.Title;
      SourceAuthor := p.Author;
      Extracted := p.ExtractPages([2, 3, 4]);
    finally
      p.CloseDocument;
    end;

    try
      { The extracted document carries a FRESH document information
        dictionary - producer and creation date of the extraction - rather
        than the metadata of the source. Re-open it to carry those over. }
      Extracted.Position := 0;
      p.OpenDocument(Extracted);
      try
        p.Title := SourceTitle;
        p.Author := SourceAuthor;
        p.Subject := 'Extracted from ' + ExtractFileName(ASourceFile);
        p.SaveDocument(ATargetFile);
      finally
        p.CloseDocument;
      end;
    finally
      Extracted.Free;
    end;
  finally
    p.Free;
  end;
end;

See Document metadata for the full set of fields.

Putting it together

A realistic assembly job uses every section above at once: merge the parts, drop what should not ship, reorder the result, straighten the landscape pages, and split a summary off — then save once.

Assemble a report from several documentsPascal
procedure TForm1.AssembleReport(const ACoverFile, ABodyFile, AAppendixFile,
  ATargetFile, ASummaryFile: string);
var
  p: TTMSFNCPDFLib;
  Info: TTMSFNCPDFPageInfo;
  Summary: TMemoryStream;
  Landscape: TArray<Integer>;
  BodyLastPage, AppendixFirstPage, I: Integer;
begin
  p := TTMSFNCPDFLib.Create;
  try
    p.OpenDocument(ACoverFile);
    try
      { 1. Append the body, then the appendix. Record the boundaries as they
        become known - appending never renumbers the pages already present, so
        an index captured here stays valid until something is deleted. }
      p.MergeDocument(ABodyFile, -1);
      BodyLastPage := p.GetPageCount - 1;
      p.MergeDocument(AAppendixFile, -1);
      AppendixFirstPage := BodyLastPage + 1;

      { 2. Drop the blank separator page the body document ends with. Deleting
        it shifts every page behind it down by one. }
      p.DeletePage(BodyLastPage);
      Dec(AppendixFirstPage);

      { 3. The appendix opens with its own index page. Move it to final
        position 1, directly behind the cover page. }
      p.MovePage(AppendixFirstPage, 1);

      { 4. Turn any landscape page upright so the report prints consistently,
        in a single RotatePages call. }
      SetLength(Landscape, 0);
      for I := 0 to p.GetPageCount - 1 do
      begin
        Info := p.GetDocumentPageInfo(I);
        if Info.Width > Info.Height then
        begin
          SetLength(Landscape, Length(Landscape) + 1);
          Landscape[High(Landscape)] := I;
        end;
      end;
      if Length(Landscape) > 0 then
        p.RotatePages(Landscape, 90);

      { 5. Split the cover and the appendix index off as a standalone summary
        sheet. ExtractPages leaves the open document untouched and hands back a
        stream the caller owns and must free. }
      Summary := p.ExtractPages([0, 1]);
      try
        Summary.SaveToFile(ASummaryFile);
      finally
        Summary.Free;
      end;

      { Every step above ran on the in-memory document; nothing reaches disk
        until SaveDocument writes the assembled result. }
      p.SaveDocument(ATargetFile);
    finally
      p.CloseDocument;
    end;
  finally
    p.Free;
  end;
end;

Note how the snippet records the page boundaries as it goes. Appending never disturbs the indices of pages that are already there, so an index captured before a merge is still valid afterwards — but a delete in the middle shifts everything behind it, which is why the appendix position is adjusted by hand before it is used again.

Common mistakes

  • Planning indices against the original order. After a delete or a move, the numbering has changed. Either use the plural methods, which resolve every index against the order before the call, or re-derive positions from GetPageCount as you go.
  • Assuming a failed step rolls the whole sequence back. Each call is transactional on its own; the operations that already succeeded stay applied to the in-memory document. Catch the failure and close without saving if you want the original.
  • Deleting every page. A PDF must keep at least one page; the delete that would empty the document raises rather than producing an empty file.
  • Calling single-page methods in a loop. Each one copies the entire document first. Collect the indices and call DeletePages, MovePages, DuplicatePages or RotatePages once.
  • Leaking the stream from ExtractPages. It is yours to free.
  • Expecting an extracted document to inherit the source metadata. It carries fresh document information and no bookmarks; copy Title and Author across yourself if the split-off file needs them.
  • Expecting merged bookmarks. Merging brings pages and their form fields, not the source's outline tree.
  • Passing a rotation that is not a multiple of 90. PDF stores only those four angles.
  • Passing -1 to MovePage. Append is an insertion concept; a move destination must be a real position in 0..GetPageCount.
  • Trying to add a page or a form field to an opened document. NewPage and FormFields raise on an opened PDF. Generate what you need into a separate document and merge it in.
  • Restructuring while a page edit is active. End or cancel the edit first — see Page overlays and underlays.

See also