Editing existing documents
Open a PDF that already exists, inspect its pages, read its text, and write it back out — without a third-party PDF library.
TTMSFNCPDFLib has always been able to generate a PDF from scratch through
BeginDocument and Graphics. It can now also
open one. OpenDocument parses an existing
file into an in-memory document you can inspect and restructure, and
SaveDocument writes the result back. The parser is built into TMS FNC Core, so
no external PDF engine, DLL or command-line tool is involved on any supported
platform. Reach for this whenever the PDF you need to work with was produced
somewhere else — a scanner, a bank statement, a supplier invoice, an archived
report — and generating a fresh document is not an option. When you are building
a document from your own data, keep using the generation API instead: it is
simpler and gives you full control over the layout.
This page covers opening, inspecting, reading and writing metadata, and saving. Restructuring the page order is covered in Rearranging pages, and drawing on existing pages in Page overlays and underlays.
Opening and closing a document
Every editing session follows the same shape: open, work, save, close. Opening
parses the whole document up front, so wrap the work in try ... finally and
close in the finally — an open document holds parsed structures until it is
closed or the instance is freed.
function TForm1.CountPages(const AFileName: string): Integer;
var
p: TTMSFNCPDFLib;
begin
Result := 0;
p := TTMSFNCPDFLib.Create;
try
p.OpenDocument(AFileName);
try
if p.IsDocumentOpened then
Result := p.GetPageCount;
finally
p.CloseDocument;
end;
finally
p.Free;
end;
end;
IsDocumentOpened reports whether a document — generated or parsed — is
currently active. CloseDocument discards the in-memory document; anything not
saved is lost, which is exactly what you want when the user cancels.
A stream overload takes the document from the current stream position, which is what you need for a document that arrives over HTTP or out of a database blob:
PDF.OpenDocument(BlobStream);
Both overloads take an optional trailing Password argument. It is reserved for
future encrypted-document support and has no effect today: an encrypted PDF
cannot be opened at all, with or without a password (see
Handling failures). On TMS WEB Core only the stream
overloads of OpenDocument and SaveDocument are available, because that target
has no local file system.
Opening also replaces whatever was active before: any document already open is released, and an unfinished page edit is silently cancelled rather than raising. Finish your work on one document before opening the next.
The transactional model
Opening, saving and every structural mutation are transactional: they either complete or leave the document as it was. That is what makes a multi-step edit safe — a merge that fails halfway does not leave you with half a document.
Three consequences worth remembering:
- Nothing reaches disk until
SaveDocumentruns. Deleting, moving, rotating and merging all act on the in-memory document. Abandon the session and the source file on disk is untouched. - Saving over the file you opened is safe. The save writes a temporary sibling file first, renames the original out of the way, then swaps the new file into place and deletes the backup; if any step fails, the original is put back. So the target folder has to be writable and needs room for a second copy of the document while the save runs. Saving under a new name is still the better habit while you are developing an edit sequence.
- Saving compacts the document. Objects that became identical — typically shared fonts and resources pulled in by a merge or an overlay — are collapsed into a single object, and unreachable objects are dropped. An edited document is therefore usually smaller than the sum of its inputs rather than larger.
The stream overload of SaveDocument replaces the contents of the stream it is
given rather than appending, so pass a stream you own. It leaves the stream
positioned at 0, ready to be read back or sent on.
Reading page geometry
Before you can lay anything out on an existing page you need to know how big it
is — and a PDF page can lie about that in three different ways at once: it can
carry a crop box smaller than its media box, a rotation, and a user-unit scale.
GetDocumentPageInfo resolves all three into
one record.
procedure TForm1.ListPageGeometry(const AFileName: string);
var
p: TTMSFNCPDFLib;
Info: TTMSFNCPDFPageInfo;
I: Integer;
Orientation: string;
begin
p := TTMSFNCPDFLib.Create;
try
p.OpenDocument(AFileName);
try
for I := 0 to p.GetPageCount - 1 do
begin
Info := p.GetDocumentPageInfo(I);
{ Width and Height are the VISIBLE size in points: the crop box with
UserUnit scaling and page rotation already applied. Use them for
layout; use MediaBox / CropBox when you need the raw boxes. }
if Info.Width > Info.Height then
Orientation := 'landscape'
else
Orientation := 'portrait';
Memo1.Lines.Add(Format('Page %d: %.0f x %.0f pt (%s), rotation %d, user unit %.2f',
[I + 1, Info.Width, Info.Height, Orientation, Info.Rotation, Info.UserUnit]));
end;
finally
p.CloseDocument;
end;
finally
p.Free;
end;
end;
TTMSFNCPDFPageInfo exposes:
| Field | Meaning |
|---|---|
MediaBox |
The full page extent, in unscaled PDF user-space coordinates. Inherited from the page tree when the page does not set it. |
CropBox |
The region a reader actually displays, in the same coordinates. Also inheritable, and equal to MediaBox when the document does not declare one. |
Width, Height |
The visible size in physical points: the crop box, with UserUnit and Rotation already applied. These are the numbers to lay out against. |
Rotation |
The inherited clockwise rotation, normalized to 0, 90, 180 or 270. |
UserUnit |
The page scale factor; 1 for virtually every document, larger for oversized drawings. |
Use Width/Height for positioning and the boxes only when you specifically care
about the raw geometry. Because a 90 or 270 degree rotation swaps the two,
Width can exceed Height on a page whose crop box is portrait.
Reading page content and text
Two methods read a page back, and which one you want depends on whether you are
after words or operators. GetDocumentPageText returns what a person would read
— for search, indexing or matching an incoming document against a record.
GetDocumentPageContent returns the raw decoded content stream for when you need
the drawing operators themselves.
function TForm1.FindPagesContaining(const AFileName, ATerm: string): TArray<Integer>;
var
p: TTMSFNCPDFLib;
I: Integer;
Hits: TList<Integer>;
begin
Hits := TList<Integer>.Create;
try
p := TTMSFNCPDFLib.Create;
try
p.OpenDocument(AFileName);
try
for I := 0 to p.GetPageCount - 1 do
{ GetDocumentPageText decodes through each font's /Encoding and, when
present, its /ToUnicode CMap, so embedded subset fonts and older
documents alike read back correctly. /ActualText spans win over the
glyph codes, lines are decided from vertical position on the page,
and form XObjects are followed, so overlay and underlay text is
included. }
if ContainsText(p.GetDocumentPageText(I), ATerm) then
Hits.Add(I);
finally
p.CloseDocument;
end;
finally
p.Free;
end;
Result := Hits.ToArray;
finally
Hits.Free;
end;
end;
The two differ in more than formatting:
GetDocumentPageText |
GetDocumentPageContent |
|
|---|---|---|
| Returns | Text, one line per text line, separated by line feeds | The decoded page content stream, as TBytes |
| Font encoding | Decoded through each font's /ToUnicode CMap, falling back to the font's /Encoding — WinAnsiEncoding, MacRomanEncoding and /Differences glyph names — so both embedded subset fonts and older documents that carry no CMap read back correctly |
Raw — character codes are whatever the page uses |
| Form XObjects | Followed, so overlay and underlay text is included | Not included; the page references them rather than containing them |
That last row matters if you stamp documents: text you added with
BeginPageEdit lives in a form XObject, so it shows up in
GetDocumentPageText but not in GetDocumentPageContent.
What text extraction resolves for you
Three details decide whether extracted text is usable, and all three are resolved without any setting on your side:
- Encoding. Simple fonts are decoded through
/Encoding, including theWinAnsiEncodingandMacRomanEncodingbase encodings and any/Differencesglyph-name overrides; a/ToUnicodeCMap, when the font carries one, takes precedence over that. Composite (Type0) fonts are decoded two bytes at a time, so documents in Chinese, Japanese, Thai, Lao or Telugu come back as real text rather than glyph indices. /ActualText. When a marked-content span declares the text it really represents, that value is used instead of the glyph codes. This is how ligatures, decorative glyph substitutions and split words in well-produced documents read back correctly.- Line breaks. Lines are decided from vertical position on the page, not from where the producer happened to end a text-showing operator. A generator that emits one operator per word — or per glyph — still yields one line per visual line, and a diacritic drawn separately stays with its base letter instead of landing on a line of its own.
Note
Within a line, text comes back in the order the page draws it, which is not always reading order: a two-column layout or a table can still interleave. Treat the result as searchable content, not as a faithful transcript.
Document metadata
Retitling an archived file, or recording who approved an incoming invoice, should
not require regenerating the document — and it does not. The document information
fields follow whichever document is active: while a document is open, Title,
Author, Subject, Creator and Keywords — the same properties you set before
EndDocument when generating — read and write the metadata of that document,
and Producer, CreationDate and ModificationDate join them.
procedure TForm1.ShowDocumentMetadata(const AFileName: string);
var
p: TTMSFNCPDFLib;
Created, Modified: TDateTime;
begin
p := TTMSFNCPDFLib.Create;
try
p.OpenDocument(AFileName);
try
{ While a document is open, the metadata properties address THAT
document instead of the generator. }
Memo1.Lines.Add('Title : ' + p.Title);
Memo1.Lines.Add('Author : ' + p.Author);
Memo1.Lines.Add('Subject : ' + p.Subject);
Memo1.Lines.Add('Creator : ' + p.Creator);
Memo1.Lines.Add('Producer : ' + p.Producer);
{ Keywords is a TStrings. Opening splits the single PDF keyword string
on spaces, tabs, commas and semicolons. }
Memo1.Lines.Add('Keywords : ' + p.Keywords.CommaText);
{ CreationDate and ModificationDate are the raw PDF date strings, for
example D:20260908120000+02'00'. TryPDFDateToDateTime, declared in
FMX.TMSFNCPDFCoreLibBase, converts them and returns False instead of
raising on a date it cannot parse. A timezone offset is normalized
away, so the TDateTime it returns is UTC. }
if TryPDFDateToDateTime(p.CreationDate, Created) then
Memo1.Lines.Add('Created : ' + DateTimeToStr(Created))
else
Memo1.Lines.Add('Created : ' + p.CreationDate);
if TryPDFDateToDateTime(p.ModificationDate, Modified) then
Memo1.Lines.Add('Modified : ' + DateTimeToStr(Modified))
else
Memo1.Lines.Add('Modified : never');
finally
p.CloseDocument;
end;
finally
p.Free;
end;
end;
| Property | On an opened document |
|---|---|
Title, Author, Subject, Creator |
Read and written directly. |
Keywords |
A TStrings. Opening splits the document's single keyword string on spaces, tabs, commas and semicolons; saving joins it back with spaces. |
Producer |
The name of the writing software. Read and written on an opened document; while generating, the library stamps its own value and writing it has no effect. |
CreationDate |
The raw PDF date string. Read and written on an opened document; the generator stamps its own value. |
ModificationDate |
The raw PDF date string. SaveDocument stamps it whenever the document changed, so an explicit value is overwritten. |
Producer, CreationDate and ModificationDate are inert while you are
generating — assigning them only has an effect on an opened document. The other
four work in both modes.
PDF dates are strings, not TDateTime
CreationDate and ModificationDate are the PDF date strings exactly as the
document carries them: D:20260908120000Z from this library, or something like
D:20260908120000+02'00' from another producer, with the timezone suffix
optional and frequently absent altogether.
TryPDFDateToDateTime, declared in FMX.TMSFNCPDFCoreLibBase, converts one:
function TryPDFDateToDateTime(const AValue: string;
out ADateTime: TDateTime): Boolean;
It returns False instead of raising when the string cannot be parsed — worth
using in that shape, because a malformed or missing creation date is one of the
most common defects in documents produced by older tools. Two details to plan
around:
- It is lenient about shape. The
D:prefix is optional, and a string truncated after the year, month or day still parses, with the missing components defaulting to the first of the month, midnight, and so on. A string shorter than four digits, or one carrying an impossible date, fails. - It normalizes a timezone offset away. When the string ends in
+02'00'or-05'00', the offset is subtracted, so theTDateTimeyou get back is UTC — not the local wall-clock time the document recorded. Convert to local time yourself if that is what you want to display.
Writing metadata back
procedure TForm1.TagApprovedInvoice(const AFileName, ATitle, AApprover: string);
var
p: TTMSFNCPDFLib;
begin
p := TTMSFNCPDFLib.Create;
try
p.OpenDocument(AFileName);
try
p.Title := ATitle;
p.Author := AApprover;
p.Subject := 'Approved for payment';
{ Producer identifies the writing software. On an opened document it is
yours to set; while GENERATING, the library stamps its own value and
writing it has no effect. }
p.Producer := 'Invoice Desk 3.2';
{ Keywords can be edited in place. SaveDocument flattens the list into
the document keyword string, so the edit does not need a separate
assignment to take effect. }
p.Keywords.BeginUpdate;
try
p.Keywords.Clear;
p.Keywords.Add('invoice');
p.Keywords.Add('approved');
p.Keywords.Add(FormatDateTime('yyyy-mm', Date));
finally
p.Keywords.EndUpdate;
end;
{ The metadata edits above mark the document as changed, so SaveDocument
stamps a fresh ModificationDate itself. Setting that property by hand
is pointless - the stamp overwrites it. Assigning a property the value
it already holds is ignored and does NOT trigger the stamp. }
p.SaveDocument(AFileName);
finally
p.CloseDocument;
end;
finally
p.Free;
end;
end;
Three behaviours are worth planning around:
- Editing metadata counts as changing the document. Setting
TitleorAuthormarks the document modified just as deleting a page would, so the save writes a new file and stamps a freshModificationDate. ModificationDateis stamped on save — but only when the document actually changed. Assigning a property the value it already holds is ignored, so simply passing an archive through your application, or writing back metadata that matches what was there, does not make every file look freshly edited.- Keyword edits are flattened at save time.
Keywordsis a list you can mutate in place;SaveDocumentjoins it into the document's keyword string at that point, so no separate assignment is needed to commit the change.
Handling failures
Anything the structural editor refuses to do reports why, rather than failing
silently — so the handler you write once at the open site covers the whole
session. A document that cannot be parsed, an encrypted file, a page index that
does not exist, and a save that cannot be written all raise
ETMSFNCPDFDocument, whose Error property classifies the failure so you can
tell the user something more useful than "the operation failed".
function TForm1.TryOpen(APDF: TTMSFNCPDFLib; const AFileName: string;
out AMessage: string): Boolean;
begin
Result := True;
AMessage := '';
try
APDF.OpenDocument(AFileName);
except
on E: ETMSFNCPDFDocument do
begin
Result := False;
case E.Error of
{ Encrypted documents cannot be opened at present. The Password
parameter of OpenDocument is reserved for future support, so do
not prompt the user for one: supplying a password only turns this
into pdeUnsupportedFeature. }
pdePasswordRequired, pdeInvalidPassword:
AMessage := 'This document is password protected. Remove its ' +
'protection before editing it here.';
pdeInvalidDocument:
AMessage := 'This file is not a readable PDF document.';
pdeUnsupportedFeature:
AMessage := 'This document uses a PDF feature that cannot be edited.';
pdeReadError:
AMessage := 'The document could not be read from disk.';
else
AMessage := E.Message;
end;
end;
end;
end;
TTMSFNCPDFDocumentError has eight values:
| Value | Raised when |
|---|---|
pdeInvalidDocument |
The data is not a readable PDF: header, trailer or cross-reference table missing or corrupt. Also raised when an operation is attempted with no document open, or while a page edit is still active. |
pdePasswordRequired |
The document is encrypted and no password was supplied. |
pdeInvalidPassword |
A supplied password does not unlock the document. |
pdeUnsupportedFeature |
The document, or the requested operation, uses something the structural editor does not support — including an encrypted document opened with a password, and the generation-only calls listed under Common mistakes. |
pdeInvalidPageIndex |
A zero-based page index is outside the document's page range, or a delete would leave it with no pages. |
pdeReadError |
The document could not be read from its file or stream. |
pdeWriteError |
The document could not be written to its file or stream. |
pdeUnknown |
The failure could not be classified further. |
Only pdeInvalidPageIndex is genuinely recoverable in code — clamp the index and
retry. The rest are reporting cases. In particular, encrypted documents cannot
be opened at present, so prompting the user for a password after
pdePasswordRequired does not help: supplying one turns the failure into
pdeUnsupportedFeature. Decrypt the file with another tool first, or tell the
user the document is protected.
Putting it together
This walks a document end to end: it opens the file, reports its metadata, the geometry and first text line of every page, and saves a clean copy — combining the opening, geometry, text, metadata and error-handling sections above.
procedure TForm1.SummarizeAndCopy(const ASourceFile, ATargetFile: string);
var
p: TTMSFNCPDFLib;
Info: TTMSFNCPDFPageInfo;
Lines: TStringList;
FirstLine: string;
Created: TDateTime;
I, LineBreak: Integer;
begin
Lines := TStringList.Create;
try
p := TTMSFNCPDFLib.Create;
try
try
p.OpenDocument(ASourceFile);
except
on E: ETMSFNCPDFDocument do
begin
ShowMessage('Cannot open this document: ' + E.Message);
Exit;
end;
end;
try
Lines.Add(Format('%s - %d page(s)',
[ExtractFileName(ASourceFile), p.GetPageCount]));
{ Metadata reads from the OPENED document while one is open. }
if p.Title <> '' then
Lines.Add(' title : ' + p.Title);
if p.Author <> '' then
Lines.Add(' author : ' + p.Author);
Lines.Add(' producer : ' + p.Producer);
{ TryPDFDateToDateTime (FMX.TMSFNCPDFCoreLibBase) parses the raw PDF
date string and fails softly on a malformed one. }
if TryPDFDateToDateTime(p.CreationDate, Created) then
Lines.Add(' created : ' + DateTimeToStr(Created));
for I := 0 to p.GetPageCount - 1 do
begin
Info := p.GetDocumentPageInfo(I);
{ Page text is returned with one line feed per text line. }
FirstLine := p.GetDocumentPageText(I);
LineBreak := Pos(#10, FirstLine);
if LineBreak > 0 then
FirstLine := Copy(FirstLine, 1, LineBreak - 1);
Lines.Add(Format(' page %d: %.0f x %.0f pt - %s',
[I + 1, Info.Width, Info.Height, FirstLine]));
end;
{ Saving an opened document without mutations writes a clean copy. }
p.SaveDocument(ATargetFile);
finally
p.CloseDocument;
end;
finally
p.Free;
end;
Memo1.Lines.Assign(Lines);
finally
Lines.Free;
end;
end;
Saving an opened document you did not change is not a no-op, incidentally: the copy is rewritten through the same compacting save path, which is a cheap way to normalize a document produced by another tool.
Common mistakes
- Calling
NewPage, or touchingFormFields, on an opened document. Both belong to the generation API and raiseETMSFNCPDFDocumentwithpdeUnsupportedFeature, naming the supported route, rather than quietly writing into the wrong document — andFormFieldsraises as soon as you read the property, not on the first field you add. Generate the page or the form fields into a new document and bring it in withMergeDocument; to draw on an existing page useBeginPageEdit. - Drawing through
Graphicswithout starting a page edit. Outside an edit,Graphicsaddresses the generation canvas, so the drawing never reaches the document you opened. - Forgetting that
CloseDocumentdiscards unsaved work. Save first, then close. - Treating page indices as one-based. Every page index in the editing API is zero-based; page 1 for the reader is index 0.
- Assuming page 1 speaks for the document. Size and rotation are per page —
call
GetDocumentPageInfoinside the loop, not before it. - Saving or mutating while a page edit is open. Saving and every structural
change raise while
IsPageEditingisTrue. CallEndPageEditorCancelPageEditfirst — note thatOpenDocumentis the one exception: it cancels the edit for you instead of raising, so an unfinished edit is lost without warning. - Expecting a password to unlock an encrypted file. It will not; see Handling failures.
See also
- Rearranging pages — delete, move, duplicate, rotate, merge, extract.
- Page overlays and underlays — draw on an existing page.
- PDF Library guides — generating documents from scratch.
TTMSFNCPDFLib— full class reference.