Do PDFs Keep Old Versions of Themselves?
Often, yes — and it is why clearing a PDF's document properties can leave the old ones physically intact in the same file.
The short answer
A PDF can contain its own history. Not as a feature you opted into, but as a side effect of how the format saves changes: rather than rewriting the file, an editor is allowed to append the changed objects to the end and leave everything that came before exactly where it was.
The practical consequence catches people out constantly. You open a document, clear the Author and Title in the properties dialog, save, and the dialog now shows them empty. The file, meanwhile, still contains the original values a few kilobytes earlier — readable by anyone who opens it in a text editor.
How incremental updates work
A PDF is a collection of numbered objects plus a cross-reference table telling a reader where each object starts. At the very end of the file sits a startxref pointer to that table, followed by an end-of-file marker.
When an editor saves incrementally, it does not touch any of that. It appends the objects that changed, appends a new cross-reference table listing where those live, and appends a new startxref pointing at the new table. A reader opens the file, jumps to the last startxref, and sees the newest version of every object.
Nothing is deleted. The superseded objects are simply no longer referenced — orphaned, but still physically present in the bytes.
- One startxref means the document was written as a single revision.
- Several means the file carries that many saved states.
- The older values are not encrypted or obfuscated — they are usually plain text, or a standard compressed stream any PDF library can decompress.
This is a feature, not a bug
Incremental saving exists for good reasons. It makes saving a large document fast, because only the changes are written. It makes digital signatures possible: a signature covers a specific byte range, so appending new content preserves the ability to verify what was signed at the time.
It also underpins revision history in review workflows. The format is doing exactly what it was designed to do — the problem is only that most people saving a file have no idea it works this way.
Why clearing the properties is not enough
The document information dictionary — Title, Author, Subject, Keywords, Creator, Producer and the timestamps — is just another object. Clear those fields and your editor writes a new, empty information dictionary and points the new cross-reference table at it.
The old dictionary, with your name in it, remains in the file. The properties dialog reads the current one and shows nothing, which is precisely why this failure is so persistent: every check you are likely to perform confirms the file looks clean.
The same applies to an XMP metadata packet, and to anything else you edited out. What you removed is the reference, not the data.
How to check any PDF yourself
You do not need special tooling for the first check. Open the file in a plain text editor and search for startxref. Count the occurrences: more than one means the document has multiple revisions in it.
For the second check, search that same text view for a name or title you expected to be gone. PDF strings are often stored plainly, so a leftover author frequently appears as readable text. If it does not, that is not proof of absence — it may sit inside a compressed object stream, which needs a tool that decompresses before searching.
The PDF tool on this site reports the revision count for any document you load, and warns when it finds more than one.
What actually removes an old revision
Only rewriting the document. The file has to be reconstructed from the current object graph so that unreferenced objects are never written to the output.
- A full rewrite — what this site's PDF tool does. It parses the latest revision, drops the metadata, and serialises a fresh document, so superseded objects are not carried over.
- qpdf --linearize rewrites the file structure and discards orphaned objects, though it does not clear the metadata for you.
- Print to PDF removes everything, at the cost of the document: text can be re-flowed or rasterised, and links, form fields and accessibility tagging are usually lost.
- What does not work: editing the properties and saving again, if your editor saves incrementally. That is the trap this whole article is about.
Redaction is a separate problem — and a worse one
Metadata is not the only thing that survives an apparent deletion. Drawing a black rectangle over a paragraph adds a rectangle; it does not remove the text underneath. The words remain in the content stream and can be selected, copied or extracted by any tool that reads the page.
This is a recurring class of failure in published documents, and it is worth being precise about why it keeps happening: the visual result looks exactly like a redaction, so the person doing it has no reason to suspect otherwise.
Removing metadata and redacting content are different operations solving different problems. Doing one does nothing for the other.
A short checklist before sending a document
If a PDF is going somewhere it cannot be recalled from, these four steps cover the common failures:
- Check the revision count. More than one means the file has history in it.
- Rewrite the file rather than editing its properties in place.
- Verify afterwards by re-reading the output, not by re-opening the properties dialog.
- If anything was hidden rather than deleted, redact it properly first — that is a separate job.
Check how many revisions your PDF contains
PDFs record who made them, with what software and when — and often keep earlier drafts inside the same file. See what yours contains, then remove it.
Open Remove Metadata From a PDFFrequently asked questions
How can I tell if a PDF has more than one revision?
Open it in a text editor and count the occurrences of startxref. One means a single revision; more means the file contains that many saved states. The PDF tool on this site reports the count for you.
Does clearing document properties delete the old ones?
Not if your editor saves incrementally. It writes a new, empty information dictionary and leaves the old one in the file. The properties dialog reads the new one, so the document looks clean while the original values are still present.
Can someone actually read the old values?
Yes. Superseded objects are not encrypted or obfuscated. They are often plain text, and when compressed the compression is standard and reversible by any PDF library.
Does saving a copy remove the history?
It depends entirely on whether the application performs a full rewrite or another incremental save. Some do one, some the other, and the interface rarely says which. Verifying the output beats trusting the menu item.
Why does the format work this way at all?
Speed and signatures. Appending changes makes saving large documents fast, and it lets a digital signature keep covering the exact bytes it originally signed.
Will rewriting a PDF break its digital signature?
Yes. A signature covers a specific byte range, so any modification invalidates it — true of every tool that changes the file, not just this one.
Is this the same as failed redaction?
No, though they are often confused. Old revisions are leftover objects and metadata; failed redaction is visible content hidden behind a shape but still present in the page. They need different fixes.
Does this affect images too?
No. JPEG, PNG and WebP have no equivalent of incremental updates — editing one rewrites the file. This trap is specific to PDF.