Executive summary
All 172 public documents we checked, from 66 Irish and UK organisations, failed at least one PDF/UA-1 check: none met PDF/UA-1. 4 of them fail only the rule that asks a file to declare a conformance claim; the other 168 fail at least one rule beyond it.
Automatic tagging helps a document that has no structure. Our own earlier, unguarded rebuild damaged documents that already had some. Across 22 already-tagged documents it kept 0 of 41 image descriptions their authors had written by hand (those 41 came from 5 of the 22) and 0 of 999 table header cells. Over the same 22 documents, 20 gained headings they did not arrive with, 10 of them from none at all. Both results come from the same process.
A clean validator result is not the same as a usable document. The most successful repair we have built took failures of one rule from 1,406 to 4 across 25 documents, and it stays switched off for customers until a screen reader has been put on its output. When we put a screen reader on our own rebuild of a document, it read out web addresses, hundreds of characters at a time, where the visible page said one word. The validator was satisfied.
This paper sets out where the line sits between what a machine can fix and what it cannot, from our own measurements.
01What we found
All 172 public documents from 66 Irish and UK organisations failed at least one check, so none met PDF/UA-1. The median document failed 5 distinct checks, with a range of 1 to 14. Across the set, 42 different rules failed somewhere.
The documents were checked between 27 and 29 September 2026 with veraPDF 1.30.2 against the PDF/UA-1 profile. Counts in this section are of distinct rules failed, not instances: one rule can fail a hundred times in one document and is counted here once.
| Check failed | Documents failing | Share of the 172 |
|---|---|---|
| The file doesn't state its accessibility claim | 135 | 78.5% |
| A link has no alternate description | 88 | 51.2% |
| An annotation has no alternate description | 87 | 50.6% |
| The document declares no title | 74 | 43.0% |
| Content is neither tagged nor marked as page furniture | 60 | 34.9% |
| The viewer isn't told to show the document's title | 57 | 33.1% |
| An image has no text description | 41 | 23.8% |
| A font isn't embedded in the document | 40 | 23.3% |
| The document is missing its XMP metadata | 34 | 19.8% |
| The tab order isn't set to follow the document structure | 33 | 19.2% |
The first row deserves a caveat rather than a headline. Any PDF that was never built to the standard fails it by definition, because the file simply doesn't declare a claim. It tells you the file was not built to claim PDF/UA-1; it does not tell you a reader is blocked.
We are deliberately careful about what that second number means. A document that fails another rule is not necessarily one a reader is blocked by. Several of the rules that fail often describe a file that is incorrectly assembled rather than a document somebody cannot read: the missing XMP metadata stream, the viewer not being told to show the title, font internals, tab order.
The same care applies to the rows that look most alarming. Half these documents contain a link with no alternate description, but a tagged link with visible text is normally read by its text, so that rule failing does not by itself mean a reader hears nothing useful. It bites on a link whose visible text carries no meaning on its own, or one that is not tagged as a link at all. Similarly, four in ten declare no title and a third are not set to show the title rather than the filename; it is where those two fail together that a reader is told a filename instead of what they have opened.
02What a machine fixes reliably
A machine is reliable on the things that are declarations about the file rather than statements about its content: whether an XMP metadata stream exists, whether the viewer is told to display the title rather than the filename, whether the tab order is set to follow the structure. These are present or absent, there is one correct setting, and nothing is being decided.
Where a fix needs words, the machine can propose and a person has to confirm. A title has to say what the document is. A language declaration has to be the language the document is actually in. Both are checkable by machine and neither is decidable by one.
There is one row in that table we could clear tomorrow and choose not to. The commonest failure in the corpus, at 78.5%, is the file not stating a conformance claim, and a conformance claim is a single piece of metadata. We write it only when it is the one rule still failing and the file then passes. Otherwise, writing it would mean asserting that the document conforms, which is the one thing a tool must never assert on a customer's behalf.
A fix that consists of claiming to be fixed is not a fix.
03What a screen reader heard
On a corpus of 25 already-tagged documents from 14 different producers, adding the missing Link elements took failures of that rule from 1,406 to 4. Of the four left, two have no tagged text beneath them and are left for a person by design; two are a case the repair does not handle. By the validator's measure that is finished work. It stays switched off for customers until a screen reader has been put on its output.
That condition comes from what happened when we put a screen reader on our rebuild of one document, a different stage of the same pipeline. It read out web addresses.
Not descriptions of links: the addresses themselves. Three filter controls on one page, labelled in the visible document as “Background”, “Nature” and “Butterfly”, were each announced as 300 to 400 characters of URL. A sign-out link ran to roughly 600. A heading was read as the full address followed by “link, heading level one”. A sighted reader sees one word. A screen reader user heard a paragraph of query string.
The cause was in our pipeline, not in the standard. The tagging stage writes a link's own web address into the field meant for its description, and the pipeline stripped placeholder descriptions only from images, never from links. The original document's links were announced as a bare “link”, which is unhelpful. The rebuild's were announced as their raw addresses, which is worse, because it is unhelpful at length. The pipeline now gives each link the words it shows on the page, and keeps the address only where the page shows none.
The validator was satisfied at every stage. Every link carried a description, which is all the rule asks, so the rule passed on exactly the property that made the links unusable.
“The validator no longer fails it” and “a person can now read it” are different claims, and the gap between them is not a rounding error.
04What a machine damages
Rebuilding the structure of a document that already has one destroys the parts a person put there. We found this in our own pipeline and measured it across 22 already-tagged documents. It is why we no longer rebuild a tagged document without checking first. Everything in this section describes an unguarded rebuild.
| What the authors had | Kept |
|---|---|
| 41 image descriptions, written by hand (from 5 of the 22 documents) | 0 |
| 999 table header cells (from 2 of the 22) | 0 |
| 137 footnotes and endnotes | 0 |
| 10 tables of contents | 0 |
The header cells were not deleted. They became ordinary data cells, which is worse: the table still looks like a table, and a screen reader can no longer say which column a number sits in. One document lost 633 header cells on its own. One annual report went from 1,503 table cells to 330. One legislative text went from 61 headings to 14.
Two documents in that corpus are web pages printed to PDF, and they are worth describing because a rebuilt version of one looks alarming until you count what is on the page. One of those pages carries 2,825 link annotations with 106 distinct targets, and draws 571 images, 136 of them 16-pixel icons; the other draws 548. The rebuild tagged what the printout had put there. The gap between an author's 14 tagged figures and the hundreds of images actually drawn is a property of how the file was produced, not evidence of a rebuild inventing anything.
Why nothing caught it
PDF/UA-1 does not fail a table that has no header cells. It has no way to know the table ever had any. So a document that loses 633 header cells produces no new failed check, and a guard that watches for a pass turning into a fail sees nothing at all.
The document scores the same or better, and reads worse. Any tool whose only measure is the validator's verdict will report success.
The same process helps weak documents
One annual report went from 0 headings to 93, a jury report from 0 to 36. Only two did not gain: the legislative text above, and one document that arrived with 18 headings and left with 18.
This is not a bad process. It is a process applied where it does not belong.
05What we do about it
The rule cannot be “tag everything” or “never tag”. It has to be decided per document: rebuild only when the rebuilt structure keeps every image description, header cell, note, contents list and heading the author wrote. Where it does not, keep what the author made and apply only the safe fixes. Where even that regresses, refuse the document and say so.
Our preservation guard compares the author's tree with the rebuilt one. It checks image descriptions by matching their text rather than counting them; a count would miss the loss, because one rebuild carried 2,999 descriptions of its own and kept none of the author's 13. It does not yet compare tables, cells or reading order, which is a real limit and the most important thing on our list.
Its evidence today is eight documents checked locally and 32 unit tests. No production job has been observed running it. A separate component, the integrity checker inside the in-place repair, has been tested adversarially by breaking 25 real structure trees on purpose: it caught 24 of 24 dropped parent-tree entries, 24 of 24 double claims and 17 of 17 deleted headings. Those are that checker's results, not the guard's.
06What needs a person
Some faults have a right answer that only a human knows. A machine can write a description for every image in a document and satisfy the check completely, without any of the descriptions being true.
Table headers are the sharper case. On our corpus a repair produced 992 suggested header cells and 88 suggested heading levels. Accepting all of them clears the failing checks: one report went from 23 failures of the header rule to none, another from 26 to none.
What that measures is that the checks stop failing. Nobody has verified that a single one of those 992 headers is correct. And because the standard does not fail a table with no headers, a wrong header produces no failed check at all. It produces a screen reader announcing the wrong column name, confidently, for every value in the table.
A wrong header is worse than a missing one.
A missing header leaves a reader knowing something is absent. A wrong one tells them something false and gives them no reason to doubt it.
So suggestions have to arrive as suggestions. The design we are building to shows the text each one proposes to promote, requires a person to accept or reject it, and has no button that accepts nine hundred at once. Our editor already lets a person set alt text, reading order, table headers and language. These machine suggestions are not in it yet, and until they are, they are not offered to anyone.
The same applies to reading order, to whether a heading is really a heading, and to whether an image is decorative or carries meaning. Two rules in the standard can be checked by a machine and fixed only in the source document: content that is neither marked as meaningful nor as decoration, and fonts whose characters cannot be mapped back to text.
07Fixing at source
Fix at source where you control the source. It is the better engineering answer: a document built with structure from the start needs no repair, loses nothing, and costs nobody a judgement call after the fact.
It does not answer the position most organisations are in, for four reasons.
The back catalogue is already published. What is in scope is a question for each organisation's own advisers: the EU Web Accessibility Directive and the European Accessibility Act each exclude some older office documents, from different dates, and the UK's regulations set their own. Whatever the scope, fixing the source changes what you publish next week. It does not change the document somebody downloads today.
The source is often not yours. A payslip, an invoice, a utility bill, a policy booklet, a patient letter, a fund factsheet: these are generated by software the publishing organisation bought and cannot modify. The organisation carries the duty and the vendor controls the structure. Telling a council to fix its documents at source is telling it to rewrite somebody else's product.
Authoring well is not the same as authoring conformantly. In our set, no document met PDF/UA-1, including the ones that arrived already tagged. Four of the ten commonest failures above are things an author cannot see and modern tools still routinely omit: the declared title, the XMP metadata stream, the conformance claim, and the setting that shows the title rather than the filename.
The human decisions do not disappear. Alt text, header scope and reading order do not become machine-decidable because they are made earlier. They become cheaper, which is a real argument for doing it at source, and they do not go away.
So: fix at source wherever the source is yours, check everything either way, and keep a record that says which documents a person still has to look at. Repair is what you do about the documents you have already published and the software you do not own.
08Documents that cannot be repaired in place
Digitally signed documents. Making a PDF accessible means rewriting its structure, and a rewrite that replaces the file invalidates the signature, which exists precisely to prove the bytes have not changed. The format itself is not the obstacle: PDF supports incremental updates that leave an earlier signed revision intact. Our tooling is. The libraries we build on cannot write an incremental update, so we refuse the file and ask for an unsigned copy, then have the accessible copy signed. We would rather say that than return a document whose signature we have quietly broken.
Forms with Adobe usage rights. These carry a signature that switches on extended features in Adobe Reader, typically the ability to fill in a form and save it. Remediating the document removes those features. That is a genuine trade and it belongs to the customer: a form that a disabled person can read but nobody can submit may or may not be an improvement, and the vendor is not the one who should decide. So we ask rather than proceed.
Scans with no text layer. A page that is only a photograph of text contains nothing to tag. In our random samples of published libraries, 21 of 199 documents that opened, about one in ten, had no text layer at all. That figure comes from six organisations' libraries and should be read as an indication, not a rate. These documents need optical character recognition and then a human pass over the result, which is a different job from remediation.
09What a record must say
“This document is accessible” is not a statement anyone can check. Someone will eventually ask what was tested, what changed, and who decided the parts a machine cannot decide. A green tick answers none of the three.
A verdict only means something alongside the things that produced it:
- which standard, and which part of it. PDF/UA-1 and PDF/UA-2 are different documents with different requirements, and we check against PDF/UA-1;
- which list of criteria, at which version. Ours holds 119 for a current check: 107 from PDF/UA-1 and 12 from WCAG 2.2, three of those twelve being human judgements, drawn from a catalogue of 128;
- which validator, at which version;
- what still fails, named in full rather than summarised;
- who accepted the judgement calls, and when.
That list has a consequence most tools avoid. If the criteria list can change, a record made under the old one has to keep pointing at the old one. We hold nine WCAG 2.1 criteria we no longer check against, solely because records signed earlier refer to them. A record that silently re-resolves to today's list is not a record of anything.
It also means a document can be checked twice, months apart, and produce two different verdicts that are both correct. Without the version of the criteria list beside each verdict, that looks like an error. With it, it is just the standard moving.
Ask any tool what it did not assess.
A tool that cannot answer has not distinguished between what it checked and what it claimed.
10Method and limits
We attempted 463 documents and report on 172. Everything dropped is listed here, and the figures sum to the number attempted.
| Outcome | Documents |
|---|---|
| Reported on | 172 |
| Lost to a bug of ours: a process-handling fault on 26 September, fixed the next day | 140 |
| Checked before the fix that stores every failure, so only three were kept | 69 |
| Link did not return a PDF | 54 |
| Scanned image, no text layer, so not validated | 25 |
| Refused by the site's robots.txt and not fetched | 2 |
| Network error, timeout or HTTP error | 1 |
| Attempted | 463 |
The 140 were a single failure of ours inside an eight-minute window, on three organisations' libraries. They say nothing about those documents and we have not reinstated them.
What this does not show
- The set is weighted toward a few publishers.
- 64 of the 172 are one document chosen per organisation; the other 108 are random samples from six organisations' libraries. Every per-document percentage above is pulled toward those six.
- It is mostly British.
- 145 documents from 39 British organisations against 27 from 27 Irish ones. The British documents published on gov.uk domains span 24 organisations; the domain alone does not say whether a publisher is local or central government.
- We did not verify what kind of organisation each publisher is.
- Our classification comes from the shape of the domain, not from checking.
- It is three days wide.
- 27 to 29 September 2026. Nothing here shows a trend.
- The damage measurement is a different corpus.
- The 22 documents in the rebuild comparison were hand-picked across several countries to cover a range of producers, and the rebuild was reproduced locally without the validator. They are not a sample of anything. The link-repair corpus of 25 is a third set. The screen-reader test was run on one document from the 22, as rebuilt.
- We did not check for signatures.
- Section 08 describes real limits of the format and of our tooling, not a measurement of this corpus. We do not know how many of the 172 were signed, because we did not look.
- Some figures are too small to report.
- We record which software produced each document, but only for 13 of the 172, with a largest group of three. There is nothing to say about producers yet, so we have said nothing.
- The validator reads a re-saved copy.
- veraPDF is given a decrypted copy of each file with its visible content unchanged, not the downloaded bytes.
What we counted
Every count of checks in Section 01 is of distinct rules failed, never instances. One rule can fail a hundred times in a single document; counted our way that is one. The repair figures in Sections 03 and 06 (1,406 to 4; 23 and 26 to none) are counts of individual failures of one rule. This matters when comparing against any other published figure, because the same document can be described as failing 5 checks or 98 times, and both can be true.
No organisation is named anywhere in this paper, and none will be. Every one of the 463 documents was published for the public to read, and any site whose robots.txt asked us not to fetch it was dropped rather than worked around.