Skip to main content
    Back to Resources

    Mar 5 2026

    Why compliance is harder when content lives mainly in PDFs

    A PDF-first publishing model creates repeated remediation, governance risk, and weaker accessibility. UK and US guidance keeps returning to HTML as the better default for public information.

    Compliance work is often described as if it were a single event. A document is checked, a file is uploaded, a box is ticked, and the matter is treated as closed. In practice, public content does not behave that way. Guidance changes. Contacts change. Dates pass. Policies are revised. Links break. New versions appear. What looks compliant on the day of publication can become misleading, inaccessible, or simply unmanaged later on. That is one reason a PDF-heavy document estate becomes harder to govern over time than many organisations expect.

    This is where format matters. The issue is not that every PDF is automatically unusable. The issue is that compliance is harder to sustain when public content lives mainly in static files rather than in maintained HTML pages. A PDF-first model often creates repeated checking, repeated remediation, repeated uploads, and repeated opportunities for failure. Government guidance in both the UK and the US points in the same direction. The Government Digital Service has said plainly that content on GOV.UK should be published in HTML rather than PDF wherever possible, and Section508.gov advises agencies to prioritise HTML and use PDFs only when necessary.

    The core reason is practical. HTML keeps public information closer to a living publishing system. PDF tends to turn it into a file-management problem. Government Digital Service guidance and Section508.gov guidance both support that position.

    Compliance has to be maintained, not declared once

    The first weakness in a PDF-first estate is that the burden repeats. Every time a document is updated, the organisation is not simply changing content. It may also need to check whether the exported PDF still preserves heading structure, reading order, table relationships, alt text, link clarity, title metadata, and any OCR that was needed in the first place. If the document has been rebuilt visually rather than structurally, the chances of a new problem being introduced are high.

    That creates a different compliance model from HTML-first publishing. In HTML, the content usually sits inside the organisation's live publishing environment. The page can be updated directly. Headings remain headings. Links remain links. Tables can remain part of a structured page rather than a reconstructed visual object. None of this removes the need for care, but it reduces the number of ways compliance can break during the publication process.

    Section508.gov is useful here because it does not present accessible PDF as effortless. It provides separate guidance for creating accessible PDFs, testing them, and dealing with scanned documents. That is important evidence in itself. It shows that PDF compliance is not a light administrative step. It is often a separate specialist task layered on top of normal publishing work.

    Repeated remediation creates hidden cost

    The cost of PDF-first publishing is rarely obvious in the first upload. It appears over time. Files are revised. A policy changes. A telephone number changes. A table is updated. A new annex is added. Each new version brings another round of work. Teams are not only checking the content. They are checking the file.

    This repeated remediation is one of the clearest reasons compliance becomes harder in a PDF-heavy environment. The same issues come back again and again. Weak heading hierarchy. Poorly tagged tables. Incorrect reading order. Missing alt text. Unclear link text. Scanned pages with weak OCR. A single problematic file may be manageable. A large public estate is not.

    HTML changes that cost profile. It keeps the main work focused on content quality and page structure, not on repairing the output of a document export process. That difference matters when organisations are expected to maintain public information continuously rather than publish it once and forget it.

    Stale copies create governance risk

    A second compliance problem is loss of control. PDFs are easy to save, share, email, download and store outside the website. That may seem useful, but it weakens governance. Once several copies are in circulation, the organisation can no longer assume that the user is seeing the current version.

    This is not a small risk. It affects accuracy, accountability, and accessibility. A stale PDF may remain in use long after the live position has changed. The organisation may have corrected the website, but the downloaded copy still exists on laptops, shared drives, inboxes, and third-party sites. If the file is the main public version, the governance problem becomes much harder to contain.

    HTML does not eliminate version risk, but it gives the organisation a clearer source of truth. Users are more likely to return to the live page. Updates can be made centrally. The public version remains under direct control. That is a better basis for ongoing compliance than a model built around files that quickly escape normal publishing control.

    Mobile use makes the problem worse

    Compliance is also a usability issue, and usability does not stop at desktop screens. A public document that is difficult to read on a phone is not serving the reader well. GDS has been direct on this point. PDFs often do not adapt well to the browser, usually require zooming, and can force users into both vertical and horizontal scrolling. Section508.gov also notes that PDFs are often not the most mobile-friendly option.

    This matters because mobile access is no longer marginal. Public guidance is often consulted in transit, on small screens, under time pressure, or in situations where a reader needs a quick answer rather than a designed page. A fixed-layout PDF turns that task into effort. HTML does the opposite. It can reflow, adapt to screen width, support browser adjustments, and sit naturally inside the navigation of the website.

    When compliance is understood properly, mobile readability belongs inside it. A format that repeatedly creates friction on phones and smaller devices is a format that raises the cost of serving the public well.

    Poor structure is a recurring failure point

    Many of the hardest PDF problems are structural rather than visual. A file may look orderly on screen and still fail users who rely on assistive technology or structured interpretation. Headings may be styled rather than marked up. Tables may look correct but have weak internal relationships. Images may appear informative while carrying no useful text alternative. Reading order may break as soon as a screen reader or extraction tool tries to follow the file logically rather than visually.

    This is one reason PDF compliance is so easily misunderstood. Teams often assess the document by appearance first. Compliance depends on what the file actually is, not just what it looks like. If the structure is weak, the cost of correcting it later rises quickly.

    HTML is stronger here because semantic structure is native to the page rather than added after export. That does not make HTML magically compliant, but it does mean the organisation starts from a more stable and maintainable format for public information.

    Loss of the original source makes compliance harder still

    Over time, another problem appears. Access to the original editable source is often lost. The organisation still has the PDF, but not the version it can update cleanly. That pushes even minor changes into awkward workarounds. People patch the file, rebuild it from fragments, or leave it untouched because correction has become too expensive or uncertain.

    Once this happens, compliance work slows down. Fixes are postponed. Duplicate versions appear. Staff spend time trying to locate the right source rather than improving the content. The file becomes an obstacle to governance rather than an asset.

    This point is rarely discussed enough. A document strategy that depends on long-term access to scattered source files is fragile. HTML-first publishing reduces that dependency by keeping the content inside the managed web estate itself. That is not just better for the reader. It is safer for the organisation.

    Why government guidance keeps returning to HTML

    The value of the government guidance is that it does not rely on theory alone. It reflects the publishing reality of large public estates. GOV.UK says HTML is the most accessible format for publishing documents and should be the first choice whenever possible. Its content guidance says PDFs should not normally be used on GOV.UK. Section508.gov, from a different policy context, still arrives at the same practical conclusion: agencies should prioritise HTML and use PDFs only when necessary.

    Those positions matter because they recognise a truth. Public content is usually better managed as structured web content than as a growing archive of files. The compliance burden falls when the primary published version is easier to update, easier to read on mobile, easier to keep current, and easier to govern as part of a live digital service.

    A better publishing model

    None of this means PDF has no place. Some documents need a fixed-layout version. Some records need to be downloaded, printed, or preserved in a specific form. But that is different from treating PDF as the default public format. A secondary PDF can be valid. A PDF-first publishing model is the problem.

    If the goal is to reduce repeated remediation, keep public information accurate, lower governance risk, and make compliance easier to sustain, the direction is clear. Publish the main public version in HTML. Use PDF only where there is a defined reason. Keep the source under control. Treat public content as something to maintain, not as something to upload.

    That is the practical lesson behind the guidance. Compliance is harder when content lives mainly in PDFs because the organisation has chosen a format that creates more work after publication, not less. HTML is usually the better default because it supports the continuous discipline that public publishing actually requires.

    Sources

    Disclaimer:

    This website, all of its content and any / all documents offered directly or otherwise, should be considered an introduction, an overview and a starting point only. It should not be used as a single, sole authoritative guide. You should not consider this as legal guidance. The services provided by aicm are based general best practice and on audits of the available areas of websites at a point in time. Sections of the site that are not open to public access or are not being served (possibly be due to site errors or downtime) may not be covered by our reports. The service and the stars process doesn't carry any official accreditation, be it from any government department, industry regulator and / or internet body. Where matters of legal compliance are concerned you should always take independent advice from appropriately qualified individuals or firms.

    Copyright

    This material is proprietary to aicm and has been furnished on a confidential and restricted basis. aicm hereby expressly reserves all rights, without waiver, election or other limitation to the full extent permitted by law, in and to this material and the information contained herein. Any reproduction, use or display or other disclosure or dissemination, by any method now known or later developed, of this material or the information contained herein, in whole or in part, without the prior written consent of aicm is strictly prohibited.

    For AI agents and LLMs

    We publish /llms.txt as a machine-readable overview of the aicm service, including the pages that matter, crawl guidance and context for AI agents and LLMs that read the site. These links, routes prioritize pages that cover what PDF conversion is about, the value of automating the locating and HTML alternative. Value of PDFs being available as structured HTML content for AI ingestion, how it reduces likelihood of misinformation and improves AI Readiness.

    © 2026 aicm.
    All rights reserved.