PDF discovery is often discussed badly. Many text-based PDFs can still be indexed, and Google says as much. The real issue is not whether a PDF can appear in search. It is whether the content is published in a form that search systems can understand, preview and connect to wider web context with minimal friction.
Structured HTML usually gives public content a better chance of being found, interpreted, previewed and used because it carries richer structure, supports clearer page-level signals, and fits more naturally with how modern search systems surface information (Google Search Central, File types indexable by Google).
That distinction matters because discovery is no longer only about whether a file exists online. It is about whether the content is easy to crawl, easy to understand, easy to present in search, and easy to use once someone arrives. Public information that lives mainly in PDFs can still be found, but it often carries less of the page-level structure, preview control and machine-readable context that HTML pages can provide directly.
That is one reason government guidance continues to push publishers towards HTML as the default format for public content (GOV.UK, Content types; Government Digital Service, Why GOV.UK content should be published in HTML and not PDF; Section508.gov, Create Accessible PDFs).
The real issue is not visibility alone
Google's current documentation makes two points that should shape this discussion. First, many text-based file types, including PDF, are indexable. Second, for AI search experiences such as AI Overviews and AI Mode, Google says the same technical requirements still apply: a page must be indexed and eligible to be shown with a snippet in Google Search.
That means the goal is not just to put a document online. The goal is to publish content in a form that search systems can reliably crawl, understand, preview and present (Google Search Central, File types indexable by Google; Google Search Central, AI features and your website).
This is where HTML has the advantage. HTML pages can expose headings, navigation, page titles, metadata, structured data, crawlable images, internal links and snippet controls within the same environment that users actually read. They can also be updated without republishing static files, which matters when public information changes.
Google does not say that structured data or page markup guarantees visibility, and it does not promise that any given page will appear in AI features. But its guidance is clear that search features rely on indexed, technically accessible, user-visible page content (Google Search Central, AI features and your website; Google Search Central, Introduction to structured data markup in Google Search).
Why HTML carries stronger discovery signals
HTML is designed for the web, and that has practical consequences. Google explains that structured data is added to the page the information applies to and should describe visible page content. It also explains that snippets are generated primarily from page content, with meta descriptions used when they provide a better summary.
Those controls exist naturally in web pages. They are part of how publishers shape how content is understood and previewed in search (Google Search Central, Introduction to structured data markup in Google Search; Google Search Central, Control your snippets in search results).
This matters because discovery is influenced by more than raw indexability. Search systems need clues about what a page is, how its parts relate to one another, what should be shown to users, and which elements are central. HTML gives a publisher far more control over those signals. Headings can be nested properly. Navigation can be explicit. Breadcrumbs, article markup and other structured data can be implemented on the same page as the visible content. Snippet behaviour can be shaped with page-level markup.
None of this guarantees prominence. All of it improves clarity (Google Search Central, General structured data guidelines).
By contrast, a PDF is often a file that has been exported from another source. Even when Google can index the text, the document still tends to arrive with less flexible structure and fewer page-level signals. A PDF cannot carry in-page structured data in the way an HTML page can. It is also more limited as a search presentation object. The result is not that discovery becomes impossible. The result is that the content is harder to enrich, easier to flatten, and more dependent on extraction rather than on web-native structure.
Why this now matters more in AI search
Google's guidance on AI search experiences is disciplined. It does not introduce a separate set of rules for AI visibility. Instead, it says the same foundations still matter: technical eligibility for Search, indexability, snippets, and helpful, reliable, people-first content.
That is an important point. AI search does not remove the need for strong web publishing. It increases the value of it (Google Search Central, AI features and your website).
In AI Overviews and AI Mode, Google says it may issue multiple related searches and identify supporting web pages while a response is being generated. That suggests a wider and more dynamic retrieval process than a single classic search result. In that environment, structurally clear pages have an advantage. They are easier to interpret as pages, easier to preview as pages, and easier to connect with related content across a site.
HTML does not win because it is fashionable. It wins because it is the native format of the web that these systems are built to work with.
The practical case is straightforward. HTML reduces friction in the path from published content to discovered content. It gives search systems more direct signals and gives users a better landing experience when they arrive. If a public body, regulator, university or publisher wants its content to be found and used, that is the outcome that matters.
Discovery is also about what happens after the click
GDS makes a point that is often missed in search discussions. A PDF can break the user's sense of place. Depending on the browser or device, it may open in a new tab, a separate viewer, or as a download. The user is taken away from the surrounding site context and may lose the navigation that would help them continue their journey. That matters even more when someone lands directly on the file from a search engine (Government Digital Service, Why GOV.UK content should be published in HTML and not PDF).
An HTML page behaves differently. It keeps the content within the navigation, linking and orientation of the site. It is easier to move from one page to the next, easier to explore related content, and easier to continue from summary to detail. Discovery should not be measured only by whether a document is located. It should also be measured by whether the user can make use of it once found. In that sense, HTML usually performs far better.
Mobile use is part of discovery, not separate from it
GOV.UK and GDS are explicit that PDFs are weak on mobile. PDFs generally do not reflow to fit the browser. They often require zooming and horizontal as well as vertical scrolling. GOV.UK's content guidance says HTML is more accessible on mobile devices and more suited to assistive technology because text can reflow as users zoom and page elements are tagged in a more web-native way. Section508.gov also says PDFs are often not the most accessible or mobile-friendly option, and that agencies should prioritise HTML.
This is not a side issue. A public page that is awkward to read on a phone is less likely to be used well, less likely to be explored further, and more likely to be abandoned. Search, discovery and use are linked. If a result leads to a cramped, fixed-layout file that is difficult to navigate on a small screen, the publishing format itself is working against the purpose of the content.
HTML makes public content easier to enrich and maintain
Google's structured data guidance also points to a broader advantage of HTML. Web pages can be marked up, tested, improved and maintained at scale in ways that exported files rarely can. Google recommends JSON-LD for structured data partly because it is easier to implement and maintain with fewer user errors. That is a quiet but important point.
Discovery works better when content is maintained as living web content rather than as a collection of static documents (Google Search Central, Introduction to structured data markup in Google Search).
The same applies to snippet control and search presentation. Google says publishers can influence snippet behaviour with page-level markup such as meta descriptions, nosnippet and max-snippet rules. Those are web controls. They are part of a publishing environment where the page itself is the managed object. That makes HTML more adaptable as search evolves. It also makes it more useful for organisations that need public content to remain accurate, current and discoverable over time (Google Search Central, Control your snippets in search results).
The government position is already clear
The most practical part of this discussion is that the policy direction is not hypothetical. GOV.UK says HTML should be the first choice for publishing documents whenever possible. Its content-type guidance says PDFs should not be published on GOV.UK in most circumstances and, where a PDF is used, an accessible version must be published with it. GDS states plainly that information in a PDF is harder to find, use and maintain than HTML content.
Section508.gov uses almost the same operational language, saying agencies should prioritise HTML and use PDFs only when necessary (GOV.UK, Content types; Government Digital Service, Why GOV.UK content should be published in HTML and not PDF; Section508.gov, Create Accessible PDFs).
Those positions are not based on a single technical objection. They reflect a broader judgement about usability, accessibility, responsiveness and long-term digital management. Discovery sits within that same picture. Public content has a better chance of being found and used when it is published as structured HTML, because HTML supports the way modern search systems and modern users actually interact with information.
PDF still has a place, but not as the default
There are still cases where a PDF may be needed. A signed record, a print-ready document, an archival output, or a downloadable companion copy can all be legitimate. The point is not that PDF should disappear. The point is that PDF should not be the primary format for public content that people are expected to find, read and use online. When the website version matters, HTML should usually come first (GOV.UK, Creating and updating pages).
That is the disciplined conclusion. Google can index many text-based PDFs. AI search features do not impose a special set of HTML-only rules, but the web still runs on pages, not just files. Public content performs better when it is published in the format that search systems, browsers and users are all built to handle most effectively. In practice, that means structured HTML.
Sources
- Google Search Central, File types indexable by Google
- Google Search Central, AI features and your website
- Google Search Central, Introduction to structured data markup in Google Search
- Google Search Central, General structured data guidelines
- Google Search Central, Control your snippets in search results
- Government Digital Service, Why GOV.UK content should be published in HTML and not PDF
- GOV.UK, Content types
- GOV.UK, Creating and updating pages
- Section508.gov, Create Accessible PDFs
