Skip to main content
    Back to Resources

    Aug 17 2026

    RESEARCH PAPER | PDF AUTHORITY AND AI SOURCE SELECTION

    Where Authority Sits

    The importance of PDFs versus web pages, and why they are more authoritative.

    A research paper from AAAnow, The Digital Confidence Company. Sources are cited inline and listed in full, with the research areas used, in the appendix.

    Executive exposure

    CONTROL GAP

    Today, in 85% of cases, you are not in control of the conversation [AirOps, 2025]. Future competitiveness and managing compliance: if this is key to you, control has to become yours.

    The conversation is the one AI now holds about you. Your customers, your investors, and your regulators ask it questions. It answers for you in seconds, from whatever it can reach. You are not in the room when it speaks.

    85%Not in control

    This is already a buying conversation. 94% of B2B buyers use AI in their purchasing [Forrester, 2026]. Consumers now rank it among their 2 most influential purchase sources, its expected to influence more than $260 billion in e-commerce [IAB, 2025].

    When the answers come out wrong, all to oftern AI takes the blame for hallucinating. The truth sits closer to home. The machine repeats the estate it was given.

    None of this is news inside your building. Your digital teams have shouted about it, and fought it for years with tools that could not cover the ground. The estate outgrew them, and then the reader changed. The digital debt has come due.

    EXECUTIVE RISK

    What goes wrong, and the liability lands at the top

    1

    Safety instructions, misdelivered.

    A capital equipment company publishes the procedure for working on its machinery. In a California example, the detail of what to do arrived wrong, and the equipment was never isolated before the work began. An engineer acted on bad information carrying your name.

    2

    Chemical handling, out of order.

    The instructions were specific: what can be transported with what, how it travels, how it is stored, and the order for mixing. AI retold the sequence wrongly, with potentially dangerous impacts. Under health and safety rules, the first 2 patterns here can end with executives going to prison.

    3

    Annual reports, misread wholesale.

    A public company's report is poorly constructed, and the structural errors mean the figures are misquoted. Wrong numbers feed into other systems, and the knock-on is considerable: the analysts, the aggregators, and the answer engines carry the error forward.

    4

    Design spend, working against you.

    A report that cost tens of thousands to produce goes out missing its fonts, and the content arrives garbled. AI misunderstands the document and misdelivers it, and the piece built to strengthen shareholder confidence now undermines it.

    5

    The door left open.

    An annual report published without titles and the information AI needs to index it accurately, by an organisation that also blocks AI from reading it. Asked about the company, the systems find what they can. Anyone who publishes their own version and passes it off as yours has just been handed the conversation.

    No single role below your board sees the whole of this estate. Marketing holds a slice, legal another, operations a third, and partner sites carry copies of their own. Control of what the organisation says online is an authority question before it is a technical one, and authority of that kind comes from the top or it does not arrive at all.

    Minimising the exposure, starts here

    You understand the importance of controlling what goes online, and that far more rigour must be applied. This is not about preventing or stifling innovation or slowing things down; it is the reality of applying a level of regulatory control around what gets published. This is about understanding, organisation-wide, that you cannot just dump whatever you want online, or share PDFs and put them out wherever. Control of this is fundamental to future competitiveness as well as managing compliance.

    It is not just about what is being said today; it is about understanding the governance of the organisation. That is the key point here. Lack off governance, lack of control, means you are self-perpetuating the problem of misinformation and losing the battle for control of your conversation to AI and to other people. 3 stages take that control back, and they run in this order for a reason: nothing can be put in order before it is counted, and nothing should be converted before it is put in order.

    01

    Understand.

    Get an accurate picture of the landscape first. Automate the inventory from the start, whether that means 10 PDFs or 10,000, and take it beyond the main website, because the real trouble sits in the outliers: the exceptions nobody planned for.

    02

    Order.

    Condense the duplicates into a single document, retire the copies, and point them at that source, so you know where you are. Then map the usage, human and AI alike: which documents are being used, by whom, and how.

    03

    Deliver.

    Convert PDF to HTML so AI takes your information accurately, and inaccurate information stops being passed. Where PDFs are key to the organisation, they stay in use as the authoritative source, on a hybrid approach: the right ones properly checked and fully deliverable to AI, quality and compliance corrected and then verified.

    The paper behind this page explains how the machines came to read this way and how any organisation can test it. The part that cannot be delegated sits here: the conversation about your organisation is running today, and control of it belongs at the top, with you.

    1. Introduction

    Organisations publish the same knowledge in 2 forms: on web pages, and in documents. AI systems now read both, and they do not weigh the 2 equally.

    AAAnow research places the difference at 1.44 to 1.8 times in favour of the document. AI systems regard a PDF page as 1.44 to 1.8 times more authoritative than a web page [AAAnow, 2026]. This paper sets out where that weight came from, why it persists, and how the position can be tested by anyone who wants to check it.

    The number is stated at page level, with the content held constant: the same words, once as a document, once as a page. What is being weighed is the form the knowledge takes, and the weighing is done by the systems themselves.

    The explanation sits in how these systems were built. The models behind them learned language from 2 supplies of text: the open web taken at scale, and curated collections assembled for quality. The curated side leant heavily on document collections: scientific papers, books, court records, patents, and medical literature [Gao et al., 2020; Brown et al., 2020]. The open web was admitted only after severe cleaning, and the standard used for that cleaning was resemblance to the curated, document-grade prose [Brown et al., 2020].

    Content that lives in documents earned its weight through character rather than through its file extension. It tends to come from institutions, to pass through review before release, and to be dense with fact, structure, and citation. It also arrives without the advertising, navigation, and filler that web cleaning exists to strip out [Raffel et al., 2020].

    For years a barrier held that weight in check: machines read documents badly, jumbling columns and losing tables. Multimodal models removed the barrier by reading pages as pages [OpenAI, 2023]. The most heavily weighted class of content became directly readable at the moment AI search and AI agents began selecting sources for people. Selection is now where authority is exercised: research across 21,311 brand mentions in AI answers found 85% drawn from pages the brand does not control [AirOps, 2025].

    The implication for leadership is direct. A published estate holds both formats, and the balance of authority between them is now a property of that estate. What an organisation places in documents, and the condition those documents are in, shapes how AI systems read it and then represent it to the people asking questions about it.

    The paper covers 4 things in sequence: the record of what these systems were trained on, the character of the content that earned the weight, the change that made documents directly readable, and the behaviour of the systems now choosing sources. The finding sits on top of that record, and the test sits behind the finding.

    None of this requires a technical reader. The mechanism is set out in plain terms, the sources are public and cited in full, and a Q&A of the 20 questions most asked sits at the back, with the complete reference list and research areas in the appendix.

    The paper closes with a straightforward test any organisation can run, using public models and a handful of matched pages, to see the position for itself.

    2. What an AI system sees

    A language model reads text as tokens: short fragments of characters, a few letters at a time [Brown et al., 2020]. At that level a PDF page and a web page look identical. The stream carries words, and nothing in it announces a file format. No switch inside the model detects a PDF and raises its weight.

    That fact is the right starting point for this paper, because it rules out the lazy version of the claim. AI systems do not favour a file type. They favour what reaching them through that file type has come to mean.

    Throughout this paper, authority means the weight an AI system gives content when selecting, citing, or relying on it. That weight is set in three places: in which text reached the model during training, in how that text was sampled once it had, and then, downstream, in how the systems built on top of models pick their sources at the point of answering.

    The first 2 belong to the training record, which is public and documented. The third belongs to the behaviour of AI search and AI agents, which is measurable. Authority in short, is earned upstream of the model and exercised downstream of it.

    Authority, so defined, is a behavioural measure. It records what a system relies on when it answers, and reliance leaves evidence: a citation, a choice between conflicting sources, a page opened first. Section 6 shows that this evidence is now studied in its own right, and section 8 turns it into a test.

    The question has also acquired a deadline. While AI systems answered no one, the weighting inside them was an academic curiosity. Once they began answering questions and acting for people at scale, it became a property of how organisations are seen.

    The distinction carries the paper, so it is worth stating carefully. A claim that machines recognise and favour a file type would collapse under expert reading in a paragraph, and it should. The claim made here is narrower and stronger: the content the world publishes as documents was chosen, weighted, and held up as the standard when these systems learned language, and the systems now selecting sources act accordingly.

    Read this way, the position becomes an evidence question rather than a slogan. The sections that follow walk through the record: what was collected, what was weighted, what was filtered out, what changed when machines learned to read pages properly, and what the systems selecting sources do now. Section 7 states the AAAnow finding that quantifies the result. Section 8 describes how anyone can test it.

    3. How authority entered the models

    The training record of the major models is a story of 2 supplies. One is the open web, crawled at vast scale. The other is a set of curated collections, gathered because their text was worth more per word. The curated side was comprised of collections that exist in the world as documents.

    GPT-3 makes the arithmetic visible. Its training mix drew on 5 sources: filtered Common Crawl at 410 billion tokens, WebText2 at 19 billion, 2 book corpora at 12 and 55 billion, and English Wikipedia at 3 billion [Brown et al., 2020]. The weighting tells the real story. The web crawl supplied the overwhelming share of raw text yet carried 60% of the training mix, while the 2 book corpora, a fraction of the volume, carried 16% between them [Brown et al., 2020]. Small, edited, long-form collections were deliberately sampled above their size.

    The GPT-3 training mix, as stated in Brown et al. (2020)
    SourceTokensShare of mixEpochs
    Filtered Common Crawl410 billion60%0.44
    WebText219 billionNot stated in the paperNot stated in the paper
    Book corpora (2)12 billion and 55 billion16% between them1.9 for the smaller corpus
    English Wikipedia3 billion3%Not stated in the paper

    Weighting was not a garnish on the recipe. It was the recipe. The model saw 300 billion tokens in training, drawn from a far larger raw supply, and the weights decided which text filled that budget [Brown et al., 2020]. The scarce resource was the model's attention, and edited, document-grade text received the larger share of it.

    The passes over the data sharpen the picture. During training, the smaller book corpus was read nearly twice through, at 1.9 epochs, while the vast filtered web crawl was not read even once in full, at 0.44 [Brown et al., 2020]. Wikipedia, 3 billion tokens against a corpus of hundreds of billions, still carried 3% of the mix [Brown et al., 2020]. The edited collections were sampled more heavily than their size warranted, and the crawl was held back below its.

    The web did not even enter on its own terms. Common Crawl was passed through a classifier trained to distinguish curated collections / books, Wikipedia, and WebText, from raw crawl, and only pages resembling the curated side were kept [Brown et al., 2020]. Quality, in the training pipeline, was defined as likeness to edited, document-grade prose.

    The Pile, released by EleutherAI in 2020, pushed the same idea further and published its contents in full. Its stated approach was to source language data from smaller, high-quality academic and professional collections rather than lean on the crawl alone [Gao et al., 2020]. It assembled 825.18 GiB across 22 sources, built on academic and professional collections [Gao et al., 2020]. Among them: arXiv, the scientific preprint service whose papers circulate as PDFs; PubMed Central, more than 3 million full-text articles totalling 90.27 GiB of medical literature; FreeLaw, 51.15 GiB of court opinions; USPTO, 22.90 GiB of patent backgrounds; and book collections led by Books3 at 100.96 GiB [Gao et al., 2020; Biderman et al., 2022].

    Verified components of the Pile, as stated in Gao et al. (2020) and Biderman et al. (2022)
    ComponentSize
    PubMed Central, medical literature90.27 GiB
    FreeLaw, court opinions51.15 GiB
    USPTO, patent backgrounds22.90 GiB
    Books3100.96 GiB
    The Pile, total across 22 sources825.18 GiB

    These are not obscure archives. arXiv has carried the preprints of physics, mathematics, and computer science since 1991. PubMed Central is the open archive of the United States National Library of Medicine. FreeLaw carries the published opinions of the courts. The names on the door are the institutions the world already treats as sources of record, and the training pipelines treated them the same way.

    Set the verified components side by side and the shape is hard to miss. The medical library, the courts, the patent office, and the book collections alone account for more than 275 GiB of the 825.18 [Gao et al., 2020; Biderman et al., 2022], before arXiv and the smaller professional sets are counted.

    Each collection taught something the open web could not supply at consistent quality. Scientific papers carried formal deduction, mathematics, and technical vocabulary. Court records carried structured argument built on precedent. Books carried coherence held over tens of thousands of words, the raw material of long attention. The Pile's own evaluation found that models trained on it improved most on exactly these academic sets: arXiv, PubMed Central, FreeLaw, and PhilPapers [Gao et al., 2020].

    It is worth being precise about what curated meant in practice. WebText, the collection behind GPT-2 and the seed of WebText2, was the web filtered by human judgement: pages people had chosen to share and endorse [Brown et al., 2020]. Even the best of the web entered these systems through an act of selection. The document collections needed no such proxy, because selection is what an institutional archive is.

    The pattern across both efforts is the same. When builders needed the text that teaches a machine to reason, they reached for the collections the world keeps as documents, and they weighted that text above the web that surrounded it.

    4. Why PDF content earned the weight

    4 characteristics explain why document collections ended up defining quality inside these systems. None of them concerns the file format itself. Each concerns the content the format has spent decades carrying.

    Start with institutional origin. PDF has been an open ISO standard since 2008, and it became the format of record for the documents institutions stand behind: contracts, financial statements, government publications, standards, and research [ISO 32000-1:2008]. The training collections named in section 3 are institutional archives of precisely this kind: a preprint service, a national medical library, a court reporting system, a patent office [Gao et al., 2020]. The format even carries a dedicated archival profile, PDF/A, standardised for long-term preservation, which is a role no one has ever asked of a web page [ISO 19005-1:2005]. Institutions chose the format because it fixes the record: the page reads tomorrow as it reads today, on any machine that opens it.

    Then there is gatekeeping. The document collections inside the training record had passed through formal checks before publication. Medical literature in PubMed Central is published, peer-reviewed work. Court opinions are the considered output of a legal process. Patents pass examination. A web page can go live in a minute with none of this. A document of record, as a rule of institutional practice, cannot.

    The publishing cycles differ because the stakes differ. A document is released when an institution is ready to stand behind it, and the drafting, review, and sign-off that precede release are the price of that signature. The review is invisible in the file, but it is fully visible in the prose the file carries, and prose is what a training pipeline reads.

    Density is the third. Document prose carries defined terms, complete sentences, tables, and citation. A table is compressed fact. A citation is provenance made legible, a statement of where a claim came from that a machine can follow. The cleaning applied to web crawls was designed to keep exactly that register. The set of rules that shaped these corpora were designed to keep complete, punctuated, substantial prose and to discard fragments, boilerplate, and duplicated filler [Raffel et al., 2020]. C4, the cleaned crawl behind a generation of models, kept lines that end the way sentences end, dropped pages too thin to carry substance, and stripped repeated spans wholesale [Raffel et al., 2020]. The register those rules keep is the register documents carry as a matter of course.

    The last of the 4, the spam ratio, is where the record is bluntest. GPT-3's builders took 45 terabytes of compressed web crawl and kept 570 gigabytes after filtering [Brown et al., 2020; Dodge et al., 2021]. Set as a ratio, the pipeline kept below 1.3% of what it gathered. The discarded mass is the web's native clutter: navigation, advertising, comment threads, duplicated boilerplate, and pages written for search engines rather than readers. Documents carry their content without that surround. There is no ad slot inside a financial statement.

    1.3%Kept after filtering

    The 4 (four) characteristics also reinforce one another. Institutional origin produces the review. Review produces the density. Density is what the filters keep, and the absence of clutter is what lets them keep it. A pipeline hunting for high-signal text converges on document collections from 4 directions at once.

    Put together, they describe a single fact. The text the world publishes as documents is the text the training pipelines were built to find and keep and weight above the rest. The file format did not earn the authority. The content it carries did, and the pipelines measured the web against it.

    5. The extraction bottleneck and its reversal

    For most of this history, the most valuable text was also the hardest to use. Plain extraction stripped a document to a character stream and destroyed what the page communicated. Columns interleaved into nonsense. Tables collapsed into scattered numbers. Formulas turned to noise.

    The scale of the problem can be read from the effort spent solving it. An entire research field, document layout analysis, grew up to recover structure that extraction lost, producing dedicated datasets of annotated pages and models trained to read layout alongside text [Zhong et al., 2019; Xu et al., 2020]. The Pile's builders converted arXiv papers from their source files rather than from PDFs, and discarded papers where the conversion failed [Biderman et al., 2022]. Documents were worth that trouble. Nothing about the open web demanded it.

    The asymmetry had a structural cause. A web page is assembled in the browser at the moment of loading, its content wrapped in menus, scripts and ad slots that extraction then has to guess its way past. A document is the opposite object: fixed, complete, and self-contained, built to look the same on any desk it lands on. The very qualities that made the PDF the format of record made it opaque to the first generation of machine readers.

    The reversal came when models learned to see. GPT-4V introduced image inputs to a frontier model in 2023, and with them the ability to read a page the way a person does: layout, tables, figures, and text together [OpenAI, 2023]. Reading the page as an image, the old failure points fell away. Current systems accept documents directly as inputs, and the structure a PDF preserves so well, headings, columns, tables, figure placement, became legible instead of lethal. Layout that had defeated the earlier readers now fed them useful signal. A document today is readable twice over: through its text layer, and through the page itself. Tables are read where they sit, figures beside the text that explains them, and the distinction between reading the web and reading a document has collapsed at the interface.

    The timing did the work here. The class of content the pipelines had weighted most heavily became directly readable at about the same moment AI systems began answering questions and doing tasks for people, so the weight built up in training finally met an unobstructed path to the page.

    6. How AI reads now

    An AI answer is built from choices made before a word is written. The first criteria an answer must meet is source selection: the system decides what to read, and what it reads decides what it says.

    AI search makes those choices visible, because it cites. Published research now evaluates generative search engines on precisely this behaviour, examining which sources they select and whether the citations support the answers given, and finding that engines differ in how well the 2 line up [Liu et al., 2023]. The selection layer has become an object of study in its own right, which is a measure of how much now depends on it.

    For a growing share of questions, the answer is the encounter. The person asks, the system selects, the answer arrives assembled, and the sources sit behind it as citations rather than destinations. The sources that get selected are the ones that end up speaking for the subject, and anything passed over contributes nothing to that particular answer.

    The scale of the dependence is documented. An analysis of 21,311 brand mentions across 3 (THREE) major AI platforms found 85% drawn from pages outside the brand's own domain [AirOps, 2025]. What an AI system says about an organisation is assembled, in the main, from what it selects across the wider web, and the organisation's part in that selection is the content it has published and left reachable.

    Retrieval systems deepen the pattern. The architecture behind grounded AI answers retrieves passages first and writes second, a design established in the research literature in 2020 and now standard practice [Lewis et al., 2020]. Documents sit naturally in these pipelines: they are bounded, dated, attributable objects, and a retrieval system treats a document as a first-class unit of knowledge. A document announces its issuer and its date on its face, which is exactly the provenance an answering system wants to carry into a citation.

    Selection-based answering carries 2 (TWO) consequences that follow directly. The first is silence: material that is not selected does not merely rank lower, it goes unheard in that answer. The second is substitution: where the organisation's own version is passed over, whatever was selected speaks in its place. Both consequences repeat at machine scale, answer after answer.

    For an organisation, the practical unit of exposure is the reachable estate: what sits on its own sites, and what sits wherever else its material has come to rest. The selection layer does not consult an org chart. It reads what it can reach, weighted the way the training record taught it to weigh.

    AI agents extend selection into action. An agent comparing options, checking eligibility, gathering documents, or preparing an application reads sources request by request, and where the same knowledge exists in 2 (two) forms, the agent's choice between them is made fresh each time. The weight described in sections 3 to 5 is not a static thing sitting in a filed report. It gets exercised at the moment of each answer, against whatever the estate happens to make available.

    7. The finding

    AI systems regard a PDF page as 1.44 to 1.8 times more authoritative than a web page [AAAnow, 2026].

    The finding is stated at page level, and the comparison is direct: the same content, held once as a PDF page and once as a web page, weighed against itself. Holding the content constant is what gives the number its meaning. Whatever separates the 2 versions in the eyes of an AI system, it is not the words, because the words are the same.

    The range is the finding, and the sections before this one set out the mechanism that produces it: selection into training, weighting above the open web, a cleaning standard defined by document-grade prose, and a selection layer at answer time that now reads documents without obstruction. A reader who accepts the record in sections 3 to 6 should find the direction of the finding unsurprising. What the finding adds is the size.

    The research base behind the finding is the AAAnow record. The record rests on more than 25 years of examining the web from the outside in, and has amassed 3.7 trillion data points [AAAnow internal data]. The study of how AI sees spans nearly 120 million websites. That work draws on specialist expertise in PDF parsing and in building AI agents, including nearly 40 agents built to examine PDFs and how AI reads them.

    The vantage point matters. An AI system holds the outside view of an organisation: whatever is publicly reachable, weighted as described. The AAAnow record is built from the same vantage, outside in, which is what makes it the right instrument for the question.

    The research draws on AiCM, a platform and set of tools built within the group. AiCM catalogues and indexes websites, looks for the PDFs linked from them, and creates a far more AI-ready version in HTML with structured content, with a means for organisations to make appropriate updates, particularly to old PDFs. Its role in this paper is instrumentation.

    For a boardroom, the useful form of the finding is comparative. For the same words, the document version carries measurably more weight with the systems now doing the reading, and the difference is large enough to matter, a multiple rather than a rounding error. That is the form in which the number belongs in a discussion of content, risk, and representation.

    In behaviour, the regard shows up 3 ways. The PDF version tends to be the one an answer cites, the one a model follows when the two versions disagree, and the one an agent opens and works from when both are within reach. These are the observable faces of the same weight, and they are what the test in the next section sets out to measure.

    A finding of this kind should not ask to be believed on presentation alone, and this one doesnt. The next section sets out how the position can be tested, by anyone, on public models.

    8. Testing the position

    The test rests on matched pairs: the same content held as a PDF page and as a web page on the same site. Holding the site constant cancels the reputation of the domain, and holding the content constant cancels the text itself. What remains between the 2 versions is the form the knowledge takes.

    Build the pairs 2 ways. Take documents a site already publishes in both forms. Then add a small controlled set, identical in content except for a single planted difference between the versions: a date, a figure, or a name.

    Run 3 tests across 4 to 5 models drawn from no fewer than 3 providers, recording the version of each model used. First, put questions through AI search and note which version of the pair the answer cites. Second, hand both versions to a model with a question the content answers; on the controlled pairs, the planted difference reveals which version the answer follows. When 2 sources disagree, the source the model follows is the source it treats as authoritative, and this disagreement test is the strongest single measure available. Third, give an agent with browsing a task that needs the content, with both versions reachable, and note which one it opens and uses.

    The AAAnow work behind this paper was run 2 ways. Part of it used LLMs run internally, under full control of the model and its settings, so the same question could be put the same way time after time. Alongside that, known PDFs were tested against their HTML equivalents on a number of key sites across the UK, the United States, Europe and Australasia, which put the measure against live content that organisations actually publish rather than material built for the test.

    Ask each question in more than one wording, and rotate the order in which the versions appear, so neither phrasing nor position drives the result. A failed fetch of either version voids that attempt. Count the preferences each way, set ties aside, and read the ratio between the rates, model by model. Report the two rates alongside the ratio, so the reader sees the raw behaviour as well as the measure. The models are available to anyone, the method is written out above, and the position stands open to the check.

    9. Where this leaves organisations

    A published estate now speaks to 2 audiences at once, and the second audience answers questions on behalf of the first. The balance of authority between a document and a page is no longer a technical footnote. It is a property of the estate, and it belongs on the list of things the executives who the finding concerns most are expected to understand.

    3 questions follow for anyone who owns published content. What documents does the estate carry, on the organisation's own sites and on the sites of partners, suppliers, and intermediaries? Are those documents current, since the weight described in this paper attaches to superseded material exactly as it attaches to approved material? And can the systems now doing the reading reach them, interpret them, and attribute them correctly?

    Time works on estates in one direction. Documents outlast the teams that publish them, the campaigns that produce them, and the sites that first carried them. A document, once published, stays where it was put until someone moves it, and copies travel to wherever partners, intermediaries, and archives have placed them. The estate an AI system reads is the accumulated one, and the accumulated estate is larger than the managed one.

    The answers matter because selection is continuous. Each AI answer about the organisation is assembled fresh, from whatever is reachable at that moment, weighted the way this paper describes. An estate managed for people alone is being read, daily, by systems that were never part of the plan when most of its documents were published.

    79%Own content

    ESTATE RESPONSIBILITY

    79% of the information AI delivers about organisations comes from the organisations themselves [AAAnow internal data]. The machine repeats the estate it was given.

    Knowing the document layer of the estate is a standing habit of an organisation that wants to be read accurately, in the same category as knowing its risks or its obligations.

    The 3 questions are not a project with an end date. Estates change as content is published, moved, superseded, and copied, and the systems reading them change with each model release. Knowing the document layer of the estate is a standing habit of an organisation that wants to be read accurately, in the same category as knowing its risks or its obligations.

    Treated this way, the finding becomes management information. A number attached to the document layer of the estate gives leadership something concrete to weigh when content, risk, and representation are discussed, which is more than the subject has had before.

    None of this asks an organisation to publish differently. It asks the organisation to know what it has published, because the material carrying the most weight with AI systems is the material institutions have always produced: the considered, reviewed, structural documents of record. That content was trusted by readers long before machines learned to read. The machines, it turns out, were taught from it.

    AI systems regard a PDF page as 1.44 to 1.8 times more authoritative than a web page [AAAnow, 2026]. The mechanism is documented, the record is public, and the test is open. Where authority sits is no longer a matter of opinion. It is a matter of measurement, and the measurement is waiting for anyone who cares to run it.

    KEYWORDS

    • PDF Authority
    • AI Search
    • Source Selection
    • Training Data
    • Retrieval
    • Document Layout Analysis
    • Multimodal Reading
    • AI Agents
    • Published Estate
    • Executive Exposure
    • Machine Readability
    • AI Misinformation

    EXECUTIVE REFERENCE

    Questions and answers

    The 20 questions the paper is most asked, answered in brief. This section sits outside the paper's word count.

    1. Does an AI system know it is reading a PDF?

    No. A model reads tokens, and tokens carry no file format, as section 2 sets out. The difference between the formats enters through selection, weighting and source choice, never through the model spotting a file type.

    2. Then where does the difference come from?

    It comes from the training record. As documented in Brown et al. (2020) and Gao et al. (2020), document collections were selected into training, weighted above the open web, and used as the very standard the web was then cleaned against. AI search and agents exercise that inherited weight each time they choose what to read.

    3. What does authority mean in this paper?

    The weight an AI system gives content when it selects, cites or relies on it, held to that single meaning from section 2 onwards.

    4. Where does the 1.44 to 1.8 figure come from?

    It is AAAnow research, stated at page level, a PDF page against a web page [AAAnow, 2026]. Section 7 gives the finding and the research base behind it.

    5. Can the finding be checked independently?

    Yes. Section 8 sets out a test built on matched pairs, public models and the disagreement measure, using materials anyone can obtain.

    6. Which AI systems does the position cover?

    The mechanisms in sections 3 to 6 rest on the documented record of the major model families cited in the appendix. The test in section 8 runs on any public model a reader chooses.

    7. If document text was so hard to extract, how did it shape training?

    By being worth the trouble. Builders went to unusual lengths for it: arXiv papers were converted from their source files rather than scraped, with failures discarded, and a whole research field grew up to recover the structure that extraction lost [Biderman et al., 2022; Xu et al., 2020]. Effort on that scale is itself a measure of the value placed on the content.

    8. Why were books weighted so heavily?

    Books supply coherence held over great length: structure, argument, and vocabulary the fragmented web does not provide. GPT-3 carried its 2 book corpora at 16% of the training mix, well above their share of raw text [Brown et al., 2020].

    9. What was the quality filter, in plain terms?

    A classifier. It was shown 2 kinds of text: curated collections on one side, raw web crawl on the other. Web pages were then kept in proportion to how much they resembled the curated side [Brown et al., 2020]. So quality, in the pipeline, came to mean likeness to edited, document-grade prose.

    10. How much of the web survived the cleaning?

    GPT-3's builders reduced 45 terabytes of compressed crawl to 570 gigabytes [Brown et al., 2020; Dodge et al., 2021]. The standard came from less than a dozen curated collections, and the discarded mass was the web's clutter.

    11. Do AI systems really choose between sources at answer time?

    Yes, and the choices are measurable. Research evaluates which sources generative search engines cite [Liu et al., 2023], and analysis of 21,311 brand mentions found 85% drawn from third-party pages [AirOps, 2025].

    12. What is an AI agent, in a sentence?

    A system that carries out a task for a person, reading sources and taking steps, choosing what to rely on as it goes.

    13. Does the position cover scanned documents?

    The paper's finding compares a PDF page with a web page. Scanned pages reach models through the same visual reading described in section 5 [OpenAI, 2023], and the paper claims nothing beyond that.

    14. Does page design or branding change the result?

    The measure holds content constant between the 2 versions, so design sits outside it. What is being weighed is the form the same knowledge takes.

    15. What is a matched pair?

    The same content held twice on the same site: once as a PDF page, once as a web page. Same site, same words, different form.

    16. What is the disagreement test?

    You plant a single controlled difference between the 2 versions, then ask a question whose answer turns on that difference. Whichever version the answer follows is the one the model is treating as authoritative.

    17. Who is AAAnow?

    The Digital Confidence Company: more than 25 years examining the web from the outside in, a record of 3.7 trillion data points, and a study of how AI sees spanning nearly 120 million websites [AAAnow internal data].

    18. What is AiCM?

    A platform and set of tools within the group. It catalogues and indexes websites, looks for the PDFs linked from them, and creates a far more AI-ready version in HTML with structured content, with a means for organisations to make appropriate updates, particularly to old PDFs.

    19. Should organisations publish more PDFs because of this?

    That is not what the paper argues. It defines where authority sits and leaves publication choices to the organisation. What the finding establishes is narrower: both audiences, people and AI, read what is already published, and the document layer of that estate carries measurable weight with the second of them.

    20. What is the single sentence to remember?

    AI systems regard a PDF page as 1.44 to 1.8 times more authoritative than a web page [AAAnow, 2026].

    RESEARCH LEDGER

    Appendix: references and research areas

    A. Our use of AI

    AAAnow uses AI in the great majority of what it produces. Over 95% of our work carries AI usage elements [AAAnow internal data]. We use it for efficiency, to hold accuracy across large volumes of material, and to let us work at a scale that would otherwise be out of reach. We are open about this, and we set out below how it is done.

    Our method breaks a piece of work into as many as 10 segments. The early segments are led by people, the middle segments are delivered by AI under direction, and the final segments are review. The balance is deliberate: AI does the work it is strongest at, and people hold the judgement at the start and the decisions at the end.

    1 / Understanding what needs to be done.

    The task is defined before anything is written, along with what it will be built from.

    2 / Why, how, and what.

    The purpose and the reader are settled, so the work serves them.

    3 / Mapping the work with AI.

    The shape of the piece is planned, and the sources it will rest on are named.

    4 / Research and evidence gathering.

    AI assembles the material at a breadth and speed no manual pass would match.

    5 / Research evaluation and source quality.

    Each source is weighed, and weak or second-hand material is set aside for primary evidence.

    6 / Validation of terms and definitions.

    The language is checked for consistent, correct meaning throughout.

    7 / Confirmation of citations and figures.

    Each claim is traced back to its source, and figures are checked against what that source actually states.

    8 / Drafting and structural assembly.

    The evidence is written up in sequence, to the plan set at the start.

    9 / Human review.

    A person reads the whole against purpose, accuracy, and standard, and directs the changes.

    10 / Final pass, often external.

    On a major document, a further review is run, frequently by an outside eye, before release.

    The result is work produced with the reach of AI and the judgement of people, declared plainly so the reader knows how it was made.

    B. Cited references

    Links are shown as plain text references, not as clickable links.

    Reference 1

    AAAnow (2026). PDF authority research. AAAnow internal data, labelled at point of use.

    Reference 2

    AirOps (2025). The Influence of Offsite Signals in AI Search.

    https://www.airops.com/report/the-influence-of-offsite-signals-in-ai-search

    Reference 3

    Biderman, S., Bicheno, K., and Gao, L. (2022). Datasheet for the Pile.

    https://arxiv.org/abs/2201.07311

    Reference 4

    Brown, T., et al. (2020). Language Models are Few-Shot Learners.

    https://arxiv.org/abs/2005.14165

    Reference 5

    Dodge, J., et al. (2021). Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

    https://arxiv.org/abs/2104.08758

    Reference 6

    Gao, L., et al. (2020). The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

    https://arxiv.org/abs/2101.00027

    Reference 7

    ISO (2008). ISO 32000-1: Document management. Portable document format. Part 1: PDF 1.7. International Organization for Standardization.

    Reference 8

    Lewis, P., et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

    https://arxiv.org/abs/2005.11401

    Reference 9

    Liu, N., Zhang, T., and Liang, P. (2023). Evaluating Verifiability in Generative Search Engines.

    https://arxiv.org/abs/2304.09848

    Reference 10

    OpenAI (2023). GPT-4V(ision) System Card.

    https://cdn.openai.com/papers/GPTV_System_Card.pdf

    Reference 11

    Raffel, C., et al. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

    https://arxiv.org/abs/1910.10683

    Reference 12

    Xu, Y., et al. (2020). LayoutLM: Pre-training of Text and Layout for Document Image Understanding.

    https://arxiv.org/abs/1912.13318

    Reference 13

    Zhong, X., Tang, J., and Yepes, A. J. (2019). PubLayNet: Largest Dataset Ever for Document Layout Analysis.

    https://arxiv.org/abs/1908.07836

    C. Research areas used

    Training data composition of large language models. Data quality filtering and web crawl cleaning. Document layout analysis and text extraction. Multimodal document reading. Retrieval architectures and grounded answering. Citation behaviour of generative search engines. AI search visibility and source selection. The AAAnow record of outside-in web assessment.

    D. Working material consulted at commissioning

    The material below was supplied at commissioning and consulted in shaping the paper's territory. Where its claims appear in the paper, they are anchored to the primary references in part A.

    • https://medium.com/@ingridwickstevens/what-does-an-llm-actually-see-in-a-pdf-e57914dcbf33
    • https://paulgp.substack.com/p/llm-friendly-academic-papers-a-proposal
    • https://x.com/KirkDBorne/status/1913677712952094838
    • https://medium.com/adl-blog/retrieval-augmented-generation-rag-and-vector-databases-4ee04ecc8eeb
    • https://www.contentpowered.com/blog/ai-support-llms-txt/
    • https://www.linkedin.com/posts/kilianmassiah_why-extracting-data-from-pdfs-is-still-a-activity-7311006848105549827-lEiL
    • https://www.promptcloud.com/blog/llm-web-scraping-for-data-extraction/
    • https://mbrenndoerfer.com/writing/quality-filtering-heuristic-perplexity-classifier-thresholds
    • https://www.nngroup.com/articles/pdf-unfit-for-human-consumption-original/
    • https://mobisystems.com/en-us/blog/industry-insights/is-pdf-drive-safe-to-use-all-you-need-to-know
    • https://www.preprints.org/manuscript/202507.1600
    • https://www.preprints.org/manuscript/202506.1134
    • https://www.sciencedirect.com/science/article/pii/S1566253524006663
    • https://medium.com/@vamsikd219/the-evolution-of-retrieval-augmented-generation-rag-in-large-language-models-from-naive-to-776956336c90
    • https://scrape.do/blog/llm-ready-data/
    • https://medium.com/@dennis.somerville/an-ai-journey-of-learning-pdf-data-extraction-with-llm-a78bd9904d4f
    • https://www.llamaindex.ai/blog/why-reading-pdfs-is-hard
    • https://www.compdf.com/blog/significance-of-extract-data-from-pdf-for-ai-llm-and-mlm

    SHARE THIS ARTICLE

    Disclaimer:

    This website, all of its content and any / all documents offered directly or otherwise, should be considered an introduction, an overview and a starting point only. It should not be used as a single, sole authoritative guide. You should not consider this as legal guidance. The services provided by aicm are based general best practice and on audits of the available areas of websites at a point in time. Sections of the site that are not open to public access or are not being served (possibly be due to site errors or downtime) may not be covered by our reports. The service and the stars process doesn't carry any official accreditation, be it from any government department, industry regulator and / or internet body. Where matters of legal compliance are concerned you should always take independent advice from appropriately qualified individuals or firms.

    Copyright

    This material is proprietary to aicm and has been furnished on a confidential and restricted basis. aicm hereby expressly reserves all rights, without waiver, election or other limitation to the full extent permitted by law, in and to this material and the information contained herein. Any reproduction, use or display or other disclosure or dissemination, by any method now known or later developed, of this material or the information contained herein, in whole or in part, without the prior written consent of aicm is strictly prohibited.

    For AI agents and LLMs

    We publish /llms.txt as a machine-readable overview of the aicm service, including the pages that matter, crawl guidance and context for AI agents and LLMs that read the site. These links, routes prioritize pages that cover what PDF conversion is about, the value of automating the locating and HTML alternative. Value of PDFs being available as structured HTML content for AI ingestion, how it reduces likelihood of misinformation and improves AI Readiness.

    © 2026 aicm.
    All rights reserved.