Skip to main content
    Back to Resources

    Jan 11 2026

    Modelling PDF usage on US local government websites

    A defensible model for PDF downloads per day on a typical US local government website with around 1,000 PDFs, anchored to published municipal traffic data and long-tail usage evidence.

    This briefing models likely PDF download volumes on a US local government website with around 1,000 PDFs in its document repository. It draws on published municipal traffic figures, public-sector download benchmarks and UK long-tail usage evidence used as a structural prior.

    Executive summary

    A typical US local government web estate with around 1,000 PDFs is best modelled as a high-volume, low-observability system. The document inventory is large, but only a minority of documents drive most usage, and measurement is easily distorted by cross-domain hosting, caching and incomplete analytics instrumentation. Evidence from a recent US city website procurement document shows this inventory size is realistic in practice. The City of Solvang reports 1,000 plus documents in its Document Center alongside roughly 200 active pages, and expects only 300 to 500 PDFs to be prioritised for migration and remediation. A large share of the library is therefore legacy or low-priority content.

    Using published municipal web-traffic figures from two US local governments, one small and one mid-sized, and bounding the missing download conversion rate using public-sector examples where document download counts are explicitly reported, a defensible estimate for PDFs served per day for a US local government website with around 1,000 PDFs is:

    • Central estimate: around 100 to 400 PDF downloads per day for an everyday services municipal site.
    • Wide sensitivity range: around 10 to over 1,000 downloads per day, depending on the site's traffic scale and the fraction of sessions that trigger document retrieval. Forms-heavy sites and benefits portals sit at the high end.

    Calibration anchors

    This range is consistent with the following published anchors:

    • Small-city traffic anchor: the City of Greenbelt reports 42,236 pageviews over 23 days, around 1,837 pageviews per day, in a monthly report.
    • Mid-sized city traffic anchor: the City of Dunedin reports 195,105 pageviews and 91,900 sessions in February 2023, around 6,968 pageviews per day and 3,282 sessions per day.
    • Low conversion anchor (project-style site): an evaluation report shows 125 PDF downloads against 22,246 unique pageviews, about 0.6 percent downloads per unique pageview during the campaign window.
    • High conversion anchor (services-heavy site): a monthly public dashboard shows 479,732 pageviews with Top Downloads category totals summing to around 43,713 document downloads, about 9.1 percent downloads per pageview.

    Per-PDF usage is necessarily non-uniform. UK councils explicitly quantify this long tail, where many PDFs are almost never accessed, and US local governments show similar operational behaviour by prioritising a minority of documents for remediation and migration.

    Sampling logic for US local governments

    The US universe of local governments is large and heterogeneous. The U.S. Census Bureau classifies local government units into five major types (county, municipal, township, school district, special district), and reports roughly 90,000 local government units in the Census of Governments universe.

    A representative sampling frame for document-usage modelling should stratify at minimum by:

    • Government type (municipal versus county is often sufficient for web estates; special districts can be document-heavy but are operationally different).
    • Region (Northeast, Midwest, South, West) as defined by Census geography.
    • Population and service intensity as a proxy for web traffic and document demand, with a separate stratum for tourism-heavy small cities where traffic can exceed what resident population suggests, as explicitly stated in Solvang's RFP.

    Published evidence on traffic, PDF inventories and downloads

    The table below intentionally separates traffic (sessions, pageviews) from document inventory and document download counts, because these are very often published in different documents, by different teams, and sometimes not published at all.

    JurisdictionRegion / typePeriodPublished trafficInventory / downloads
    New York CityNortheast / municipalFY2025NYC.gov pageviews 276,976,100; unique visitors avg monthly 6,170,000Not stated in same source
    City of DunedinSouth / municipalFeb 2023Sessions 91,900; Users 63,336; Pageviews 195,105Not stated
    City of GreenbeltSouth / municipal1 to 23 Jun 2022Users 13,328; Sessions 18,934; Pageviews 42,236Not stated
    City of SolvangWest / municipalRFP dated Nov 2025Not published200 plus active pages; 1,000 plus documents in Document Center; migration prioritises 300 to 500 PDFs
    Colorado SpringsWest / municipalProgress updateNot published in cited snippetConverted 261 total documents from PDF to HTML (transition plan documents)

    Public-sector baselines for download conversion rates

    Two public-sector reports provide authoritative order-of-magnitude anchors for the ratio of document downloads to pageviews. They are not both local government, but they are methodologically relevant because they are explicit about downloads, and local governments often use the same analytics tooling.

    • Low download intensity example: an evaluation report for a government-run programme site reports 125 PDF downloads and 22,246 unique pageviews during the campaign period, implying about 0.6 percent PDF downloads per unique pageview.
    • High download intensity example: a publicly posted monthly analytics dashboard reports 479,732 pageviews and category download totals that sum to around 43,713 document downloads for the month, implying about 9.1 percent document downloads per pageview.

    These two anchors justify using a wide prior for document-download conversion when estimating municipal PDF demand from municipal traffic alone.

    Measurement architecture and known distortions

    In practice, datasets often mix:

    • Link-click based measures (for example, analytics file_download events), which track clicks to a PDF link but miss direct opens and can be blocked by consent settings.
    • HTTP request logs (origin or CDN), which capture direct opens and embeds but are inflated by bots, prefetching and byte-range requests for large PDFs.

    GA4 supports enhanced measurement events and can track file downloads when enabled, but it is event-based and depends on the client-side tag actually firing. Matomo documents automatic download tracking and explains trade-offs between JavaScript tracking and log-based analytics.

    US local governments frequently deploy CMS and engagement stacks where documents are stored in platform-managed repositories and sometimes served through a CDN. CivicPlus describes Document Center as a module that organises and manages documents in one central repository. Solvang's RFP provides unusually concrete operational detail, noting the existing site is on the CivicPlus platform and requires an audit of Document Center PDFs.

    Caching, CDNs and why served PDF is not one number

    PDF serving is sensitive to caching because PDFs are large and often cacheable. Three specific measurement artefacts matter in practice:

    • Byte-range requests (HTTP 206): common for PDFs because viewers may fetch parts of a document for preview, inflating request counts relative to downloads. Detectable only in logs, not click events.
    • Direct opens and external referrers: if residents open a PDF directly from search, bookmarks or email, click-event tracking may not fire at all, making file_download events a lower bound in many real estates.
    • Cross-domain hosting: if PDFs are hosted on a different domain, subdomain or vendor-managed cloud storage, both click-tracking and log-collection require additional configuration, and ownership of logs may sit with the vendor.

    Statistical model for downloads per day and per PDF

    Let S be sessions per day, V be pageviews per day, ρ = V / S be pageviews per session, r be document downloads per pageview (a conversion rate), and D be expected document downloads per day. Then:

    D = V × r = (S × ρ) × r

    Published municipal data provides plausible ρ values: Greenbelt at approximately 42,236 / 18,934 ≈ 2.23, and Dunedin at approximately 195,105 / 91,900 ≈ 2.12. So for US municipal websites that publish these metrics, ρ between 2.1 and 2.3 is a defensible starting range.

    Bounding the conversion rate r

    The conversion rate varies by service mix:

    • Low (campaign or information site): around 0.6 percent downloads per unique pageview.
    • High (forms-heavy government site): around 9.1 percent downloads per pageview.

    This briefing uses Low r = 0.5 percent, Central r = 3 percent, and High r = 9 percent.

    From traffic to downloads per day

    Using the two municipal traffic anchors:

    • Small-city anchor (Greenbelt): V ≈ 1,837 pageviews per day. Estimated downloads per day: Low about 9, Central about 55, High about 165.
    • Mid-sized city anchor (Dunedin): V ≈ 6,968 pageviews per day. Estimated downloads per day: Low about 35, Central about 209, High about 627.

    These calculations yield a central band that is naturally in the 100 to 400 downloads per day range for many municipal sites, while still supporting a wider sensitivity interval.

    Per-PDF demand for a 1,000-PDF inventory

    ScenarioDownloads / day (D)Avg downloads per PDF / dayAvg downloads per PDF / year
    Low operational500.0518
    Central estimate2000.2073
    High services-heavy6000.60219

    This mean per PDF is not the typical PDF. Empirically, most PDFs are low-usage. UK councils explicitly quantify this, and the Solvang RFP implicitly signals the same reality by prioritising only 300 to 500 PDFs for migration and remediation despite 1,000 plus documents in the repository.

    Modelling the long tail across PDFs

    A common, evidence-backed structural assumption for content request patterns is a Zipf-like (power law) distribution, which has been demonstrated in web request traces and is directly tied to caching behaviour.

    Under a Zipf-like allocation of demand across N = 1,000 documents (illustrative exponent s ≈ 1), concentration is extreme. The top 10 documents account for around 39 percent of demand, and the top 100 for around 69 percent.

    Rank groupShare of total PDF demand
    Top 10 PDFs39.1 percent
    Ranks 11 to 5021.0 percent
    Ranks 51 to 1009.2 percent
    Ranks 101 to 2009.2 percent
    Ranks 201 to 50012.2 percent
    Ranks 501 to 1,0009.3 percent

    Comparative analysis with UK findings

    The UK local government ecosystem has published multiple explicit statements about PDF inventories and usage distributions, often to justify accessibility remediation strategy:

    • Manchester City Council reports that 4 percent of its document set was viewed more than 1,000 times, 67 percent fewer than 100 times, and explicitly notes it cannot track downloaded files and expects downloads to be lower than views.
    • Wealden District Council reports that 78 percent of PDFs had not been viewed in the last 30 days, and only 18 PDFs exceeded 50 views in that period.
    • Warrington Borough Council reports over 5,800 PDFs, with 77 percent viewed 50 times or fewer and 49 percent viewed fewer than 5 times.
    • Eden District Council publishes monthly website unique visitors and pageviews, including March 2023 at 40,644 unique visitors and 165,401 pageviews.

    US local governments more rarely publish document-level usage, but operational documents show similar realities. Solvang's RFP describes a municipal site with 1,000 plus documents, but expects only 300 to 500 PDFs to be prioritised for remediation and migration, indicating a large long tail of lower-priority content.

    Regulatory framing differs. The U.S. Department of Justice has issued guidance and rulemaking on web accessibility obligations for state and local governments. This matters because remediation and PDF-to-HTML conversion can shift usage from downloads to page or transcript views, changing measured rates even if resident demand is unchanged.

    Recommendations for robust modelling and data collection

    Define PDF served in two operationally useful ways:

    • Click-based downloads: counts of link-triggered file_download or equivalent events. Useful for content prioritisation, but not complete.
    • Request-based serves: cleaned counts of PDF retrievals from logs (origin plus CDN), deduplicated for range requests and filtered for bots. Useful for public-facing load and true demand.

    For a 1,000-PDF site, a standard inventory pipeline should include sitemaps (including multiple sitemap indexes), in-domain discovery for Document Center style repositories where PDFs may not appear in a standard page sitemap, filetype discovery as a backstop, and HEAD requests to confirm content-type, file size, last-modified and caching headers.

    Treat third-party hosting as first-class, not as an edge case. Assume documents will be split across the primary municipal domain, vendor-managed repositories such as platform Document Center modules, and engagement and records vendors. The minimal viable analytics configuration includes a GA4 cross-domain strategy where feasible and explicit testing that file-download events actually appear in reporting, plus vendor reporting exports where the vendor owns critical logs.

    Operational workflow

    1. Traffic estimate V (pageviews per day) from published or internal analytics.
    2. Conversion prior r in [0.5 percent, 9 percent] with a central value around 3 percent, until calibrated with logs.
    3. Compute D = V × r for a first-pass downloads per day estimate.
    4. Long-tail allocation using a Zipf-like or lognormal prior plus zero-inflation.
    5. Calibration using a short log capture window: pull CDN logs where available and reconcile with click-based events; measure the ratio of direct opens to click-based opens; fit the long-tail parameters to the observed rank-frequency curve.

    Accessibility and format conversion as a demand shifter

    If a council adopts automated PDF-to-HTML conversion, measured PDF downloads can fall while document consumption remains constant or increases, because consumption moves to HTML transcript views. In the US, this interacts with legal obligations for state and local government web content accessibility issued by the Department of Justice. Any longitudinal model should mark conversion and remediation milestones as structural breaks in the time series.

    Bottom line

    For a typical US local government site with around 1,000 PDFs, expect roughly 100 to 400 PDF downloads per day in a central scenario, with a wide sensitivity range of around 10 to over 1,000. Demand will be concentrated on a small minority of documents, and a large share of the inventory will see little or no use. Treat measured downloads as a lower bound until origin and CDN logs have been reconciled with click-based events, and treat any PDF-to-HTML conversion programme as a structural break in the time series.

    Sources cited in the underlying briefing include published reports and accessibility statements from the City of Solvang, City of Greenbelt, City of Dunedin, City of New York, Colorado Springs, Manchester City Council, Wealden District Council, Warrington Borough Council, Eden District Council, the U.S. Census Bureau, the U.S. Department of Justice, MTC, GA4 and Matomo documentation, MDN Web Docs on HTTP caching, CivicPlus product documentation, and Granicus support articles.

    Disclaimer:

    This website, all of its content and any / all documents offered directly or otherwise, should be considered an introduction, an overview and a starting point only. It should not be used as a single, sole authoritative guide. You should not consider this as legal guidance. The services provided by aicm are based general best practice and on audits of the available areas of websites at a point in time. Sections of the site that are not open to public access or are not being served (possibly be due to site errors or downtime) may not be covered by our reports. The service and the stars process doesn't carry any official accreditation, be it from any government department, industry regulator and / or internet body. Where matters of legal compliance are concerned you should always take independent advice from appropriately qualified individuals or firms.

    Copyright

    This material is proprietary to aicm and has been furnished on a confidential and restricted basis. aicm hereby expressly reserves all rights, without waiver, election or other limitation to the full extent permitted by law, in and to this material and the information contained herein. Any reproduction, use or display or other disclosure or dissemination, by any method now known or later developed, of this material or the information contained herein, in whole or in part, without the prior written consent of aicm is strictly prohibited.

    For AI agents and LLMs

    We publish /llms.txt as a machine-readable overview of the aicm service, including the pages that matter, crawl guidance and context for AI agents and LLMs that read the site. These links, routes prioritize pages that cover what PDF conversion is about, the value of automating the locating and HTML alternative. Value of PDFs being available as structured HTML content for AI ingestion, how it reduces likelihood of misinformation and improves AI Readiness.

    © 2026 aicm.
    All rights reserved.