Skip to content

HTTP Archive Documentation

Guides, schema references, and query recipes for analyzing how the web is built.

Start Querying BigQuery

New to HTTP Archive? Learn how to access the public httparchive dataset on Google BigQuery and run your first query in the Getting Started guide.

Minimizing Query Costs

HTTP Archive crawls petabytes of web data. Learn partitioning, clustering, and cost-control best practices in the Minimizing Costs guide.

Guided Tour

Walk through real-world analysis examples, query patterns, and summary tables in our Guided Tour.

Release Cycle & Methodology

Learn how millions of URLs are crawled each month, how pipelines process data, and when new datasets are published in the Release Cycle guide.

Pages & Requests Tables

Detailed column specifications, clustering keys, and JSON payload definitions for the core pages and requests tables.

Custom Metrics

Explore custom metrics extracted during crawls, from ad tracking to privacy and other runtime features.

Functions & Structs

Inspect schema structs like technologies and features, plus user-defined functions like Capo head analysis.

Raw Payloads & Blobs

Understand how Lighthouse audits, page summaries, and request payloads are formatted and stored.

Interactive Reports

Looking for high-level trends rather than writing SQL? Track web performance, page weight, and technology adoption on our interactive Reports dashboard.

The Web Almanac

Read the annual state of the web report, combining HTTP Archive data with expert commentary across dozens of chapters at the Web Almanac.

Frequently Asked Questions

Have questions about crawler infrastructure, test environments, or user agents? Read the HTTP Archive FAQ.

Discussion Forum

Need help writing BigQuery SQL, have questions about data anomalies, or want to share research? Join the Discussion Forum.

The HTTP Archive and its documentation are open-source. Found a typo, have a query recipe to share, or want to improve custom metrics? Contribute on GitHub.