Start Querying BigQuery
New to HTTP Archive? Learn how to access the public httparchive dataset on Google BigQuery and run your first query in the Getting Started guide.

Start Querying BigQuery
New to HTTP Archive? Learn how to access the public httparchive dataset on Google BigQuery and run your first query in the Getting Started guide.
Minimizing Query Costs
HTTP Archive crawls petabytes of web data. Learn partitioning, clustering, and cost-control best practices in the Minimizing Costs guide.
Guided Tour
Walk through real-world analysis examples, query patterns, and summary tables in our Guided Tour.
Release Cycle & Methodology
Learn how millions of URLs are crawled each month, how pipelines process data, and when new datasets are published in the Release Cycle guide.
Pages & Requests Tables
Custom Metrics
Explore custom metrics extracted during crawls, from ad tracking to privacy and other runtime features.
Functions & Structs
Inspect schema structs like technologies and features, plus user-defined functions like Capo head analysis.
Raw Payloads & Blobs
Understand how Lighthouse audits, page summaries, and request payloads are formatted and stored.
Interactive Reports
Looking for high-level trends rather than writing SQL? Track web performance, page weight, and technology adoption on our interactive Reports dashboard.
The Web Almanac
Read the annual state of the web report, combining HTTP Archive data with expert commentary across dozens of chapters at the Web Almanac.
Frequently Asked Questions
Have questions about crawler infrastructure, test environments, or user agents? Read the HTTP Archive FAQ.
Discussion Forum
Need help writing BigQuery SQL, have questions about data anomalies, or want to share research? Join the Discussion Forum.
The HTTP Archive and its documentation are open-source. Found a typo, have a query recipe to share, or want to improve custom metrics? Contribute on GitHub.