Methods
Everything a reader needs to re-run this site: inclusion criteria, the frozen fields and the rules that derive them, the index definition, the sources and their terms, the corrections policy, the right of reply, and the limits of what the data supports.
Who is accountable
Editor: Jeremy Mays. Publisher and funder: Paciva AI. Paciva is an AI company publishing an index that measures coverage of AI companies. The conflict is stated first so every measurement can be read against it. There is no editorial board. There is no opinion column. Contact: editor@paciva.ai.
What is in and what is out
In: any article, broadcast segment, podcast episode, government record, filing, or research publication, in any language, that matches the published topic query (AI safety, AI risk, existential risk, frontier model, rogue AI, AI agents, superintelligence, artificial general intelligence, AI regulation) and is captured by a source listed on this page. Out: social media posts (except as a cited primary when the author is the subject), paywalled text beyond the headline, and any item whose publisher URL cannot be resolved. Wire pickup is merged by title similarity so one story counts once. The query string is published on this page and versioned with the taxonomy.
Fields and the rules that derive them
No field is a judgment about a story, an outlet, or a person, and no language model emits any field value. Each rule is published and any reader can re-run it.
Word sets for the verb classes: observed language observed, confirmed, demonstrated, detected, documented, found, recorded, occurred, completed, exfiltrated, accessed, downloaded, uploaded, published, signed, introduced, filed, raised, closed; assessed language assess, assessed, assesses, likely, attributed, attribute, attributes, consistent with, suspected, believed, estimate, estimates, estimated, indicates, suggests, suggestive; conditional language could, may, might, potential, potentially, possible, possibly, would, if, plausibly, risk of, threat of. One set matched gives that class; two or more gives Mixed; none gives No qualifying verb found.
Definition, weights, and sensitivity
Four hourly component series: coverage clusters (GDELT DOC, all languages, wire duplicates merged), broadcast minutes (GDELT TV Explorer), discourse volume (Hacker News, Wikipedia pageviews, podcast episodes), and front page prominence (GDELT Frontpage Graph). Each is divided by its own 90 day median and multiplied by 100, so 100 means the median day. Default weights: 0.4, 0.2, 0.2, 0.2. The index is the weighted mean of the four. It is undefined until 30 days of baseline exist. A composite hides its components, so the components are always shown beside it and the Workbench lets you set any weights. The sensitivity table below fills in once a baseline exists and reports the reading under five alternative weightings so a reader can see how much the number depends on the choice.
| Weighting | Coverage | Broadcast | Discourse | Front page | Reading |
|---|---|---|---|---|---|
| Default | 0.4 | 0.2 | 0.2 | 0.2 | awaiting baseline |
| Equal | 0.25 | 0.25 | 0.25 | 0.25 | awaiting baseline |
| News only | 1 | 0 | 0 | 0 | awaiting baseline |
| Broadcast heavy | 0.2 | 0.5 | 0.15 | 0.15 | awaiting baseline |
| Discourse heavy | 0.2 | 0.15 | 0.5 | 0.15 | awaiting baseline |
How data moves
A Cloudflare Worker cron fires every 15 minutes on offset minutes (7, 22, 37, 52) and enqueues one fetch job per source. A queue consumer fetches each source, deduplicates by canonical URL and title cluster, derives the fields above, and writes to a database. GDELT is queried once per topic per cycle with exponential backoff and a block window, because its throttling is undocumented; when it blocks, the site keeps the last good state and the header says how old it is. Headlines in languages other than English are translated by Cloudflare Workers AI (m2m100) and always shown beside the original, labeled machine translation. That is the only model call in the pipeline. After each cycle the current state is written into the HTML and the site is redeployed, so a crawler sees a complete page. One immutable snapshot per day is stored and served at a dated URL. The editorial tables (source registry, primary source records, headline pairs, corrections, replies) live in a spreadsheet the editor maintains, published as CSV and re-read every cycle, so the source universe is a public, diffable document.
Per source terms
Our analysis layer (the derived fields, the index, the snapshots) is CC BY 4.0. A permissive license on our layer does not change upstream terms, so each source's terms are listed. Cite as: GDELT Project (2026), processed by the AI Risk Coverage Monitor. Excluded on purpose: Reddit, NewsAPI.org free tier, Perigon free, MediaStack free, Brave default tier, Bing, and any Google News scraper, each for a licensing or storage reason recorded in the repository.
| Source | Terms as applied here |
|---|---|
| GDELT Project | Unlimited use, redistribution permitted, citation and link required |
| Publisher RSS feeds (OPML in the repo) | Headline, link, date only; no description text stored; robots.txt and per feed terms honored |
| Google News RSS | Discovery only; every item resolved to the publisher URL before storage |
| Congress.gov API | US government work, not subject to copyright |
| LegiScan Public API | CC BY 4.0 |
| Federal Register API | Discovery; govinfo.gov PDF cited as the record |
| FEC API and Senate LDA | US government records |
| arXiv | Metadata CC0; PDFs never mirrored |
| OpenAlex | CC0, polite pool with mailto |
| Hacker News (Algolia and Firebase) | Counts and links only |
| Wikimedia pageviews | Public attention proxy, counts only |
| Podcast Index | Core index free for any use |
| YouTube Data API | Curated channel playlists only, never search |
| Metaculus, Polymarket, Kalshi | Read only; prices labeled forecasts |
| MIT AI Risk Repository v4 | CC BY 4.0, taxonomy mapping |
| Cloudflare Workers AI, m2m100 | Headline translation only, labeled machine translation |
Policy
A correction is dated, states what was published and what is correct, and is shown on the item it corrects and on this page. Corrections are never silent. Anyone can request one at editor@paciva.ai. Corrections log: none yet.
For any named person or organization
Submit a reply and it publishes verbatim beside the item within 24 hours, with no editing. The contact field is used only to confirm you are who you say; it never publishes and it is not stored on this site.
What the data does not support
The Monitor does not independently verify third party reporting. A headline gap is a count of words, not a finding about accuracy; a short headline that paraphrases well can score high, and the count is shown so that a reader can check it against the two texts. The verb class rule is a word list, and a document can use conditional words for a confirmed event or observed words for a hypothesis. Volume counts depend on the topic query and on GDELT's coverage of each language. Machine translation of headlines can be wrong. Market forecasts are estimates by the firms that sell them. Nothing here describes anyone's reasons for anything.
Immutable, dated, downloadable
One snapshot per day in JSON, byte identical on re-fetch, served at a dated URL that never changes. A citation pins to a snapshot, not to the live page. The taxonomy is versioned separately from the data (taxonomy v1, data seed-2026-09-12), so a change to a rule never rewrites history. Current: seed-2026-09-12.json. Live state: state.json.
No personal data
The site sets one cookie, a random identifier used to give you a pseudonym for notes. It holds no email, no name, no IP address, and no analytics profile. Right of reply contact addresses go to the editor's spreadsheet and never to this site's database. Requests to remove personal data from a note or a reply go to editor@paciva.ai and are handled within 72 hours.