Methodology & Sources

Last updated: September 2026

Why this page exists

newscrate aggregates content from other organizations rather than reporting it directly. Every figure and event you see here should be traceable back to where it came from — this page is that trace. If something looks wrong, this is also the first place to check whether it's a source issue or a bug on our end (and the contact form reaches a real person either way).

News articles

Pulled from each publisher's own public RSS/Atom feed on a 5-minute cycle, deduplicated by article URL. We store the title, a short excerpt from the feed's own summary, and a link back to the original — never the full article text, and we never rehost or republish anyone's reporting. Currently active feeds:

SourceCategory
BBC NewsWorld
NPRWorld
The GuardianWorld
DWWorld
France 24World
The Wall Street JournalWorld
The TelegraphWorld
Fox NewsWorld
Al JazeeraGeopolitical
Foreign PolicyGeopolitical
The VergeTech
Ars TechnicaTech
TechCrunchTech
NASAScience
Nature NewsScience
The GuardianBusiness
CNBCBusiness

A feed going offline or changing its URL only reduces coverage from that one source — it doesn't affect ingestion of the others. This list will grow over time.

The "Category" column above is each feed's default, not a fixed label — every article's title and summary is checked against a keyword list per category (tech, science, business, geopolitical), and reclassified when one clearly wins, so a business story from a general World feed lands in Business rather than staying tagged World just because of which feed it arrived on. A feed's default is only kept when no category's keywords clearly apply.

Alongside these, GDELT's DOC 2.0 API — a real-time index of a much larger set of global news sources — is queried for articles matching a geopolitical-crisis keyword filter (coups, insurgency, strikes, ceasefires, sanctions, and similar) and stored under the same "Geopolitical" category as Al Jazeera and Foreign Policy above, since it's a broader net than our hand-picked feed list.

Every incoming article's title and summary is checked against a static list of country names and common aliases; a match tags the article with that country and plots it at that country's rough centroid on the Live Map's News layer, linking straight back to the source — a free, keyword-based, country-level signal.

Source bias labels

Articles from our own curated feed list above carry a small colored label — Left, Lean Left, Center, Lean Right, or Right — next to the source name, plus a "source balance" bar showing the mix of labels among the articles currently on screen. This is a hand-set lookup based on general, widely-cited public consensus about each outlet's overall editorial leaning, not licensed data from a commercial ratings service (AllSides, Ad Fontes, etc.) — we don't have access to and can't redistribute those proprietary ratings, so treat the label as a rough orientation, not a precise score. Trade/science publications (The Verge, Ars Technica, TechCrunch, NASA, Nature News) aren't politically oriented in the way a general news outlet is, so they're left unlabeled rather than forced onto the scale. GDELT's much broader, algorithmically-selected set of sources is also left unlabeled — it spans far too many outlets to rate individually, and a fabricated rating would be worse than none. Country editions' own local sources (see "Country editions" below) are hand-picked the same deliberate way as the sources above, so they're marked "curated" rather than "aggregated" - but not yet bias-labeled, since we don't have researched editorial-leaning data for them yet.

When a story is corroborated by two or more labeled sources (see "Duplicate coverage" below), a small per-story bias bar appears next to it, showing the lean split across everyone covering that specific story - not the whole page's aggregate above. If every one of those labeled sources leans the same direction with none from the opposing side, it's tagged "One-Sided": a real, honest signal, built entirely from our own hand-set labels and duplicate-detection above, that only one side's framing of that story is represented in our own curated list - worth knowing regardless of which side it is.

Duplicate coverage

When multiple outlets cover the same story - our own curated feeds, GDELT's broader net, or a mix of both - only the first one seen shows in the main list, with an "Also covered by N other sources" note rather than 2-3 near-identical headlines back to back. Detection uses a text embedding (Cloudflare Workers AI) compared against other articles from the last 48 hours; a high similarity score marks the later one as a duplicate of the first. GDELT's much larger volume is capped to a fixed number of its newest articles per 5-minute cycle rather than embedding every one it fetches, to keep this predictable and cheap - anything beyond the cap simply isn't checked that cycle, the same trade-off already made elsewhere for GDELT's own record limit. This is a similarity threshold, not perfect judgment - it's deliberately set conservatively (biased toward missing a real duplicate over wrongly merging two distinct stories), so occasionally the same story may still appear more than once.

Within each day on the homepage (day boundaries always stay chronological, so nothing old resurfaces above today's news), articles are ranked by a small weighted score rather than pure publish order or a flat "most-covered wins" rule - modeled on two long-established, public ranking algorithms rather than invented from scratch: Hacker News' time-decay shape (score divided by (age in hours + 2)1.8, the same gravity constant HN itself uses, so a boost fades predictably rather than pinning one story to the top all day) combined with Reddit's log-compressed corroboration count (log₁₀(1 + other-source count), so going from 1 to 2 corroborating sources matters far more than going from 8 to 9). A category weight is layered on top, favoring geopolitical and world coverage over tech/business/science in this mixed view specifically - picking a single category filter shows the same articles in that category regardless, since the weight is identical for every item once only one category is on screen.

Confidence tiers

An article from GDELT's broader net (rather than our own hand-picked feed list above) carries a small "Aggregated" label next to its source name - a human chose every feed on our own list, while GDELT selects sources algorithmically from a much larger pool. Neither is hidden or excluded; the label just says which selection method produced that particular article, so you can weigh it accordingly. Curated-feed articles carry no label at all, on the theory that the default case doesn't need flagging - only the exception does.

Conflict & unrest events (Live Map)

Event data on the Live Map page comes from UCDP (the Uppsala Conflict Data Program)'s Georeferenced Event Dataset, a free, no-key academic dataset of political violence, geolocated to a specific place and date. It classifies events into State-based violence, Non-state violence, and One-sided violence, released on a periodic cadence.

Marker size scales with reported fatalities. Event data on this site is used under UCDP's own terms of use and attribution policy; we are not affiliated with or endorsed by them.

Hazards (earthquakes, cyclones, floods, droughts, volcanoes, wildfires, tsunamis)

Earthquakes come from the USGS public GeoJSON feed, filtered to magnitude 4.5 and above so a global feed isn't dominated by the thousands of minor tremors that occur daily. Other hazard types come from GDACS (Global Disaster Alert and Coordination System), which assigns each event a categorical alert level (Green/Orange/Red) rather than a magnitude. Wildfires additionally draw on NASA FIRMS satellite fire detections, capped to the highest-confidence, highest-intensity detections since the raw feed runs to thousands of points a day globally. All shown on the Live Map, filterable by hazard type and independently of the conflict-event layer. The Live Map also offers a globe projection as an alternative to the default flat map (lazy-loaded only if you switch to it, so it doesn't slow down the default view).

Flight emergencies

From adsb.lol, a free, community-fed ADS-B aggregator: aircraft actively broadcasting one of the three recognized emergency squawk codes (7500 hijack, 7600 radio failure, 7700 general emergency), updated in near real time.

Internet outages

From Cloudflare Radar's outage annotations, plotted at the affected country's centroid (Radar reports by country, not exact location). A national- or regional-scale internet disruption is often a useful crisis signal in its own right — deliberate shutdowns, infrastructure damage, or unrelated technical failures all show up the same way here, so treat this as "something happened to connectivity," not a claim about the cause.

GDELT-flagged events

Beyond the keyword-filtered article search described above, we also pull GDELT's raw Global Knowledge Graph (GKG) export directly — a free, no-key file published every 15 minutes with real per-article geolocation. Records are kept only if they're tagged with themes suggestive of conflict or unrest, then capped to the most negative-toned matches per cycle: a broad automated signal for "something newsworthy and tense happened here."

Reference infrastructure (nuclear plants, military bases)

Static reference layers, not live feeds — nuclear power plant locations from GeoNuclearData (sourced from WNA/IAEA) and military base locations from HIFLD, the U.S. government's Homeland Infrastructure Foundation-Level Data. Fetched once rather than every cycle, since this kind of infrastructure doesn't change day to day.

Prediction markets

Live odds from two sources, kept on the same footing rather than picking one as primary: Polymarket's public Gamma API and Kalshi's public market-data API (a CFTC-regulated exchange). Both are filtered to markets whose titles look geopolitically relevant — both platforms cover everything from sports to entertainment, and this filter keeps the Monitor page focused. A market's probability is the current price traders are willing to pay, not a forecast newscrate is making itself.

A multi-candidate race (a presidential election, say) isn't one market — both platforms structure it as a separate Yes/No contract per candidate. Those are grouped back together into one ranked card showing each candidate's odds, rather than scattering a single race across several unrelated-looking rows in the list.

Cyber advisories

From CISA's Known Exploited Vulnerabilities catalog — vulnerabilities with confirmed evidence of active exploitation, including state-linked campaigns. Treated as a cyber-conflict signal alongside the physical-world data above.

Trending entities & sentiment

From APITube's News API, which attaches named-entity extraction (organizations, people, places) and sentiment scoring to each article it indexes. Refreshed periodically rather than every cycle, and shows the organizations, people, and places mentioned most often across a recent sample of war/conflict/sanctions coverage specifically (not a general trending-across-all-news measure — the "Trending" row on the homepage and Monitor's own listing say so on hover), alongside the average sentiment of the coverage mentioning them — a tone reading on the reporting itself, not a judgment newscrate is making about the entity.

Elevated badges

FX rates, prediction-market prices, and trending-entity mention counts on Monitor are snapshotted every 30 minutes into a rolling history; once a currency, market, or entity has at least 20 snapshots, an "Elevated" badge appears when its current value sits 2 or more standard deviations from its own 30-day average - a plain statistical outlier flag, not a prediction or a judgment about whether the move is good or bad. Nothing shown before 20 snapshots have accumulated (roughly half a day), so a badge is never based on too little history to mean anything. Prediction markets also show their 24-hour probability-point swing directly, flagging moves of 15 points or more.

A market moving 15+ points in 24h is checked against newscrate's own articles for anything matching its topic in the last 48 hours; if nothing turns up, it's tagged "Divergence" - traders are pricing in a real move on something this site hasn't picked up in the news yet, whether that's because it genuinely hasn't been reported, or just because the check didn't match the market's exact wording. This runs on the same 30-minute cron cycle as everything else on this page, not live when someone happens to open Monitor, so it's checked (and cleared again, once a move fades or coverage catches up) on schedule regardless of traffic.

Two severity tiers, not one - "Elevated" (outlined) for 2-3.5 standard deviations, "Highly Elevated" (solid) at 3.5+, the same escalation idea as GDACS's own Green/Orange/Red hazard-alert levels used elsewhere on this site. A 2-stddev move and a 5-stddev move aren't equally unusual, and one flat label was throwing that difference away.

Country pages use the same statistical approach for conflict-event activity: a country's last-24h event count is compared against its own 30-day baseline, and a status line appears only when it's a genuine outlier (2+ standard deviations, with enough history to mean something). One honest caveat - the baseline is built only from windows where at least one event occurred, not from every 30-minute snapshot including the quiet ones, so it reads as "unusually active even by this country's already-active periods" rather than "active vs. total silence."

Natural-language Archive search

Archive's "describe what you're looking for" box (below the plain keyword search) sends your request to Gemini, which extracts a short keyword query plus an optional category/country - the actual search still runs through the same full-text search as everything else, Gemini never sees or ranks the articles themselves. What it interpreted your request as is always shown alongside the results, so you can see exactly what got searched for rather than a black box. Unlike situation summaries, each request is a one-off, deliberate action and isn't cached.

AI-generated situation summaries

Country pages that show unusually elevated conflict-event activity (see "Elevated badges" above) may include a short AI-generated summary, via Google's Gemini API. The model is given only a small structured object - counts of events, curated vs. aggregated articles, and nearby infrastructure for that country - and asked to describe what those counts show in plain language; it isn't given article text, isn't asked to predict anything, and isn't told to explain causes. Every summary is labeled "AI-generated summary" and paired with the confidence note and caveats the model itself produced about its own inputs. A country with nothing notable, or a failed/unconfigured AI call, simply shows no summary - the structured counts elsewhere on the page are the complete, correct picture either way; the summary is a convenience on top, never a replacement.

Generated on the same 30-minute cron cycle as everything else on this page, for any country with conflict-event activity in the last day - so a summary is usually already current by the time you visit, not generated fresh by whichever visitor happens to be first after something changes. A country notable only for being unusually quiet falls outside that sweep's candidate list and still generates on first visit instead.

Top Coverage (homepage)

The card at the top of the homepage is deliberately not AI-written - two earlier attempts at having Gemini summarize an arbitrary cross-category sample of headlines into one "headline" either overstated how important something was, or, once told to hedge, read as vague filler. The actual useful signal - which story multiple independent outlets happen to be covering right now, and what's recent in each category - is real, deterministic data newscrate already has, so the card just selects and organizes it rather than paraphrasing it. A "Top story" only appears when it's tagged World or Geopolitical AND at least one other source is independently covering the same thing (the same detection used for "Also covered by" elsewhere) - restricted to those two categories on purpose, not just down-weighted: tech-blog coverage of the same product announcement reliably clears our similarity threshold since outlets covering the same press event tend to phrase it near-identically, while genuinely serious stories covered by differently-phrased outlets often don't, so a discount alone couldn't stop a well-corroborated tech story from winning by default on a quiet day for real news. A quiet cycle with nothing corroborated in those two categories simply shows no top story rather than picking outside them. Below it, a few of the most recent headlines in each category are listed as-is, using the original outlets' own titles - nothing rewritten, nothing summarized, no hallucination risk.

Breaking news ticker

The scrolling red bar shown site-wide (pause it by hovering) flags an article as breaking on recency alone - published in the last 45 minutes by one of our own curated sources (not GDELT's much larger, algorithmically-selected net). A first version of this required several independent sources to already be corroborating a story, but that measures "confirmed," not "breaking" - real breaking news starts with one source, by definition. The trade-off is an honest one: a single hand-picked outlet's initial report can still turn out wrong or early, the same as any breaking banner anywhere else. If other curated sources do pick up the same story afterward, that shows as a small "+N sources" tag rather than being required up front. A quiet period with nothing that fresh means the ticker is simply absent, not filled with an arbitrary headline to avoid looking empty. It refreshes every 5 minutes, matching the site's fastest ingest cycle.

Saved articles

The ☆ on any article saves it to a "Saved" list, viewable via the toggle next to the category filter. This is stored only in your own browser's local storage - there are no accounts, nothing is sent to newscrate, and it won't follow you to another device or browser.

Submitting an article

Found something from a source not on our list? The "Found an article from a source we don't have yet?" box on the Archive page takes a URL, fetches that one page, and screens it with Gemini using only its title and meta description - never full content, and never anything beyond that one page. A submission that reads as spam, a non-article page, or gibberish is quietly received and goes no further; nothing is published automatically either way. One that passes emails a review link to the site's maintainer, who chooses one of four outcomes: publish this one article, add the source going forward without publishing this particular article, both, or reject - a real decision by a person, every time, never an automatic publish. An ongoing source approved this way is fetched on the same 5-minute cycle as every hand-picked feed above, going forward.

Country editions

newscrate.cc's own article list is one global, hand-picked set of sources. Separate hostnames (starting with Malta at mt.newscrate.cc) run the exact same site - the same ranking, Top Coverage, breaking ticker, Archive search, and RSS feeds - against a different, smaller list of that country's own local sources instead, so a Malta headline getting far less attention on the global site because it's a small country isn't competing against BBC/Reuters volume on its own edition. One shared codebase and database, not a separate deployment per country: every article is tagged with which edition it belongs to, and everything that reads from the article table (the list itself, duplicate/corroboration detection, Top Coverage, the breaking ticker) is scoped to match. The Monitor dashboard (prediction markets, currency stress, cyber advisories, trending entities) is deliberately not split by edition - those signals aren't meaningfully "local" the way news coverage is, so every edition shares the same global Monitor. A small "[Country] edition" tag next to the logo, linking back to the global site, says when you're on one.

Currency stress

Daily ECB reference rates via Frankfurter, covering every currency it tracks (~30) rather than a fixed curated subset, with a search box to find any of them by code or name, and a choice of which currency to measure everything else against. A proxy signal, not a claim of causation — currencies move for many reasons.

RSS feeds

/feed.xml is a personal RSS feed of newscrate's own articles, filterable to a category and/or a source (e.g. /feed.xml?category=geopolitical), so you can subscribe to a slice of the site in your own reader instead of checking back. Archive's filter panel and every country page link to the feed for whatever you're currently viewing.

Source status

The Status page shows the live up/down state of every ingest source, checked on each cron run and overwritten each time (not a historical uptime log). If something on the site looks stale or missing, that's the first place to check whether it's a known source issue rather than a bug.

Refresh cadence & limits

Everything on this page refreshes on a 5-minute cycle via a scheduled job — there's no manual curation deciding what shows up first. We cap how much we pull per cycle to stay within reasonable limits for a small, free site.

Public API

Everything newscrate shows is also available as JSON, read-only, no key or account needed. This is the same fused data the site itself renders — not a raw passthrough of the underlying sources, which you can already reach directly (they're all linked above). Be a reasonable neighbor: this runs on a small free-tier setup, so cache what you can and don't hammer it.

EndpointWhat it returns
/api/articlesNews articles. Filters: category, country, source, from/to (date), hours, limit/offset.
/api/searchFull-text article search. q plus the same filters as above.
/api/eventsConflict/unrest events. Filters: country, type, hours, limit.
/api/hazardsEarthquakes, hazards, flight emergencies, outages, GDELT-flagged events.
/api/marketsPrediction markets (Polymarket, Kalshi).
/api/sitesReference infrastructure (nuclear plants, military bases).
/api/entitiesTrending entities and sentiment.
/api/briefingThe homepage Top Coverage card's data (top story, headlines by category).
/api/breakingCurrent breaking-news ticker items.
/api/situation?country=A country's fused situation object, plus an AI summary when notable (see above).
/api/history?type=&days=Trend history for fx/market/entity/cyber/country_events; add stats=1 for the mean/stddev baseline behind the Elevated badges.
/api/statusSame data behind the Status page.
/api/cyber, /api/fxCyber advisories, currency rates.
/feed.xmlRSS, not JSON - filterable the same way as /api/articles.

Questions or a source to suggest

Use the contact form — "Suggest a source" is one of the options there for exactly this.