Web News Scraping (stdlib-only toolkit)

Status: Active Tags: Tooling NLP Agentic Created: 2026-08-30 Last updated: 2026-09-14 Related: hermes-cron-operations, cloudflare-pages-deploy

Definition

A set of five dependency-free Python scripts (stdlib only: urllib, re, json, xml.etree) for researching fresh web news without any search API key. Lives in /home/romain/workspace/. Built 2026-08-30 by the Daily News cron session as an API-key-free fallback when the hosted web_search tool path is slow or unavailable.

How It Works

ScriptInputOutput
bing_search.pyquery(s)Bing HTML web search scrape; freshness via qft=interval%3d%227%22 (24 h) / %228%22 (7 d); parses <li class="b_algo"> blocks → (title, url, snippet, age)
bing_news.pyquery(s)Bing News vertical scrape; parses news-card divs (triple-regex fallback for attribute order) → (title, source, url, snippet, age)
ld_headlines.pyURL(s)fetches page HTML, extracts JSON-LD nodes of type Article/NewsArticle/ReportageNewsArticle/WebPage/BlogPosting → deduped (headline, date, url)
ld_local.pylocal HTML file(s)same JSON-LD extraction offline (for saved/archived pages)
rss_scan.pyRSS 2.0 / RDF (DW) / Atom file(s)parses all three formats, multi-format date parsing, prints items newer than N hours, newest first
rss_probe.py, rss_batch*.pylive RSS URLsbatched multi-outlet RSS probing (added 2026-09-14 by the news cron) — probe outlet feeds for freshness, then batch-dump the fresh ones to disk for the agent to summarize

Common patterns:

  • Browser User-Agent spoofing (Chrome 126 string) on every request.
  • Regex-based HTML parsing, no requests/bs4 — each parse is wrapped so failures are skipped, not fatal.
  • Relative age strings (“N hours ago”) are captured as-is; RSS dates are normalized to UTC before the freshness cutoff.
  • CLI-first: each script is runnable directly (python3 bing_search.py 1 "query"), which makes them usable from agent terminal calls without import plumbing.

Key Parameters

ItemValue
Location/home/romain/workspace/{bing_search,bing_news,ld_headlines,ld_local,rss_scan}.py
DependenciesPython stdlib only (no venv needed)
Freshness filter (Bing)qft=interval="7" = 24 h, interval="8" = 7 days
RSS cutoffrss_scan.py <hours> <files...>

When To Use

  • Research step of news/deals cron jobs when web_search is too slow, rate-limited, or the job must avoid hosted-tool latency (see hermes-cron-operations — the news job has died twice on API timeouts).
  • Verifying a story’s real outlet coverage via JSON-LD of a known article page.
  • Bulk freshness check of a directory of saved feeds with rss_scan.py.

RSS-first research mode (as of 2026-09-14)

  • The Daily News cron’s research step is now RSS-first: Google/Bing/Reuters/NYT are blocked by the sandbox network, so the 09-14 run built rss_probe.py + rss_batch*.py and pulled all 30 cards from live RSS of these 19 outlets: France24, The Hill, UN News, TASS, RT, Global News (Canada), NPR, TechCrunch, Ars Technica, Wired, CNBC, MarketWatch, Climate Home, Inside Climate News, Electrek, Phys.org, ScienceDaily, Variety, The Hollywood Reporter.
  • Pattern: probe each outlet’s feed for items fresh enough, batch-dump the fresh items to disk (spill research to files — keep context lean per the cron hardening rules), then have the agent summarize per category.
  • This supersedes the earlier curl-against-hub-pages fallback (09-03) as the primary research path; web_search remains unavailable in cron sandboxes.

Risks & Pitfalls

  • Regex HTML parsing is fragile: Bing class names (b_algo, news-card) change without notice — if a script suddenly returns 0 items, assume markup drift, not empty results.
  • Headline/snippet only — no full-article extraction; use ld_headlines.py on the article URL for structured metadata, but body text still needs a fetch+strip pass.
  • rss_scan.py strips <!DOCTYPE> before ElementTree parsing (required for feeds that carry it); other strict-XML quirks will surface as PARSE FAIL lines.
  • Heavy use of the same UA may get IP-throttled by Bing; keep request counts modest (the cron pattern is a handful of queries per run).
  • ld_headlines.py includes WebPage type — expect some noise beyond real articles; ld_local.py deliberately omits WebPage.

Sources

  • /home/romain/workspace/bing_search.py, bing_news.py, ld_headlines.py, ld_local.py, rss_scan.py (created 2026-08-30 04:14–04:25 by cron session cron_eef1a69519af_20260830_041056)
  • Cron job eef1a69519af (Daily News Refresh + Deploy) prompt + failure report cron/output/eef1a69519af/2026-08-30_04-41-46.md