Hermes Agent Cron Operations (this host)

Status: Active Tags: Tooling Infrastructure Agentic Created: 2026-08-30 Last updated: 2026-10-01 Related: web-news-scraping-stdlib, cloudflare-pages-deploy, local-llm-stack

Definition

Operational patterns, limits, and observed failure modes for scheduled Hermes Agent cron jobs running on this host (profile romain). The recurring theme: agent-driven research jobs are bounded by several hard timeouts, and jobs must be written so they never sit idle.

How It Works

Hard limits (observed failures)

LimitSymptomObserved incident
Cron idle timeout: 600 sTimeoutError: Cron job '...' idle for 600s (limit 600s) — job killed if no tool activity for 600 sDaily News job, 2026-08-30 04:41 (after building the scraping toolkit, session went idle); 2026-09-04 04:42 (reported as “idle for 603 s”; last activity a completed patch call — the run had made substantial progress and left partial artifacts in the workspace)
Non-streaming API call: 240 sRuntimeError: Non-streaming API call timed out after 240s with no responseDaily News job, 2026-08-22 15:07
Stale-token floor (provider)Long agent runs killed when the provider’s stale check (default 180 s on the qwen3 slug) fires mid-runFixed 2026-08-30: providers.custom.stale_timeout_seconds=900 in config — note the CLI splits dotted model names, so the key is provider-wide, not per-model
LLM API connection errorRuntimeError: Connection error. — API connection fails before the agent does any work; output .md contains only prompt + error (~4 KB), no partial stateDaily News job, 3 consecutive days: 2026-08-31, 09-01, 09-02 (05:41)
web_search loop guardrailloop_web_search_cap fires after ~50 non-progressing web_search calls — agent must change strategy (direct curl / local toolkit), not retry the same queryDaily News re-run 2026-08-30 11:11 — agent then fell back to curl, wrote content, built and deployed successfully
Response length limitRuntimeError: Response truncated due to output length limit — the run dies when a model response hits an output length cap; the output .md contains only prompt + error (same shape as a connection error), no partial report. First observed; root cause not yet pinned (long final report vs. provider cap — both plausible)Daily News fire-now re-run 2026-09-05 18:26 (the same morning’s scheduled run had died on idle 601 s); site unaffected
Cron sandbox: no network / no web_search toolThe cron sandbox lost all external network access (all interfaces NO-CARRIER, DNS failure) and the hermes_tools Python module is not installed in the cron venv. The agent cannot research news or deploy to Cloudflare Pages. Observed on 2026-09-25 — a fundamental environment regression; likely triggered by a recent change (Python 3.14 upgrade, Hermes Agent version update, or container network config).
Generic request timeoutRuntimeError: Request timed out. — a generic HTTP-level request timeout (not the 600 s idle limit, not the 240 s API timeout). The run produced no partial state. First observed 2026-09-28 — the error string is too terse to classify further; may be a network-level timeout or a provider-side timeout.

Job-authoring best practices

  • Self-contained prompts: absolute paths, explicit workdir, exact file schemas, exact verify steps. Cron sessions have no conversation history.
  • no_agent for deterministic work — don’t pay LLM latency for fixed pipelines.
  • Pin model + provider at hermes cron create — the cronjob tool cannot set them later.
  • [SILENT] in the prompt for monitor-style jobs with nothing to report.
  • Keep total research steps few: 600 s idle + 240 s API timeouts mean a 6-category × N-query research job is near the ceiling; the 2026-08-30 news run built a local fallback (web-news-scraping-stdlib) precisely to cut hosted-search latency.
  • Job output + failure reports land in ~/.hermes/profiles/<p>/cron/output/<job-id>/ — check the latest .md there when a job “fails” to see the exact error.

Prompt-hardening patterns (validated 2026-09-03/04)

Both the Daily News (eef1a69519af) and Weekly Deals (7f0264357282) prompts were hardened after repeated timeout deaths, and verification (fire-now) runs then completed end-to-end: news 2026-09-04 01:56 (deployed, verified 200/30 cards first try) and deals 2026-09-03 23:35 (+9 deals, deployed, HTTP 200). The patterns that made the difference:

  1. Pre-flight validation step (STEP 0) before any research — e.g. import-check the content files and auto-repair (fix_quotes.py) if broken, so a killed prior run’s half-written state is healed before new work starts.
  2. Atomic file writes — write to <file>.new, validate (ast.parse / JSON parse), only then mv over the original. A run killed mid-write can no longer leave a broken file behind.
  3. Context-bloat rules — research ONE unit (category/angle) at a time and write results to disk immediately; cap search queries (~15–20); spill raw research notes to a scratch file (/tmp/deals_research.md) instead of holding search dumps in context; keep assistant turns short.
  4. Explicit tool fallbacks — state the goal, not a required tool (curl when web_search is absent).

Evidence: the same jobs had died on 240 s API timeouts (context bloat, deals 08-29) and 600 s idle (news 08-30) before hardening; after it, both verification runs finished. Caveat: hardening reduced but did not eliminate failures — the 09-04 04:42 news run still died on the idle limit (see below).

Gateway / service ops

  • The gateway rewrites its systemd unit from a template on every start — direct edits to the unit file silently revert. Persist changes with drop-in files: <unit>.service.d/*.conf.
  • Active log files (agent.log, gateway.log, errors.log, weixin-relay.log) must never be deleted; old *.log* get gzipped by the daily maintenance job.
  • Profile topology: default + jiayi profiles are dormant (no gateway/creds); all WeChat/DingTalk credentials exist only in the romain profile (WEIXIN_* in its .env; weixin adapter in the gateway, account a8a52773).

Key Parameters

ItemValue
Cron idle limit600 s
Non-streaming API timeout240 s
Stale timeout (fixed)providers.custom.stale_timeout_seconds=900
Output dir~/.hermes/profiles/romain/cron/output/<job-id>/
Jobs registry~/.hermes/profiles/romain/cron/jobs.json

Current jobs (snapshot 2026-09-14)

IDNameSchedule (CST)
eef1a69519afDaily News Refresh + Deploydaily 04:10
7f0264357282Weekly Deep Deals Research + DeploySat 04:30
4f90ecb2a611Weekly Memory CompactionSun 03:00
c6def452f00fDaily BESS News Feed Refreshdaily 05:15
2cd6faaa0f8aDaily Wiki + Workspace Maintenancedaily 05:00
efcedb24cce0wiki-index-to-honchoevery 360 min
532b734cea6foutlook-email-pollevery 360 min

When To Use

  • Authoring or debugging any cron job on this host.
  • When a cron job fails with a timeout: identify which of the three limits above fired (the error string in the output .md tells you).
  • When changing gateway service config — use drop-ins, never the unit file.

Risks & Pitfalls

  • A job that does one very long terminal command (e.g. a 10+ min deploy) is fine (tool is active) — the idle limit counts time between tool completions; what kills jobs is the model stalling between tools or a hung non-streaming call.
  • The stale-timeout fix is provider-wide: raising it affects every model served by providers.custom (see local-llm-stack).
  • The jobs table above is a snapshot — re-check jobs.json before relying on IDs.
  • A connection-error failure is the cleanest kind: nothing was executed, so a plain re-run is safe (no half-written files to reconcile). Contrast with an idle timeout, which can leave content files half-written.
  • An idle kill can land mid-write even on a run that was progressing (2026-09-04 04:42 news run: killed “idle 603 s” after a patch call, leaving rewritten content files, a new news_assemble.py, and partial per-card JSONs in news_cards/). The live site was unaffected (previous deploy still up), but before re-running or resuming, check the workspace’s on-disk state — atomic writes (.new → validate → mv) make the broken-file case impossible, though partial new artifacts can still accumulate.
  • The maintenance job itself is not immune to the idle limit. It died on idle timeouts two days running — 2026-09-04 06:08 (“idle 600 s”) and 2026-09-05 06:02 (“idle 603 s”, last activity a completed patch) — each time leaving wiki work on disk without the index/log/commit steps (09-05: a whole new page realtime-voice-assistant.md + three user-layer summary edits, uncommitted, missing from both index.md files and both log.md files). Recovery recipe, applied 2026-09-06: ① git -C /media/data/wiki-common status for uncommitted/untracked pages; ② compare each layer’s index.md against the filesystem (find wiki -name '*.md' vs index entries); ③ read the interrupted run’s cron/output/<id>/<latest>.md for what it intended; ④ append the missing index entries + log entries (log LAST, as always); ⑤ commit. Keep wiki-editing turns short (one page per tool call, no long synthesis passes between calls) so the run never sits idle >600 s mid-batch. | web_search is not always present in cron sandboxes. Two distinct failure modes: (1) 2026-09-03 — no web_search binary but network worked; agent succeeded via curl (caching ~45 articles to /tmp/news/). (2) 2026-09-25 — complete environment isolation: no hermes_tools module and no network at all (all interfaces NO-CARRIER, DNS fails). The sandbox cannot research news or deploy to Cloudflare Pages. Root cause: likely a recent environment change (Python 3.14 upgrade, Hermes Agent version update, or container network config). Prompts should not hard-require a specific search tool — state the goal and allow fallbacks (web-news-scraping-stdlib).
  • Job scripts can rot independently of the cron runtime — not every repeated cron failure is a timeout. outlook-email-poll (532b734cea6f) has failed every run since 2026-09-11: first FileNotFoundError: .outlook_credentials.json, later ImportError: cannot import name 'TENANT_ID' from 'outlook_graph' — email_poller.py imports a constant the module never defined (code drift, exits immediately, exit 1). Diagnose by reading the newest run report in cron/output/<id>/; the error string tells you whether it’s a runtime limit vs. a script bug. (Full state in users/romain wiki: summaries/outlook-email-pipeline.md.)
  • wiki-index-to-honcho (efcedb24cce0) has a silent skip bug: the Honcho API caps message content at 25,000 chars and the index script does not chunk, so pengcheng-fashion-exhibition.md (109 KB) has never been indexed (HTTP 422 string_too_long). Worse, the skip logic keys on mtime only — a file whose last attempt errored is skipped forever until its mtime changes. Fix (not yet applied; script lives under the default profile’s ~/.hermes/cron/scripts/ so the romain agent’s cross-profile guard blocks edits): chunk long files into ≤ ~24,000-char messages and change the skip condition to status == "indexed" AND mtime matches.

Sources

  • cron/output/eef1a69519af/2026-08-30_04-41-46.md (idle timeout) and 2026-08-22_15-07-24.md (API timeout)
  • cron/output/eef1a69519af/2026-09-04_01-56-52.md (hardened-prompt success) and 2026-09-04_04-42-13.md (idle 603 s, last activity a patch call)
  • cron/output/7f0264357282/2026-08-29_18-00-59.md (240 s timeout, context bloat) and 2026-09-03_23-35-08.md (hardened-prompt success, +9 deals)
  • cron/output/eef1a69519af/2026-09-05_18-26-21.md (response-truncation error) and 2026-09-06_05-10-04.md (scheduled success after 09-05 double failure)
  • cron/output/2cd6faaa0f8a/2026-09-04_06-08-05.md and 2026-09-05_06-02-52.md (maintenance job idle kills, partial wiki state left behind)
  • cron/jobs.json, profile memories/MEMORY.md + USER.md (compacted 2026-08-30)
  • logs/agent.log cron scheduler entries, 2026-08-30 and 2026-09-04 (run tags cron_<job>_<ts>, per-call latency + tool activity)