{"data":{"items":[{"id":"b7013511-2bc9-4b34-94cc-06ca3f657c7c","excerpt":"Models, Providers & Plans Megathread — June 2026 — **LAST UPDATED:** June 25, 2026\n**Sourced from:** 31+ r/hermesagent threads, 290+ community comments, [May 2026 Models Megathread](https://www.reddit.com/r/hermesagent/comments/1tgbsuz/) by u/digitalnomadpdx\n\n**Scope:** Cloud APIs, subscription plans, provider comparis","url":"https://www.reddit.com/r/hermesagent/comments/1ufrtsf/models_providers_plans_megathread_june_2026/","role":"pricing","weight":1.2431388,"occurredAt":"2026-06-26T00:45:39.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"hermesagent","intent":"pricing_complaint","painScore":0.25513104,"sentiment":0.5473251,"confidence":0.99044544,"matchedPatterns":["recommend","how_can_i","too_expensive","free_tier","missing_feature","workaround"],"statement":"Too expensive as primary.","title":"Models, Providers & Plans Megathread — June 2026","body":"**LAST UPDATED:** June 25, 2026\n**Sourced from:** 31+ r/hermesagent threads, 290+ community comments, [May 2026 Models Megathread](https://www.reddit.com/r/hermesagent/comments/1tgbsuz/) by u/digitalnomadpdx\n\n**Scope:** Cloud APIs, subscription plans, provider comparison, model selection, and routing strategies. For local/self-hosted models, see the [Mac/MLX Megathread](https://www.reddit.com/r/hermesagent/comments/1uc7rw5/) and [r/hermesagent Local Models Guide](https://www.reddit.com/r/hermesagent/comments/1stiwug/).\n\n---\n\n## TL;DR — What Should I Use?\n\n| Decision | Community Pick | Runner-Up | Notes |\n|----------|---------------|-----------|-------|\n| **Best overall (paid API)** | DeepSeek v4 Pro (direct) | DeepSeek v4 Flash | Pro for orchestrator, Flash for workers/auxiliary. $60 got one user 8B tokens. |\n| **Best value subscription** | OpenCode Go ($10/mo) | Minimax $10 token plan | OpenCode Go = \"~$60 API credit for $10.\" Minimax = \"virtually unlimited\" background agent work. |\n| **Best premium subscription** | Nous Portal ($20) | OpenAI Codex ($20 ChatGPT) | Fixed monthly cost, no billing surprises |\n| **Best free model (no catch)** | owL-alpha (OpenRouter) | Nemotron 3 Super 120B (free) | \"Absolute beast at tool usage.\" Best free tier for Hermes agentic tasks. |\n| **Best coding model** | GPT-5.5 (via Codex) | DeepSeek v4 Pro | GPT-5.5 is the \"undisputed king\" for complex coding. |\n| **Best orchestrator model** | GPT-5.4-mini | DeepSeek v4 Flash | Fast, cheap, handles 90% of routing/web search/light tasks. |\n| **Best provider for predictable billing** | OpenCode Go + Minimax stack | Nous Portal | Subscriptions = no surprise $100 days. Go+Minimax = $20/mo for near-unlimited. |\n\n---\n\n## PART 1: Cloud Provider Comparison\n\n### Tier 1 — Community Favorites\n\n| Provider | Pricing | Model Access | Best For | Watch For |\n|----------|---------|-------------|----------|-----------|\n| **DeepSeek (direct API)** | Flash: ~$0.22/M in, $0.20/M out (cache reads $0.004/M). Pro: ~$0.44/M in, $0.87/M out. Real-world: $0.30–$1.30/day for Flash, $2–6/day for Pro. | v4 Pro, v4 Flash, Coder | Primary orchestrator, heavy coding, cost-sensitive workflows | Direct API 4-5x cheaper than via OpenRouter. Throttling/503 errors during US peak hours (single-digit t/s). |\n| **OpenCode Go** | **$10/mo** ($5 first month). ~$60 worth of credits at list prices. | DeepSeek Flash/Pro, Minimax M3, MiMo 2.5 Pro, GLM, Kimi, and more | **Best overall value subscription.** Covers 90% of agent workload. No concurrency limits. | Lacks some multimodal models (Gemma 4). Weekly/monthly caps. |\n| **Nous Portal** | $20/mo subscription | Hermes models, DeepSeek, Qwen, routing | All-in-one convenience, predictable billing | Some models cost extra credits; check usage dashboard. Using non-free models exhausts $20 quickly. |\n| **OpenAI Codex** | $20/mo (ChatGPT Plus BYOK) | GPT-5.5 (undisputed king for coding), GPT-5.4-mini (ultra-fast orchestrator) | Complex coding, deep reasoning. Best used as \"senior fixer\" — not as daily driver. | Burns weekly rate limits fast if overused. OAuth only. Shares ChatGPT rate limits. |\n\n### Tier 2 — Strong Alternatives\n\n| Provider | Pricing | Model Access | Best For | Watch For |\n|----------|---------|-------------|----------|-----------|\n| **Minimax** | $10/mo token plan (\"virtually unlimited\"). 15K req/week on high-speed models, 1.5K/5hr. | M3 (as good as GPT-5.5-low per community), M2.7 | Background agent work, auxiliary tasks, stable everyday use | M2.7 reliable but uncreative. M3 is better. Prone to looping without guardrails. |\n| **Xiaomi MiMo** | $6/mo direct token plan. ~$13/mo annual (2.4B tokens/yr). | MiMo 2.5, MiMo 2.5 Pro | Strong agentic intelligence, vision, coding. \"Steal at current price.\" | No caching — burns tokens faster than DeepSeek. Sometimes over-eager. |\n| **Kimi/Moonshot** | Pay-per-token via OpenRouter or direct | K2.6 (very intelligent, strong tool calling), K2.7 | Best open-source Hermes main model. \"Built my entire Debian server stack.\" | Strict quota limits. Occasional Chinese chars in output. Tends to overthink on coding. |\n| **OpenRouter** | Pay-per-token (variable). Free tier available with $10 credit. | 200+ models. **owL-alpha (free)** is the best free model — \"absolute beast at tool usage and coding.\" | Model experimentation, fallback chains, free-tier models | Variable pricing; same model costs 4-5x more through resellers. Avoid auto-routing. Pin models to specific providers. |\n| **Gemini (Google)** | Pay-per-token / free tier | Flash 2.5 (free), Pro 2.5 | Free tier for light tasks, strong vision | Rate limits on free tier; OAuth subscription risky as BYOK |\n| **Ollama Cloud** | $20/mo ($22 credits), $100 tier | Free/open models only | Hassle-free hosted local-style models | **3 concurrent connection limit** — cron jobs crash if chatting simultaneously. Recently degraded (server busy, slow tokens). No frontier models. |\n\n### Tier 3 — Budget / Niche\n\n| Provider | Pricing | Best For | Watch For |\n|----------|---------|----------|-----------|\n| **GLM 5.1 / 5.2 (NeuralWatt)** | Free $5 credit, then PAYG. GLM-5.1 is very efficient and cheap. | Deep reasoning when speed doesn't matter. Stable, reliable. | 5.2 pricier. Painfully slow (18hrs for what GPT-5.5 does in 1hr). Prone to looping. |\n| **Anthropic Claude (sub)** | $20/mo subscription | Opus 4.7 — high-quality reasoning, code review | Agentic use explicitly discouraged by Anthropic. Too expensive as primary. Token hog. |\n| **NVIDIA NIM** | Free tier | Nemotron 3 Super 120B (free) — best emergency fallback when GPT limits hit | Smaller ecosystem; genuinely free |\n| **Grok / superGrok** | $10/mo or $30/mo X Premium | Multi-modality, voice, good tool calling | X Premium $30/mo gives only ~2hrs agent work. API gives much more. Weak at coding. |\n| **NanoGPT** | $12/mo | Only if you need uncensored models | \"Sketchy AF.\" Slow, low limits. Models overly verbose. Not recommended as primary. |\n| **Qwen OAuth** | Subscription / pay-per-token | Qwen models direct | Newer provider; fewer community data points |\n| **OpenCode Zen** | $10/mo | Curated model selection | Smaller model selection than Go. Go is the better value. |\n\n---\n\n## PART 2: Model Comparison — Which Model for Which Task\n\nCommunity consensus on which cloud models excel at which role. All accessible via the providers in Part 1.\n\n### 🥇 Tier 1 — Best-in-Class\n\n| Model | Best For | Access Via | Real-World Notes |\n|-------|----------|-----------|-----------------|\n| **GPT-5.5** | Complex coding, deep reasoning, research | OpenAI Codex ($20/mo), OpenAI API | \"Undisputed king\" — from 6B-token test. Writes economic journal articles. Burns rate limits fast. |\n| **DeepSeek v4 Pro** | Daily driver, heavy coding, multi-step synthesis | DeepSeek direct API, OpenCode Go, Nous Portal | \"Smarter than Claude Sonnet.\" $60 for 8B tokens over 2-3 weeks. Best value powerhouse. |\n| **DeepSeek v4 Flash** | Orchestrator, routing, 90% of daily tasks | DeepSeek direct API (free tier available), OpenCode Go | $0.30–$1.30/day. With `reasoning=xhigh` can exceed Pro in some domains. Cache hits at $0.004/M. |\n| **Claude Opus 4.7** | Highest-quality reasoning, code review | Anthropic subscription ($20/mo), OpenRouter | Brilliant but expensive. Token hog. Not for daily driving. Best as CLI-invoked code reviewer. |\n\n### 🥈 Tier 2 — Strong Performers\n\n| Model | Best For | Access Via | Real-World Notes |\n|-------|----------|-----------|-----------------|\n| **GPT-5.4-mini** | Ultra-fast orchestrator, routing | OpenAI Codex, OpenAI API | Handles 90% of routing/web search. Pair with GPT-5.5 for heavy lifting. |\n| **Kimi K2.6** | General daily agent, coding, tool calling | OpenRouter, Kimi direct API | \"Built my entire Debian server stack and moved 5 agents.\" Very intelligent, inexpensive. Strict quotas. |\n| **Minimax M3** | Background agent work, stable everyday use | Minimax $10 token plan, OpenCode Go | \"As good as GPT-5.5-low.\" The $10/mo \"virtually unlimited\" plan eliminates token anxiety. |\n| **MiMo 2.5 Pro** | Agentic intelligence, coding, vision | Xiaomi $6-13/mo plan, OpenCode Go | Strong multi-modality, good coding. \"Steal at current price.\" No caching — burns faster. |\n| **Gemini 3.1 Pro** | Research, multi-modal, strong in custom pipelines | Google API, OpenRouter | Mixed community reception but holds up in custom agent pipelines. |\n\n### 🥉 Tier 3 — Budget & Specialist\n\n| Model | Best For | Access Via | Real-World Notes |\n|-------|----------|-----------|-----------------|\n| **owL-alpha** | Best free model — tool usage and coding | OpenRouter (free tier) | \"Absolute beast at tool usage.\" Rate-limited for heavy multi-step loops. |\n| **Nemotron 3 Super 120B** | Emergency fallback, free coding | NVIDIA NIM (free), OpenRouter (free) | Best fallback when everything else is rate-limited. Free. |\n| **GLM-5.1** | Deep reasoning when speed doesn't matter | NeuralWatt (free $5 credit) | Stable, reliable. Painfully slow (18hr vs 1hr for GPT-5.5). |\n| **Grok 4.3** | Multi-modality, voice, general agent | Grok API, X Premium | API gives much more than $30/mo X Premium sub (only ~2hrs agent work). |\n| **Gemini 2.5 Flash** | Free tier, strong vision, light tasks | Google API (free tier) | Free. Good for light automation. Rate limits on heavier use. |\n| **Qwen 3.6** (API) | Reliable function calling, cheap API rates | OpenRouter, Qwen OAuth | Reliable tool use. Rate limits on free tier. |\n\n---\n\n## PART 3: Subscription vs Pay-Per-Use\n\n### Use a subscription ($10-20/mo fixed) if:\n\n- You want predictable billing (no $100 surprise days)\n- You use Hermes daily for extended sessions\n- You prefer \"set and forget\" without monitoring token burn\n- You're new to Hermes and don't know your usage patterns yet\n\n### Use pay-per-token (DeepSeek direct, OpenRouter) if:\n\n- Your usage is bursty (heavy days, then light days)\n- You're extremely cost-sensitive and willing to monitor usage\n- You run multiple worker profiles on cheaper models\n- You can self-manage fallback chains and routing\n\n### Hybrid strategy (most common in community):\n\n> *\"I use OpenAI $20/mo + DeepSeek v4 Flash for workers. Pro for main orchestrator. Multiple providers so I never hit a single rate limit.\"* — from \"Affordable and good Models\" thread\n\nPattern: Subscription for primary → cheap pay-per-token for workers/auxiliary. This is the most-recommended approach across 5+ threads.\n\n---\n\n## PART 4: Provider Reliability & Gotchas\n\n### DeepSeek\n- ✅ Extremely cheap direct API ($0.22/M input Flash, $0.44/M Pro; cache hits at $0.004/M)\n- ✅ v4 Pro widely considered better than Claude Opus 4.7 for agentic tasks. \"Smarter than Claude Sonnet\" — from 6B-token test.\n- ✅ Direct API includes prompt caching (80-90% cache hit rate with repeated tool schemas — absurdly cheap)\n- ✅ Real-world costs: $0.30–$1.30/day for Flash, $2–6/day for Pro, $60 for 8B tokens over 2-3 weeks (u/drwebb)\n- ❌ OpenRouter markup: same model costs 4-5x more through OR resellers\n- ❌ Pricing changes: v4 pricing restructured May 2026; monitor for future changes\n- ❌ China-based; data sovereignty concerns for some users\n- ❌ Throttling/503 errors during US peak hours (single-digit t/s). Community workaround: fallback chain to Flash or another provider.\n- 💡 **Tip:** Use direct API, not OpenRouter. Pin to `api.deepseek.com`. Cache system prompt + tool schemas for massive savings.\n\n### OpenRouter\n- ✅ Access to 200+ models from one account\n- ✅ Built-in fallback chains\n- ✅ owL-alpha on free tier — \"absolute beast at tool usage\"\n- ❌ Variable per-provider pricing — same model at different price points\n- ❌ Community warns: \"avoid silent auto-routing for anything that mutates state\"\n- ❌ Default routing can silently switch models mid-task\n- 💡 **Tip:** Pin specific model IDs and set limits. Use as fallback pool + free tier only.\n\n### Nous Portal\n- ✅ $20/mo covers multiple models\n- ✅ First-party integration with Hermes (Nous builds Hermes)\n- ❌ Some models consume extra credits beyond subscription\n- ","offTopic":false},{"id":"552fd439-d71d-468f-a445-c59b2a8770c9","excerpt":"Cloud Models & Providers for Hermes Agent — The August 2026 Megathread — **LAST UPDATED: August 23, 2026**\n\n*Compiled from ~30 r/hermesagent threads and hundreds of comments (Aug 1–23), Hacker News discussion threads, live OpenRouter pricing (pulled Aug 23), and independent benchmark data. Every price is the live API r","url":"https://www.reddit.com/r/hermesagent/comments/1vvv2x1/cloud_models_providers_for_hermes_agent_the/","role":"pricing","weight":1.0582464,"occurredAt":"2026-08-23T02:31:50.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"hermesagent","intent":"pricing_complaint","painScore":0.27801403,"sentiment":0.21875,"confidence":0.8280397,"matchedPatterns":["free_tier","workaround","product:anthropic"],"statement":"**Q: Can I use my Claude Max subscription inside Hermes?** **A:** Not directly — Anthropic subscription tokens don't feed third-party harnesses, and bridge workarounds carry ban risk (documented bans exist).","title":"Cloud Models & Providers for Hermes Agent — The August 2026 Megathread","body":"**LAST UPDATED: August 23, 2026**\n\n*Compiled from ~30 r/hermesagent threads and hundreds of comments (Aug 1–23), Hacker News discussion threads, live OpenRouter pricing (pulled Aug 23), and independent benchmark data. Every price is the live API rate unless marked as user-reported. Community disagreements are kept side-by-side on purpose. Companion to u/carlonox's request in the [Megathread Submission Request thread](https://www.reddit.com/r/hermesagent/comments/1vq28h5/) — yes, this replaces the stale recommendations.*\n\n---\n\n## Part 1: TL;DR — The Consensus Table\n\n| Decision | Community Pick | Runner-Up | Watch Out For |\n|---|---|---|---|\n| Budget daily driver | DeepSeek V4 Flash 0731 (direct API) | Ollama Cloud Pro $20 | Aug 16 peak/off-peak pricing; cache-pin your setup |\n| Best $20/mo subscription | ChatGPT Plus (Luna, near-unlimited) | Ollama Cloud Pro / Nous Portal Plus | Luna context-capped ~256k on subs |\n| Heavy agentic use ($200 tier) | Codex Pro via OAuth (1.4B tokens/mo reported) | Claude Max 20x | Claude sub tokens can't be used in Hermes directly |\n| Quality ceiling | Claude Opus 5 / Sonnet 5 (API) | GPT-5.6 Sol high | Sol burns tokens; \"eat your tokens like cookies\" |\n| Free testing | OpenRouter free tier (~14 models) | NVIDIA NIM (free, 1M ctx) | Rate limits; free models train on your data |\n| Orchestration model | GLM-5.3 (tied #1 agentic index with Opus 5) | GPT-5.6 Terra high as default | Beijing-hours congestion; token-hungry |\n| New this month worth trying | Qwen3.8-27B hosted ($0.40/$3 per M) | Meta Muse Spark 1.2 (budget) | Hosted quality varies by provider |\n\n**The August 2026 earthquake in one paragraph:** DeepSeek raised prices Aug 16 (peak/off-peak, up to 3–5× at peak), OpenRouter got bought by Stripe for $7B+, GPT-5.6 Sol dropped 50% on OpenRouter only while Luna (~80% cut) and Terra (20%) got their own cuts, Gemini killed postpay billing Aug 13, and Qwen3.8-27B landed as a hosted option for those without the VRAM (\"matches GPT-5.6 Luna on Max reasoning… for the cost of electricity\" — one local runner's take). Three of those five change everyone's math.\n\n---\n\n## Part 2: The Decision Tree\n\nStart here, follow the first line that matches you:\n\n- **\"I spend under $10/mo and want it to stay that way\"**\n  → DeepSeek V4 Flash direct + aggressive caching (off-peak hours). Below even that: OpenRouter free tier or NVIDIA NIM, accepting rate limits.\n- **\"I have a $20/mo budget and want zero admin\"**\n  → Ollama Cloud Pro if privacy (zero-retention) matters most; ChatGPT Plus (Luna) if you're already in that ecosystem; Nous Portal Plus for 300+ models on one key with Hermes-native integration.\n- **\"I run agents all day and pay-as-you-go is eating me alive\"**\n  → Flat subscription, full stop. As one poster put it: \"Agents love getting stuck in retry loops and vaporizing pay-as-you-go balances while you sleep, so locking in a flat monthly sub is just basic financial self-defense.\" Codex $200 OAuth is the current king of raw tokens (1.4B/mo reported, ~90% cache hits).\n- **\"I need maximum quality on hard tasks regardless of cost\"**\n  → Claude Opus 5 or Sonnet 5 via direct API, or GPT-5.6 Sol at high effort. Note: Opus-class is the bar everyone measures against, but one camp now argues GLM-5.3 tied Opus 5 on the agentic index at a fraction of the price.\n- **\"I want to orchestrate multi-agent Hermes setups\"**\n  → The emerging community pattern: a strong orchestrator + cheap subagents. Reported working combos: GLM orchestrating + DeepSeek Flash subagents; Terra-high default with Sol reserved for stuck-points.\n- **\"I'm price-sensitive but don't trust Chinese data handling\"**\n  → This disqualifies DeepSeek/Z.ai/Kimi/MiniMax direct. Your lane: Ollama Cloud (zero retention), Gemini free tier (mind the key-leak hygiene), or Western-hosted open models via DeepInfra/Fireworks — noting these still serve Chinese-origin weights, so the distinction is about who operates the endpoint. Free tiers generally mean your data trains something.\n\n---\n\n## Part 3: Provider Tiers\n\n### 🥇 Tier 1 — The community defaults\n\n**DeepSeek direct API** — still the price floor off-peak for most users, superb tool-calling, cache-hit rates up to 97% reported. Dissenting camp exists post-hike: some argue the peak/off-peak complexity means \"not worth it via direct API anymore\" since gateways kept old prices — weigh whether you'll actually hit off-peak windows. The Aug 16 hike introduced peak/off-peak windows (peak = 01:00–04:00 & 06:00–10:00 UTC; weekends all-off-peak). Budget win reported: $5.80 over 3 weeks with caching discipline `[user-reported]`. Watch-for: no vision; your data may train their models.\n\n**OpenRouter** — 500+ models, one key, auto-fallback, widely praised developer experience (though one HN commenter reports smart routing producing worse results at higher cost). Now owned by Stripe ($7B+, closed Aug 19). Sentiment split: praised for UX, panned for support (\"No support exists when things go wrong\") and post-acquisition anxiety (\"RIP OpenRouter… enshrittification will be inevitable\" — one voice among many, not consensus). **Critical operational tip from the burn threads: pin to ONE provider per model. Provider-mixing destroys prompt-cache hits and can double or triple your effective cost.**\n\n**ChatGPT/Codex subscriptions** — Luna at $20 is described as \"practically unlimited\" and needs ~2.8× fewer tokens than DS Flash for similar task performance. Codex $200 = 20× monthly tokens via OAuth. Watch-for: context caps (~256k–272k) on subscription tiers.\n\n### 🥈 Tier 2 — Strong alternatives\n\n**Nous Portal ($20/mo Plus)** — the official Hermes-native path: 300+ models, one subscription, tool gateway, 900k context added Aug 19. Verdicts are genuinely split: fans report 50M tokens ≈ $2 during promos; critics calculate only \"$22 in API value\" for the base plan vs Codex's \"$70–100\" `[anecdotal]`. One user-reported burn: ~€70/week through Portal API credits, mostly DS Flash 0731 `[anecdotal]`. DeepSeek V4 Flash notably did NOT raise prices there after Aug 16. If you want the zero-config Hermes experience, this is the path of least resistance.\n\n**Ollama Cloud Pro ($20)** — the privacy pick: zero data-retention policy, DS V4 Flash included. One user claims 15k+ requests/month `[anecdotal]`; another disputes the weekly-vs-monthly framing. Either way: rare billing complaints.\n\n**OpenCode Go/Zen ($10 + credits)** — 2× credit value, Hermes-fallback-friendly. Turbulent August: DeepSeek moved to a China-based provider (ZDR concerns), the 0813 model was briefly broken during pricing rollout, and rate-limit complaints surfaced. They've said they're working on matching old prices with own inference. Watch this space before committing.\n\n**Z.ai (GLM-5.3)** — tied #1 agentic index with Opus 5 (HN, Aug 18 release). The Max coding plan (~$120/mo) gets called \"the best/most affordable plan\" for heavy users. Watch-for: Beijing-hours congestion makes Hermes unresponsive during EU/US mornings; API pricing for 5.3 took days to appear after launch (now live at $1.4/$4.4).\n\n**DeepInfra** — cheap, good caching (\"handles caching much better… saves a ton on reasoning via cache reads\"). One quality complaint when accessed through OpenRouter pairing. Worth benchmarking yourself.\n\n### 🥉 Tier 3 — Budget/free lanes (all with catches)\n\n| Lane | The deal | The catch |\n|---|---|---|\n| OpenRouter free tier | ~14 free models incl. GLM-5.2 | Frequent 429s; free models generally train on your data |\n| NVIDIA NIM | Free, 1M ctx, Nemotron 3.5 family | Model quality mixed for agent work |\n| Hetzner Experiments | Free GLM/DS/Kimi/Qwen endpoints, 5M tokens/day | Signup friction, ban risk for sustained agent use, models rotating out |\n| Omniroute/Orcarouter | Free aggregators, OAuth-based | Unvetted; treat credentials accordingly |\n| Meta Muse Spark 1.2 | Cheaper than pre-hike DS Flash; transparent pricing | New, availability-limited |\n| Tencent Hy3 | \"Benches like DS Pro, costs like Flash\" | Chatty output style |\n\n⚠️ Gray-market resellers (discounted token sites, \"unlimited Claude\" bridges) exist and some users report reliability. These carry real account-ban and payment risk. Anthropic has both sued over subscription advertising and sent legal letters to harness makers bridging subs into APIs. That's your risk landscape; this megathread doesn't recommend them.\n\n---\n\n## Part 4: Model-by-Model Quick Verdicts (cloud)\n\n| Model | Best for | $/M in-out (live OR) | Community watch-fors |\n|---|---|---|---|\n| Claude Opus 5 | The quality bar | $5/$25 | Pricey; \"none is universally as good\" vs \"Opus doesn't even have quality anymore\" — both camps active |\n| Claude Sonnet 5 | Reliable default premium | $2/$10 | Newest Sonnet; cheaper than 4.6 per token |\n| Claude Sonnet 4.6 | Previous-gen default | $3/$15 | Subscription can't feed Hermes directly |\n| GPT-5.6 Sol | Hard planning, one-shots | $2/$10 (−50% on OR mid-Aug) | Token furnace; polarizing (\"beast\" vs \"overengineers everything\") |\n| GPT-5.6 Luna | Daily driver | $0.20/$1.20 | \"Feels dumb in long context\" reports vs \"seriously underrated\" |\n| GPT-5.6 Terra | Cheap default | $2/$12 | The \"Terra high default, Sol only when stuck\" pattern |\n| Gemini 3.7 Flash | Multimodal, free tier | $0.38/$1.88 | Prepay-only since Aug 13; key-leak horror stories ($82k/48h) |\n| DeepSeek V4 Flash | Budget workhorse | $0.054/$0.11 (off-peak) | Peak/off-peak windows; no vision; data training |\n| DeepSeek V4 Pro | Sanity-check tier | $0.414/$0.828 | Cache-hit pricing up to 12× differential |\n| Qwen3.8-27B hosted | Local-class quality, no GPU | $0.40/$3.00 | Hosted quality varies by endpoint provider |\n| Kimi K2.7-Code | Long autonomous runs | $0.67/$3.40 | Token-hungry; \"digest my entire month's tokens in minutes\" |\n| Kimi K3 | Open-weight flagship | $3/$15 | Now pricier than Sol per HN commenters |\n| GLM-5.3 | Orchestration, agentic | $1.4/$4.4 | Beijing-hours lag; no multimodality |\n| MiniMax M3 | Bill-over-brilliance round-the-clock | $0.30/$1.20 | Most polarizing model in the corpus: \"stupidest I've tried\" AND \"code with it all day for less than a dollar\" |\n| Grok 4.5/4.6 | Rising value pick | $2/$6 | \"Opus-level\" claims are user-reported, not benchmarked |\n| Qwen3.8-Max 2.4T | Frontier open-weight | $2/$6 | Massive; new |\n\n---\n\n## Part 5: The Cost Reality Section (numbers as reported, single-source flagged)\n\n**Horror stories (learn from them):**\n- $5 in ONE day on DS Flash via OpenRouter — root causes found in-thread: provider-mixing killing cache hits + Hermes skill bloat re-sent every cron session\n- $200+ credits burned on \"simple coding tasks\" via OpenRouter; the poster switched to subscription OAuth afterward\n- ChatGPT business plan gone in 2 days — consensus in-thread: Sol was the culprit, not Hermes\n- $25/day Gemini 3 Flash for two users + ops `[anecdotal]`\n- Key-leak cautionary tales outside Hermes: $82k in 48h from one leaked Gemini key\n\n**Budget wins (also real):**\n- DS Flash direct with caching discipline: $5.80 over 3 weeks including 140M-token sessions\n- Codex $200 OAuth: 1.4B tokens/month, never limit-hit\n- MiniMax M3: full-time coding under $1/day\n- Nous Portal promo: 50M tokens for ~$2\n- 2.5B DeepSeek tokens for $40\n\n**The three habits that separate the two columns:** pin providers, cache aggressively, cap your cron sessions' skill payloads.\n\n---\n\n## Part 6: FAQ\n\n1. **Q: Is OpenRouter still safe to use after the Stripe acquisition?**\n   **A:** Yes today; service continues normally. Long-term pricing/support direction is the open question — the community is watching, not fleeing.\n\n2. **Q: Why did my DeepSeek bill triple?**\n   **A:** Aug 16 peak/off-peak pricing. Peak is 01:00–04:00 & 06:00–10:00 UTC daily; weekends are all off-peak. Shift cron jobs to off-peak and check your cache-hit rate.\n\n3. **Q: Can I use my Claude Max subscription inside Hermes?**\n   **A:** Not directly — Anthropic subscription tokens don't feed third-party harnesses, and bridge workarounds carry ban risk (documented bans ex","offTopic":false},{"id":"8aed4bda-a2b4-4d67-9106-1cd9530b0aa8","excerpt":"🤖 Weekly OpenRouter Comparative Scan Report (2026-08-23) — # 🤖 Weekly Model Recommendations & Comparative Delta Analysis\n**Generated on:** 2026-08-23 UTC  \n**Evaluation Engine:** LLM Agent with Live Web Search Grounding  \n**Source Data:** Live [OpenRouter Model Catalog](https://openrouter.ai/models)\n\n---\n## 📑 Executive","url":"https://github.com/axiomantic/claude-threepio/issues/4","role":"request","weight":0.86192334,"occurredAt":"2026-08-23T00:40:21.000Z","sourceKey":"github","sourceName":"GitHub","credibility":0.78,"venue":"axiomantic/claude-threepio","intent":"feature_request","painScore":0.21,"sentiment":1,"confidence":0.7123333,"matchedPatterns":["recommend","missing_feature","product:anthropic"],"statement":"The current lineup overweights experimental/free options and lacks the newest stable frontier anchors.","title":"🤖 Weekly OpenRouter Comparative Scan Report (2026-08-23)","body":"# 🤖 Weekly Model Recommendations & Comparative Delta Analysis\n**Generated on:** 2026-08-23 UTC  \n**Evaluation Engine:** LLM Agent with Live Web Search Grounding  \n**Source Data:** Live [OpenRouter Model Catalog](https://openrouter.ai/models)\n\n---\n## 📑 Executive Summary\n\nOpenRouter’s 2026 catalog has shifted toward a China-led value frontier: DeepSeek V4 Flash/Pro, Qwen 3.7/3.8, Z.ai GLM 5.2/5.3, Google Gemini 3.x Flash/Pro, and Anthropic Opus/Sonnet 4.6–5 now cover nearly every tier with stronger 1M-context and better benchmark/cost efficiency than many legacy picks. Recent coverage and benchmark roundups indicate DeepSeek V4 Flash is near the top of coding leaderboards on price-adjusted performance, while Claude Sonnet 4.6/5 remains the preferred Sonnet-class coding/agentic workhorse and Claude Opus 4.8/5 remains the quality ceiling for frontier reasoning. OpenRouter’s own benchmarks dashboard and benchmark API now aggregate live model intelligence, coding, and agentic scores, making it clear that the current claude-threepio lineup is under-updated in Opus/Fable/Mythos and can be materially improved by swapping in newer GA releases such as DeepSeek V4 Pro 0813, Claude Opus 4.8, Claude Sonnet 5, GPT-5.6 Sol/Terra/Luna, GLM-5.3/5.2, Qwen3.8 Max/2.4T, and Gemini 3.1/3.7 families. Free and ultra-low-cost coverage is now much richer than before, with multiple free MoE models and very low-cost specialist options that can cleanly anchor Haiku/Sonnet subagent routing. [6][7][9][10]\n\n### 🌐 Web Search Findings & Benchmark Grounding\n\n- OpenRouter’s live benchmarks page reports 6 benchmarks and 2,465,315 task evaluations as of Aug 22, 2026, and its benchmark API aggregates Artificial Analysis, Design Arena, and OpenRouter task evals. [10][6]\n- Recent model leaderboard coverage highlights DeepSeek V4 Flash as a top cost-adjusted coding model, with low latency and strong Artificial Analysis coding performance, while GPT-5.6 Luna/Terra/Sol and Claude Sonnet 4.6/5 remain high-end agentic/coding choices. [9][5]\n- Recent release coverage shows major 2026 frontier updates across OpenAI GPT-5.2/5.4/5.5/5.6, Anthropic Claude Opus 4.6/4.7/4.8/5 and Sonnet 4.6/5, Google Gemini 3.1/3.5/3.6/3.7, DeepSeek V4 Flash/Pro, Qwen3.5/3.6/3.7/3.8, Z.ai GLM 5.x, Moonshot Kimi K2.5/2.6/2.7/3, NVIDIA Nemotron 3 family, and Poolside Laguna S/XS 2.1. [5][8][9]\n\n---\n## 📊 Tier-by-Tier Comparative Delta Analysis\n\n### 🏷️ OPUS TIER (Heavyweight Reasoning & Complex Architecture) (`claude-opus-4`)\n- **Current Recommended (in `claude-threepio`):** `DeepSeek V4 Pro (deepseek/deepseek-v4-pro) — $0.50/$1.00`\n- **Proposed Recommended:** **`DeepSeek V4 Pro 0813 (deepseek/deepseek-v4-pro-0813) — $1.12/$3.37`**\n- **Recommendation Shift Rationale:** The current recommendation is still a strong value pick, but the GA DeepSeek V4 Pro 0813 is the better default because it is the current production snapshot, retains 1M context, and aligns better with today’s frontier reasoning/coding stack. Claude Opus 4.8 remains the pure quality ceiling, but its price is far above the value optimum for the Opus tier.\n\n*This tier should present a clear spectrum from free exploratory models through a rational default and then a quality ceiling. The current lineup overweights experimental/free options and lacks the newest stable frontier anchors. DeepSeek V4 Pro 0813 replaces the older DeepSeek V4 Pro at +$0.62 input (+124%) and +$2.37 output (+237%), but the delta is justified because the newer snapshot is the GA release with better stability and a tighter fit for long-horizon reasoning. Claude Opus 4.8 should be added as the premium ceiling, because benchmark coverage and leaderboard commentary indicate Opus-class models remain at the top of general reasoning and professional work. [5][9] Retain the free/open options, but prioritize the models with strong orchestration and 1M context: Nemotron 3 Ultra free, Ox Alpha, GLM 5.2 free, and Qwen 3.7 Plus. GLM 5.3 and Qwen3.8 Max should be added as higher-ceiling alternatives because live catalog descriptions place them in complex software engineering/agentic workflows, and they represent the newest wave of large reasoning models. [9][10]*\n\n#### 🔄 Model Swaps & Lineup Adjustments\n| Action | Previous Model | Proposed Model | Price Delta | Benchmark & Engineering Justification |\n| :--- | :--- | :--- | :--- | :--- |\n| **SWAP** | DeepSeek V4 Pro (deepseek/deepseek-v4-pro) | `DeepSeek V4 Pro 0813 (deepseek/deepseek-v4-pro-0813)` | $0.50/$1.00 vs $1.12/$3.37 (+124.0% input, +237.0% output) | GA release snapshot is preferable for production: stronger stability and newer frontier tuning, while keeping 1M context and preserving the same DeepSeek reasoning family. |\n| **RETAIN** | Nemotron 3 Ultra 550B (Free) (nvidia/nemotron-3-ultra-550b-a55b:free) | `Nemotron 3 Ultra 550B (Free) (nvidia/nemotron-3-ultra-550b-a55b:free)` | $0.00/$0.00 vs $0.00/$0.00 (0% delta) | A free 1M-context MoE gives zero-cost exploration and long-context experimentation with meaningful orchestration capacity. |\n| **RETAIN** | Ox Alpha (Stealth Free) (stealth/ox-alpha) | `Ox Alpha (Stealth Free) (stealth/ox-alpha)` | $0.00/$0.00 vs $0.00/$0.00 (0% delta) | Keeps a zero-cost anonymous fallback for bursty traffic and A/B testing. |\n| **RETAIN** | GLM 5.2 (Free) (z-ai/glm-5.2:free) | `GLM 5.2 (Free) (z-ai/glm-5.2:free)` | $0.00/$0.00 vs $0.00/$0.00 (0% delta) | Free reasoning baseline with 1M-context class capability in the current catalog and strong long-horizon workflow fit. |\n| **RETAIN** | Qwen 3.7 Plus (qwen/qwen3.7-plus) | `Qwen 3.7 Plus (qwen/qwen3.7-plus)` | $0.32/$1.28 vs $0.32/$1.28 (0% delta) | Balanced cost/performance anchor and a strong value alternative for reasoning-heavy use cases. |\n| **ADD** | None | `Anthropic: Claude Opus 4.8 (anthropic/claude-opus-4.8)` | $5.00/$25.00 vs new option | Frontier quality ceiling for the tier; Opus-class models remain the best fit when users prioritize maximum reasoning and professional-grade output over cost. |\n| **ADD** | None | `Z.ai: GLM 5.3 (z-ai/glm-5.3)` | $1.40/$4.40 vs new option | Newest GLM reasoning release in-catalog; useful for long-horizon agent tasks and software engineering workloads. |\n| **ADD** | None | `Qwen: Qwen3.8 2.4T A95B (qwen/qwen3.8-2.4t-a95b)` | $2.00/$6.00 vs new option | Modern sparse MoE frontier option with a large parameter budget and strong fit for high-ceiling reasoning in the open-weight ecosystem. |\n\n#### 📋 Full Proposed Options for Tier\n| Model ID | Display Name | Live Price ($In / $Out) | Context | Status | Link | Rationale |\n| :--- | :--- | :--- | :--- | :--- | :--- | :--- |\n| [`nvidia/nemotron-3-ultra-550b-a55b:free`](https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b:free) | Nemotron 3 Ultra 550B (Free) | **$0.00/$0.00** | 1M Context | Alternative | [View](https://openrouter.ai/nvidia/nemotron-3-ultra-550b-a55b:free) | Zero-cost exploration and long-context free fallback. |\n| [`stealth/ox-alpha`](https://openrouter.ai/stealth/ox-alpha) | Ox Alpha (Stealth Free) | **Live API** | 1M Context | Alternative | [View](https://openrouter.ai/stealth/ox-alpha) | Free stealth fallback for evaluation, burst routing, and privacy-conscious usage. |\n| [`deepseek/deepseek-v4-pro-0813`](https://openrouter.ai/deepseek/deepseek-v4-pro-0813) | DeepSeek V4 Pro 0813 | **$1.12/$3.37** | 1M Context | **⭐ Recommended** | [View](https://openrouter.ai/deepseek/deepseek-v4-pro-0813) | Best value default for heavyweight reasoning with 1M context. |\n| [`qwen/qwen3.7-plus`](https://openrouter.ai/qwen/qwen3.7-plus) | Qwen 3.7 Plus | **$0.32/$1.28** | 1M Context | Alternative | [View](https://openrouter.ai/qwen/qwen3.7-plus) | High-value open-weight reasoning choice with strong long-context support. |\n| [`z-ai/glm-5.2`](https://openrouter.ai/z-ai/glm-5.2) | GLM-5.2 | **$0.97/$3.04** | 1M Context | Alternative | [View](https://openrouter.ai/z-ai/glm-5.2) | Established long-context reasoning model that remains highly competitive for agent tasks. |\n| [`z-ai/glm-5.3`](https://openrouter.ai/z-ai/glm-5.3) | GLM 5.3 | **$1.40/$4.40** | 1M Context | Alternative | [View](https://openrouter.ai/z-ai/glm-5.3) | Newest GLM frontier option for users who want the latest Z.ai release. |\n| [`qwen/qwen3.8-2.4t-a95b`](https://openrouter.ai/qwen/qwen3.8-2.4t-a95b) | Qwen3.8 2.4T A95B | **$2.00/$6.00** | 1M Context | Alternative | [View](https://openrouter.ai/qwen/qwen3.8-2.4t-a95b) | High-ceiling sparse MoE model for frontier reasoning experiments. |\n| [`anthropic/claude-opus-4.8`](https://openrouter.ai/anthropic/claude-opus-4.8) | Claude Opus 4.8 | **$5.00/$25.00** | 1M Context | Alternative | [View](https://openrouter.ai/anthropic/claude-opus-4.8) | Premium ceiling option with the strongest Anthropic-quality positioning in this tier. |\n\n---\n\n### 🏷️ SONNET TIER (Agentic Coding, Tool Use & Everyday Workhorse) (`claude-sonnet-4-5`)\n- **Current Recommended (in `claude-threepio`):** `DeepSeek V4 Flash (deepseek/deepseek-v4-flash) — $0.08/$0.15`\n- **Proposed Recommended:** **`DeepSeek V4 Flash Latest (~deepseek/deepseek-v4-flash-latest) — $0.06/$0.13`**\n- **Recommendation Shift Rationale:** The current recommended model is already excellent, but the latest DeepSeek Flash alias should replace the fixed snapshot because it tracks the newest production revision at lower price and the same 1M context. Claude Sonnet 4.6/5 should also be kept as the premium high-throughput coding ceiling in the tier.\n\n*This tier needs the strongest coding/agentic value-per-dollar and reliable tool use. The price-sensitive default should move from DeepSeek V4 Flash 0423 to the latest redirect alias, reducing input cost by 25% and output cost by 13.3% while preserving the same family and 1M context. The current lineup is already strong, but it can be improved by adding explicit specialist models that have emerged as coding/workflow leaders: Qwen3 Coder Next, KAT-Coder-Pro V2.5, Poolside Laguna S 2.1, Google Gemini 3.7 Flash, and OpenAI GPT-5.2-Codex / GPT-5.3-Codex. Benchmark coverage suggests DeepSeek V4 Flash is a top value coding model, while Claude Sonnet 4.6/5 remains among the strongest general agentic coding models. [9][5] The tier should keep free, budget, workhorse, specialist, and premium ceiling options all visible. [9][10]*\n\n#### 🔄 Model Swaps & Lineup Adjustments\n| Action | Previous Model | Proposed Model | Price Delta | Benchmark & Engineering Justification |\n| :--- | :--- | :--- | :--- | :--- |\n| **SWAP** | DeepSeek V4 Flash (deepseek/deepseek-v4-flash) | `DeepSeek V4 Flash Latest (~deepseek/deepseek-v4-flash-latest)` | $0.08/$0.15 vs $0.06/$0.13 (-25.0% input, -13.3% output) | Latest redirect keeps the same model family while automatically tracking the newest revision; lower cost makes it the best Sonnet-tier default. |\n| **RETAIN** | North Mini Code (Free) (cohere/north-mini-code:free) | `North Mini Code (Free) (cohere/north-mini-code:free)` | $0.00/$0.00 vs $0.00/$0.00 (0% delta) | Free coding-specialist fallback remains valuable for ultra-low-cost subagent routing. |\n| **RETAIN** | Laguna S 2.1 (Free) (poolside/laguna-s-2.1:free) | `Laguna S 2.1 (Free) (poolside/laguna-s-2.1:free)` | $0.00/$0.00 vs $0.00/$0.00 (0% delta) | Keeps a free agentic coding model from a code-first vendor. |\n| **RETAIN** | Gemma 4 31B (Free) (google/gemma-4-31b-it:free) | `Gemma 4 31B (Free) (google/gemma-4-31b-it:free)` | $0.00/$0.00 vs $0.00/$0.00 (0% delta) | Free general-purpose multimodal option broadens the Sonnet tier’s fallback coverage. |\n| **ADD** | None | `Google: Gemini 3.7 Flash (google/gemini-3.7-flash)` | $0.38/$1.88 vs new option | Recent Google release optimized for fast agentic workflows and coding; strong fit for tool-heavy work. |\n| **ADD** | None | `Qwen: Qwen3 Coder Next (qwen/qwen3-coder-next)` | $0.12/$0.80 vs new option | Open-weight coding specialist with excellent cost/performance for repo-scale coding and agentic development. |\n| **ADD** | None | `Kwaipilot: KAT-Coder-Pro V2.5 (kwaipilot/kat-coder-pr","offTopic":true},{"id":"73ee29d0-b23c-4d2f-99df-e169b12f3806","excerpt":"Cost & Token Optimization Megathread — Hermes Agent (June 2026) — **LAST UPDATED:** June 21, 2026  \n**Sourced from:** r/hermesagent cost threads, official Hermes context compression/caching docs, provider pricing pages (DeepSeek, Anthropic, OpenAI), GitHub issue #4379 (token overhead analysis).\n\n---\n\n## TL;DR — What's ","url":"https://www.reddit.com/r/hermesagent/comments/1ud03si/cost_token_optimization_megathread_hermes_agent/","role":"pricing","weight":1.0455577,"occurredAt":"2026-06-22T23:06:58.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"hermesagent","intent":"pricing_complaint","painScore":0.19844228,"sentiment":0.06451613,"confidence":0.8724306,"matchedPatterns":["how_can_i","free_tier","product:anthropic"],"statement":"| |----------|----------------|-------------|-----------|-------------| | **API model** | DeepSeek V4 Flash | $3-10/mo | DeepSeek V4 Pro (75% off) | $8-20/mo | | **Local model** | Qwen 3.6-27B on RTX 3090 | $0/mo (electricity only) | Qwen…","title":"Cost & Token Optimization Megathread — Hermes Agent (June 2026)","body":"**LAST UPDATED:** June 21, 2026  \n**Sourced from:** r/hermesagent cost threads, official Hermes context compression/caching docs, provider pricing pages (DeepSeek, Anthropic, OpenAI), GitHub issue #4379 (token overhead analysis).\n\n---\n\n## TL;DR — What's This Going to Cost Me?\n\n| Decision | Cheapest Option | Monthly Est. | Runner-Up | Monthly Est. |\n|----------|----------------|-------------|-----------|-------------|\n| **API model** | DeepSeek V4 Flash | $3-10/mo | DeepSeek V4 Pro (75% off) | $8-20/mo |\n| **Local model** | Qwen 3.6-27B on RTX 3090 | $0/mo (electricity only) | Qwen 3.5-9B on any GPU | $0/mo |\n| **VPS + API** | Oracle free tier + Flash | $3-10/mo | Hetzner €4 + Flash | $7-14/mo |\n| **All-in (VPS + API + tools)** | ~$15-25/mo | Real community budget | ~$30-35/mo | u/Background-Remote765 |\n| **Token overhead** | ~73% of each call is fixed | Use toolset trimming | ~40% after trimming | hermes-token-router |\n\n---\n\n## Part 1: Provider Pricing — June 2026\n\n### 🥇 DeepSeek (Best Value)\n\n| Model | Input / 1M tokens | Output / 1M tokens | Context | Notes |\n|-------|-------------------|-------------------|---------|-------|\n| **V4 Flash** | **$0.14** | **$0.28** | 1M | 18× cheaper than GPT-5.4 input. Cached: $0.014/M. |\n| **V4 Pro** | $1.74 ($0.435*) | $3.48 ($0.87*) | 1M | *75% off promotional pricing. Strongest reasoning. |\n| **V3.2** | $0.27 | $1.10 | 128K | Legacy but still viable. |\n| **R1** | $0.55 | $2.19 | 128K | Reasoning model. |\n\n**Community consensus:** V4 Flash for everyday agent work. V4 Pro for complex reasoning. The 75% off Pro pricing makes it competitive with Flash for heavy reasoning tasks.\n\n**Direct API vs OpenRouter:** Always use DeepSeek's direct API. OpenRouter adds a markup. No benefit for single-provider use.\n\n### 🥈 Anthropic Claude (Best Quality, Higher Cost)\n\n| Model | Input / 1M | Output / 1M | Context | Cached Input / 1M |\n|-------|-----------|------------|---------|-------------------|\n| **Sonnet 4** | $3.00 | $15.00 | 200K | $0.30 |\n| **Sonnet 4.6** | $3.00 | $15.00 | 200K | $0.30 |\n| **Opus 4.7** | $15.00 | $75.00 | 200K | $1.50 |\n\n**Prompt caching is critical with Claude:** Hermes automatically caches the system prompt + rolling 3-message window. Cache hits cost **90% less** on input tokens. In practice, multi-turn conversations with Claude can see 50-70% effective input cost reduction.\n\n**When Claude is worth it:** Complex multi-step coding, security-sensitive work, tasks where tool-calling reliability is paramount. \"The boring model that follows schema for 6+ tool calls beats the spicy one that talks itself into a ditch.\"\n\n### 🥉 OpenAI (Wide Ecosystem)\n\n| Model | Input / 1M | Output / 1M | Context |\n|-------|-----------|------------|---------|\n| GPT-5.4 | $2.50 | $10.00 | 128K |\n| GPT-5.5 | $5.00 | $20.00 | 272K (Codex) / 1.05M (direct) |\n| GPT-4o | $2.50 | $10.00 | 128K |\n\n### Local Models (Zero API Cost)\n\n| Setup | Hardware Cost | Running Cost | Real-World Tok/s |\n|-------|-------------|-------------|-----------------|\n| Qwen 3.6-27B Q4 | RTX 3090 ($700 used) | ~$15/mo electricity | 25-40 tok/s |\n| Qwen 3.5-9B | RTX 3060 ($200 used) | ~$5/mo electricity | 40-60 tok/s |\n| Qwen 3.6-35B-A3B | M1 Max 64GB | Laptop you own | 61 tok/s (MLX) |\n| DeepSeek V4-Flash local | 2× RTX 5090 | ~$30/mo electricity | Production speed |\n\n**When local breaks even:** If you spend >$20/mo on API tokens, a used RTX 3090 pays for itself in ~35 months on electricity alone. Add VPS savings (no VPS needed) and it's faster.\n\n### Free Tier Options\n\n| Provider | What You Get | Limits | Community Verdict |\n|----------|-------------|--------|-------------------|\n| OpenRouter free models | Multiple models, no credit card | Rate limited, inconsistent tool calling | Test/fallback only |\n| Nous Portal free tier | Bundled Hermes models | Limited context | Solid for light use |\n| DeepSeek | $5 free credit (new accounts) | One-time | Good starter |\n| Google Gemini | Free tier available | Rate limits, 503 errors | \"Nearly destroyed my entire hermes\" — u/Simple_Tune2882. Multiple users report 503s even on paid tier. |\n| Opencode Go | $10/mo → **$60 in API credits** | 6x credit multiplier; Kimi 2.6, Mimo, etc. | **Top recommendation.** \"I've been using Kimi 2.6 pretty hard and in 5 days only managed to use 11% of my monthly allowance.\" |\n| Minimax | $10/mo plan | Minimax 2.7, M3 (\"as good as GPT 5.5-low\") | \"Most overlooked model\" — multiple users |\n| NanoGPT | $8/mo | 60M tokens/week with GLM 5.1 | Extremely cheap if GLM works for you |\n| Ollama Cloud | $20/mo flat | No usage caps | \"Not even getting close to hitting the limits\"; some speed complaints |\n\n### Plan vs API: The Decision Framework\n\nFrom u/getstackfax — the clearest framework on the subreddit:\n\n**Use a plan when:**\n- You're doing heavy daily interactive use\n- You want predictable billing (no $54 surprises)\n- The plan covers models you actually use\n- You're not disciplined enough to audit your API usage weekly\n\n**Use raw API when:**\n- You route carefully between cheap models\n- You keep context small and caching tight\n- You've set budget caps and loop guards\n- You actually check your provider dashboard weekly\n\n**The golden rule:** *\"API is dangerous if the agent is allowed to wander.\"* Plans cap your downside. API requires discipline.\n\n**The dashboard audit checklist** (weekly):\n- Model actually used vs default\n- Cache hit rate (not just \"caching supported\" — verify hits)\n- Input/output token split\n- Fallback/ escalation events\n- Background/heartbeat costs\n- Total cost per useful task\n\n> *\"'Caching supported' and 'your workflow is actually getting cache hits' are different things.\"* — u/getstackfax\n\n---\n\n## Part 2: The Token Overhead Problem\n\n**73% of every Hermes API call is fixed overhead that doesn't change between requests.**\n\nFrom GitHub issue #4379 — community analysis of 6 request dumps from a Telegram + WhatsApp + Cron deployment:\n\n- System prompt: ~3,000-8,000 tokens\n- Tool definitions: ~2,000-15,000 tokens (depends on enabled toolsets)\n- Memory injection: variable (grows over time)\n- SOUL.md / USER.md: ~500-2,000 tokens\n- **Actual conversation:** only ~27% of the call\n\n### Where Your Tokens Actually Go\n\n| Component | Approximate Tokens | Variable? | Can You Trim It? |\n|-----------|-------------------|-----------|-----------------|\n| System prompt | 3,000-8,000 | Somewhat | No — core functionality |\n| Tool definitions | 2,000-15,000 | **Yes** | **Yes — disable unused toolsets** |\n| Memory | 500-5,000+ | Yes, grows | Yes — trim old memories |\n| SOUL.md | 200-2,000 | Yes | Yes — keep it compact |\n| USER.md | 200-1,500 | Yes | Yes — keep it compact |\n| Skills context | 200-3,000 | Yes | Yes — fewer loaded skills |\n| Conversation history | 1,000-50,000+ | Yes, grows | Compression handles this |\n\n### How to Cut Overhead by 40-50%\n\n**1. Toolset trimming (biggest win):**\n```bash\n# See what's enabled\nhermes tools\n\n# Disable toolsets you never use\nhermes config set toolsets.enabled \"['terminal','file','web','skills']\"\n# NOT: \"['terminal','file','web','skills','browser','computer_use','vision','image_gen','tts','discord','spotify','homeassistant','kanban','todo','cronjob','delegation']\"\n```\n\nEvery disabled toolset removes its entire JSON schema from every API call. The `browser` toolset alone can add 1,500+ tokens.\n\n**2. Compact SOUL.md and USER.md:**\n- Target 500-1,000 tokens each\n- Remove examples, verbose instructions\n- Use bullet points, not paragraphs\n- Test: does the agent still behave correctly? If yes, you haven't lost anything.\n\n**3. Memory hygiene:**\n- Run `hermes memory stats` to see memory size\n- Remove stale/duplicate memories\n- Memory grows with every session — prune periodically\n\n**4. Use the hermes-tool-router plugin (in development):**\nPredicts which toolsets are needed before each turn and only sends those definitions. Can cut tool overhead by 60-80% per turn. Currently in testing.\n\n---\n\n## Part 3: Context Compression — How Hermes Saves Tokens\n\nHermes has a dual compression system that fires automatically:\n\n### Layer 1: Agent Compressor (50% threshold)\nFires when prompt tokens reach 50% of the model's context window:\n1. Prunes old tool results (>200 chars are replaced with a placeholder)\n2. Summarizes middle conversation turns into a structured summary\n3. Preserves last N messages (tail) unmodified\n4. On subsequent compressions, **updates** the existing summary instead of re-summarizing\n\n### Layer 2: Gateway Session Hygiene (85% threshold)\nSafety net for long-running gateway sessions (Telegram/Discord). Catches sessions that escaped the agent's own compressor.\n\n### Tuning Compression\n\n```yaml\ncompression:\n  enabled: true\n  threshold: 0.50          # Fire at 50% of context (default)\n  target_ratio: 0.20       # Tail gets 20% of threshold budget\n  protect_last_n: 20       # Minimum tail messages preserved\n```\n\n**Lower threshold = more aggressive compression = lower costs but more context loss.** Default 0.50 is well-tuned. Don't go below 0.30 — you'll lose too much context.\n\n### The Summary Model Matters\n\nThe summary is generated by a separate LLM call. If the summary model's context window is **smaller** than the main model's, compression fails silently and you **lose conversation context with no warning.** Always ensure your auxiliary/compression model has at least as large a context window as your main model.\n\n---\n\n## Part 4: Prompt Caching — Anthropic's 90% Discount\n\nWhen using Claude models, Hermes automatically uses Anthropic's prompt caching:\n\n- **System prompt** — cached across all turns (breakpoint 1)\n- **Last 3 messages** — rolling cache window (breakpoints 2-4)\n- **Cache hit savings:** 90% reduction on input tokens\n- **Real-world impact:** 50-70% effective input cost reduction in multi-turn conversations\n\n**Cache-aware tips:**\n- Don't modify SOUL.md mid-conversation — it invalidates the system prompt cache\n- The rolling 3-message window re-establishes within 1-2 turns after compression\n- TTL configurable: `prompt_caching.cache_ttl: \"5m\"` (default) or `\"1h\"` for slow conversations\n\n```yaml\nprompt_caching:\n  cache_ttl: \"1h\"    # Better for Telegram where turns have gaps\n```\n\n---\n\n## Part 5: Community Cost-Saving Strategies\n\n### Strategy 1: Cheap Model for Easy Tasks, Expensive for Hard\n\nThe \"router\" pattern — most-recommended across all cost threads:\n```yaml\n# config.yaml — main model\nmodel:\n  default: deepseek/deepseek-v4-flash\n\n# For complex tasks, switch mid-session:\n/model anthropic/claude-sonnet-4\n```\n\nUse DeepSeek Flash for 80% of agent work. Switch to Claude only when tool-calling reliability matters or you hit a task Flash can't handle.\n\n**Community consensus on model tiers for agent work:**\n- 🥇 **Daily driver:** DeepSeek V4 Flash, Minimax 2.7, Qwen 3.6-27B (local)\n- 🥈 **Complex reasoning:** DeepSeek V4 Pro, Claude Sonnet 4, Kimi K2.6\n- 🥉 **Budget/light:** GLM 4.7 Flash, Gemma 4 26B, Granite 4.1 8B\n- ⚠️ **Avoid for Hermes:** Gemma 4 (\"bugged with Hermes — tool call issues\" confirmed by multiple users), Gemini (503 errors, \"nearly destroyed my entire hermes\")\n\n### Strategy 2: Local Inference + Cloud Fallback\n\nRun a local model (Qwen, Llama) as primary. Fall back to DeepSeek API when the local model struggles:\n```yaml\nproviders:\n  custom:\n    local-qwen:\n      base_url: http://localhost:1234/v1\n      api_key: not-needed\n      models: [qwen3.6-27b]\n\nmodel:\n  default: custom:local-qwen/qwen3.6-27b\n  fallback: deepseek/deepseek-v4-flash\n```\n\n### Strategy 3: Subagent Delegation for Cost Isolation\n\nDelegate expensive reasoning to short-lived subagents that use cheaper models:\n- Parent agent uses DeepSeek Flash (cheap)\n- Research subagent uses DeepSeek Flash (cheap, isolated context)\n- Only the summary comes back — no token history contamination\n\n### Strategy 4: API Direct, Not OpenRouter\n\nOpenRouter adds markup. If you use primarily one provider:\n- DeepSeek: use direct `deepseek` provider\n- Anthropic: use direct `anthropic` provider\n- OpenAI Codex: use direct `openai-codex` provider\n\nO","offTopic":false},{"id":"551a8aa7-d72c-472f-b1eb-d97c03508be1","excerpt":"Battle of the $20 (or cheaper) providers — Hi all. \n\n  \nI've been testing out different models and providers to see what is the best bang for buck you can get for around $20 if you are not running local models.\n\nI have a Hermes agent running on a VM with 6GB RAM, which I got for an absolute steal of $45 per year (check","url":"https://www.reddit.com/r/hermesagent/comments/1tewdky/battle_of_the_20_or_cheaper_providers/","role":"pricing","weight":0.9671867,"occurredAt":"2026-05-16T15:12:37.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"hermesagent","intent":"pricing_complaint","painScore":0.36,"sentiment":0.07964602,"confidence":0.7111667,"matchedPatterns":["paying_monthly","product:anthropic"],"statement":"You pay $10 per month, and essentially get $60 worth of credits.","title":"Battle of the $20 (or cheaper) providers","body":"Hi all. \n\n  \nI've been testing out different models and providers to see what is the best bang for buck you can get for around $20 if you are not running local models.\n\nI have a Hermes agent running on a VM with 6GB RAM, which I got for an absolute steal of $45 per year (check out the LowEndTalk forum for cheap VPS deals). I use it mainly to maintain a dashboard that does the following:\n\n* Gather news on specific topics from various sources. It then curates them to see if they align with my interests (eg. no sensasionalist crap), summarizes and deduplicates articles.\n* Check the latest benchmarks on different models\n* Scrape my favourite webcomics from Instagram, RSS feeds, Bluesky, whatever, so they are all in one place.\n\nIt also maintains the VPS, so I have it install docker containers for stuff I want, like Mealie or whatever.\n\nLastly, I synced my Obsidian vault where I keep a list of people with birthdays, notes etc. So it can remind me who's birthday it is and what I can buy for them, or other stuff like that. My Obsidian is also where it keeps track of my health stuff. Diet, gym log, etc.\n\nSo, I've been playing around with the following providers. In all cases except Codex and OpenRouter, I used Kimi K2.6 as my main model, and usually tried Gemma4 for some of the tools and auxiliary models:\n\n* Ollama Cloud - $20 per month\n* OpenCode Go - $10 per month\n* NanoGPT - $12 per month (I think you can get $8 if you find a ref link)\n* OpenAI Codex - $20\n* OpenRouter - Free Models only\n\nHere are my findings.\n\n# Ollama Cloud\n\nVery stable. Charges per GPU hours instead of tokens, so as models get more efficient, you actually gain mode usage. Some people say it's a bit slow, but in my experience it was never slow enough to be problematic. \n\nI actually had a hard time hitting my usage limits. I had to run my Hermes Agent, as well as 2 pretty big coding tasks simultaneously before I hit my 5 hour window limit, and this only happened once. The rest of the time, I barely cracked 25%. For Hermes alone, you will likely never hit that limit.\n\nCons, are that you are limited to 3 concurrent connections. Meaning, my example of 2 coding cases and Hermes was pushing it. If I had to chat to Hermes and a cron job fired that used a model, it errored out because I went over the limit of 3 connections. This is something to keep in mind for people running multiple agents or lots of cron jobs and such.\n\n# OpenCode Go\n\nI felt like this was ever so slightly less stable than Ollama, but not enough to be a problem or to stay away from it. Speed was fine, I honestly didn't feel much of a difference between OpenCode and Ollama. You pay $10 per month, and essentially get $60 worth of credits. \n\nOne might think $60 credits is not much, but whether it is an efficiency thing or just the fact that we aren't paying Anthropic pricing, it stretched very far. I never hit my limits. Just like Ollama, on average usage I barely got to 25-30% weekly. Unlike Ollama, you don't have concurrency limits.\n\n  \nThe con for me is that it didn't have the model I wanted for tool calls, Gemma 4. They don't have that on here. They have DeepSeek which is cheap and fast, but Gemma 4 is cheap, fast AND multimodal. Useful for curating news articles or webcomics.\n\n# NanoGPT\n\nThis one seemed sketchy AF at first. It's clearly meant for a specific crowd. It has a ton uncensored text models included in the sub, as well as uncensored image models (Qwen Image and Z Image Turbo) with 100 free image generations per day. They allow you to load up with crypto (or visa if you don't have crypto) and sign in with only a passkey, no need to enter an email or anything, allowing for a degree of anonymity. \n\nKimi on this one was VERY verbose. It thought a lot, and then would output that as messages in Telegram, meaning the chat context grew very, very fast and had to compress every couple of messages. They had Gemma 4 though (a bunch of variations), and using them for tool calls worked fine. Of this list, NanoGPT had the most models available on the sub. Usage limits seemed a lot lower than Ollama and OpenCode. Also worth noting, since the model naming on this one is a bit weird, if you are relying on your main model to maintain it's own config, you need to give it the *exact* model you want to use. If you just tell it to use \"Gemma 4\" then high chance it will take the one not in your sub and complain about you needing to top up credits first.\n\n# Codex\n\nCurrently testing. Ran it for a day and weekly usage is already at 30%. Didn't even push it that hard. Using GPT 5.5 on it. It feels like it is running an excessive number of tool calls whenever I give it a task. Doing random searches, terminal commands, notes, etc. I'll see if I hit my weekly in 3 days or not. I probably will.\n\n# OpenRouter\n\nThe standard free models are extremely unreliable and often hit rate limits. However they also frequently have preview models that work very nicely for a week or 3, and are worth at the very least using for tool calls. They recently had Tencent Hy3 for free which even now is topping the LLM Leaderboard on OpenRouter. It is very much worth having an OR API key in your back pocket that you can plug into an auxiliary function or some cron jobs to save usage when things like this happen.\n\n# Honorable Mention\n\n**Nous Portal** \\- You pay $20, you get $22 credits. Not a lot of savings. However they do have some free models from time to time as well. Right now they have Step 3.5 Flash and Deepseek V4 Flash for free. Need to top up your wallet before you can use them though. Like OpenRouter, worth having a key in your back pocket for the occasional freebie.\n\n# My plan going forward\n\nOnce this month's codex runs out, I think I will likely stick with **OpenCode Go + NanoGPT**. I will use OpenCode Go for my main model, profiles, and maybe a bit of coding, and NanoGPT for auxiliary models and free image generation. I am paying $8 per month for Nano instead of $12, not sure how I got that discount, think it was an affiliate link probably. This means, my total setup will be **$18 per month** (or $22 if you don't get a discount) and I have access to a TON of models. I then still have some credits in Nous Portal and OpenRouter on the off chance I need something very niche.","offTopic":false}],"breakdown":[{"sourceKey":"reddit","sourceName":"Reddit","count":4},{"sourceKey":"github","sourceName":"GitHub","count":1}],"total":5}}