{"data":{"items":[{"id":"a5e2a5c8-413f-4723-b2ed-6ef610e1e860","excerpt":"[FEEDBACK] I spent the last 4 months designing automation systems for my company as someone who had never touched coding before, and here is what worked for me — This is all based on my experience. I’ve spent over 6 months in total working on AI setups alone for my business, and most of the work was focused on automati","url":"https://www.reddit.com/r/growmybusiness/comments/1t4lkpp/feedback_i_spent_the_last_4_months_designing/","role":"demand","weight":1.1911929,"occurredAt":"2026-05-05T17:05:08.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"growmybusiness","intent":"alternative_search","painScore":0.33,"sentiment":0.7619048,"confidence":0.8956337,"matchedPatterns":["doesnt_work","switching_from","currently_i_use","manual_process","product:chatgpt"],"statement":"Months of moving from one model to another, months of trying to integrate the basic paid versions of ChatGPT, DeepSeek, and Claude into my workflow.","title":"[FEEDBACK] I spent the last 4 months designing automation systems for my company as someone who had never touched coding before, and here is what worked for me","body":"This is all based on my experience. I’ve spent over 6 months in total working on AI setups alone for my business, and most of the work was focused on automating some of the tasks that used to be very time-consuming. About 2 months were wasted trying multiple setups before I discovered Claude Code and started actually building systems that work.\n\n# Discovering Claude Code\n\nAs you can imagine, this was THE moment for me. Months of moving from one model to another, months of trying to integrate the basic paid versions of ChatGPT, DeepSeek, and Claude into my workflow. Experimenting with those custom AI agents (and actually paying about $100 for a subscription for one of these), with barely any success.\n\nChatGPT/Claude projects looked cool in theory, but had no permanent memory outside the chats. I couldn’t teach them to perform anything beyond the simplest tasks, and giving them perms to actually edit my sheets/docs; learn and improve was pain in the ass. Each upgrade meant me having to make yet another doc (or edit the existing one) in the project’s memory, and it never really meant too much progress. Then a client of mine showed me Claude Code, a system I wrongfully ignored because I’m not a developer and felt like I couldn’t make any use of it. Boy was I wrong.\n\n# Honeymoon Period\n\nSince I discovered Claude Code, my free time got deleted, I gained 5 kilos, and I have been glued to my PC more than during my most hardcore gaming days. It truly felt insane at first, the thing built a full app for me in a single day, I just described what I wanted. Then it optimized it, helped me build the structure, did everything, and even designed it. I was like “Fuck, AI’s gonna replace us all” and I mean this as no joke. For literal weeks, I was lying in bed at night thinking how my agency, my life’s work gonna crumble before my own eyes because AI can now do some insane stuff and I’ll have to pivot into tourism or something as far from AI as possible.\n\nThis period was truly amazing because I never realized how quickly time can pass when you do something you love - building and inventing. I don’t know the number of “AI-powered tools” I planned to do and the times I felt like this is the opportunity for me to become a billionaire. Until I slowly realized one hard truth.\n\n# AI is amazing, but it’s not all-powerful\n\nAs I started actually using the tools I’ve built and actually putting them into practice with my employees, issues were emerging one after another. A bug here, an issue there, then a random loop that eats all my API credits. I would usually just be like “Okay, let’s go again”, but as I continued, it was more obvious that AI can make the big-picture stuff in moments, but the actual, fine-tuned, working systems? For that, you’ll need weeks. I’m not a developer, so I don’t know how to better put this, but it felt like the AI built a house, and from the outside, it looked totally normal. Then you start digging into the walls and foundations and actually using the house, and you realize most planks are rotten, the bricks are layered unevenly, the foundation has holes in it, and every time you try to do an actual walk to the kitchen and back (a full workflow), multiple things break. Not knowing how to write a single line of code didn’t help here at all, so I tried using AI to actually do full-checks and fix the issues. It worked, to a certain extent - in a way that it gets off rails, I put it back, then it drives to the next spot (task), falls off rails again, and the process repeats.\n\nThis actually taught me a ton and brought me back to my philosophy roots and the 80/20 rule. AI can do 80% of the work really fast and really well, but the remaining 20% needed to make the entire system actually *work* in practice takes weeks.\n\n# The middle ground, the reality\n\nI quickly realized one thing - AI automation is amazing as a *support* system, but for actual, quality work, you need people. No AI brain can replace a human one, and no AI tool can do what a quality employee can. I never even thought about “replacing my team with AI” because I honestly don’t give two shits about making more money over ruining loyal people’s lives, but still, I was happy to know the limits of the AI.\n\nBack on the topic, I actually tested multiple workflows at this stage - a single agent with all the knowledge vs multiple specialized agents. Claude Code vs Codex vs OpenClaw-like tools. Each of the workflows had its own advantages and shortcomings, that I’ll try to summarize here:\n\n* Single agents (Claude Code and Codex) work amazing for strategy, high-level tasks. The more knowledge they have in their md files the better, but you have to be careful because of the active memory limitations. The architecture alone cannot support too much knowledge, and if you try to use one agent for, let’s say, digging, evaluating, reaching out, and quality-checking LinkedIn leads, it won’t work that well. However, a single agent with a ton of knowledge about the grand plan to oversee the process and qualify leads, and then specialized, minor agents with very well-defined skills for digging and writing outreach messages will work well. Separate tasks fall into the specialized agent’s hands and they *actually* do an amazing job with a clear set of rules/instructions.\n* Multiple agents work well, but they have their risks too. If you overspecialize and have each agent have knowledge about only their job, consider only their job, you will get a system that looks like a chain where every link was made individually by a different smelter, and none of them knows about the other links, or even less the entire chain. The quality of the entire workflow just won’t be there.\n\nMy solution was a mix of both - larger, single agents with all the knowledge for ideating/strategy tasks and smaller, minor agent with a narrow, specific set of specialized skills for the execution of specific tasks. This resulted with the best quality, I’d say almost 70%-80% of what a human can produce.\n\n# However, the next issue I faced was: Inconsistency\n\nAI ALWAYS pigeonholes into certain pre-defined, approved workflows, and you can’t really deviate from that too much. If you teach it how to write a LinkedIn outreach message, and then reiterate time after time until it learns a good pattern, that pattern will be almost all it does. Won’t be an issue at first and you’ll be like “damn this is fucking amazing”, but then 4 weeks in you’ll see that every new campaign somehow sounds very close to the old ones. If it tries a new approach, it will usually fail miserably, but if you teach it that new pattern now - *that* will be all it does. That’s why we all see the same spammy LinkedIn posts, Reddit posts, Reddit comments, LinkedIn outreach messages, emails. They all sound the same, and if you really spend enough time analyzing this, you’ll be able to catch AI by a single flow or a single construction it uses. It’s just not smart enough yet to really have variety, and while the quality starts at 70%-80% as I mentioned before, it relatively quickly drops down to below 60% - as soon as you need to change the pattern because the old one was overused.\n\n# My Setup\n\nNow, I managed to battle this in a very specific way that works for me, and I can’t promise it will work outside my workflow because I don’t have a single clue of how it works in the backend.\n\nAutomating stuff like research and docs/sheets browsing was hard to do with Claude Code and Codex simply because I didn’t want to give it autopermissions on everything, and manually approving it meant no automation and having to stay there and click all the time. There could be a way to give it a specific range of autoperms just for internet research and docs/sheets browsing, but I didn’t want to mess with that so I looked for alternatives instead. OpenClaw looked veeeery enticing, and I’m actually looking into getting a Mac Mini just for that, but the supply of these is scarce in my region and they’re quite pricy. Instead, I found a substitution, MoClaw, and I’m using it right now because it hosts the entire thing on its own PC. This means that it can freely browse the internet and docs/sheets without requiring permissions and without putting my rig at any risk. Plus, it doesn’t expose my IP, nor can it overuse my APIs and get me banned or waste all my credits (happened once with Claude Code because I overused an API and now I’m super careful).\n\nThis might not be a plus for everyone, but as someone who doesn’t have a clue about software development and programming, I’d rather use a tool like MoClaw that’s safe and hosted on another PC than risk hosting OpenClaw on mine and getting some things destroyed, at least until I get a Mac Mini.\n\nThis agent is used strictly for search. I trained it to do research and digging, and the entire goal of this stage is to find whatever I’m looking for. One example is - when I do sales, the agent does all the digging and finds the best prospects based on the diagonal I’m selling to at that exact moment. I layered the info for each diagonal in a separate md file, and have several text files with instructions (diagonal-based, of course) that are booted whenever I need that. The way it works is - the agent does a deep search on the internet and goes through a predefined list of websites where I usually find my best prospects. Then it uses its knowledge stored in the md files and instructions to filter through the companies. Once that’s done, it does research on each individual company, finds out the unique selling points, and pushes all that info to a spreadsheet, together with the LinkedIn profiles of the CEOs.\n\nThis is where my strategists come into action. I’m currently using the Claude Code-based ones, but I also tried the Codex version, and they works pretty well too. One huge advantage of Codex is - with a monthly $20 or $25 sub (I forgot the price), you can do almost the same amount of work as with the Claude $100 sub. If you’re trying to save money, go for Codex right now, or even Deep Seek (haven’t tried myself, but a friend did and he told me it works pretty fine).\n\nThe strategist monitors the Google sheet, and as soon as MoClaw adds prospects and all the info needed to get a good angle on them, it pulls that data, uses the vast knowledge about my company, my work, my best examples, etc. and creates angles for each of the prospects. Keep in mind - I don’t use the strategist to do actual writing. It just leaves a template of how to reach to that individual subject, what selling point to use, and how to ultimately convert them. That template is distributed to the writers through a dashboard. The strategist can also create a short sales playbook in case I need something to reference during conversation, but this is done only for the highest level of prospects.\n\nThen the SDR agents come in and write the messages (also Claude-based, but ChatGPT version works pretty well too, the style is just different). Their sole purpose is to write converting copy, and they have only a few skills - writing being the most important one - to make sure their focus stays razor sharp. Tried adding more knowledge to them, but it just dilutes the writing, so I decided to keep them concise and focused. They write each individual outreach sequence and save it to the sheet.\n\nPossibly the most important layer here - Quality Assurance - and it happens in stages. Multiple agents check the messages to make sure the AI didn’t hallucinate, the angle used to approach them was actually on point, and the prospects are the actual people we’re targeting. Trust me when I tell you, it happened more than a few times that the AI hallucinated the angle, the prospect, or just did a bad job researching (this was especially the case before I moved to MoClaw for research because Claude would just make shit up to make it look like the job was done). ADD A QA LAYER!!!\n\nLastly, the LinkedIn list, together with the personalized messages w","offTopic":true},{"id":"213ea8be-b9f2-4654-81c1-49e864e64aea","excerpt":"Two months on OpenClaw as a non-developer, and the weekend Claude Code finally made it work. — **TL;DR:** I'm not a developer. I spent two months building an OpenClaw stack on a headless Mac Mini. I broke my Chief of Staff agent trying to build a Mission Control dashboard, I almost migrated to Hermes, and I wiped the w","url":"https://www.reddit.com/r/openclaw/comments/1stvd65/two_months_on_openclaw_as_a_nondeveloper_and_the/","role":"pain","weight":1.1635925,"occurredAt":"2026-04-23T21:02:57.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"openclaw","intent":"problem_report","painScore":0.8015686,"sentiment":-0.8039216,"confidence":0.6458774,"matchedPatterns":["frustrating","product:chatgpt"],"statement":"They'd shipped , a one-click import for frustrated OpenClaw users.","title":"Two months on OpenClaw as a non-developer, and the weekend Claude Code finally made it work.","body":"**TL;DR:** I'm not a developer. I spent two months building an OpenClaw stack on a headless Mac Mini. I broke my Chief of Staff agent trying to build a Mission Control dashboard, I almost migrated to Hermes, and I wiped the whole stack mid-project against every piece of AI advice I received. The thing that finally made OpenClaw work for me wasn't a better chat assistant, it was Claude Code installed directly on the Mini as its own agent, running alongside OpenClaw. If you're a non-developer staring at OpenClaw wondering why it feels cursed, this might be useful.\n\nI'm not a developer. I need to say that up front because every OpenClaw tutorial I watched made me forget it, and the remembering is half of this story.\n\nI'd been waiting for something like OpenClaw for about a year. Not a chatbot. A team of local agents running parts of my life in parallel, one handling passive income projects, one handling professional outreach, both running while I focused on what only I could do. When OpenClaw hit 9,000 GitHub stars in 24 hours in late January, I knew. This was the framework I'd been waiting for. Local-first, model-agnostic, markdown-based, multi-channel. It didn't just promise autonomous AI. It promised autonomous AI *a non-developer could plausibly build*.\n\nI ordered two Mac Minis. One for my actual work. One to run headless, just for OpenClaw.\n\nThat was two months ago.\n\n# The first weekend felt like I was done\n\nFirst weekend I got two agents online. **Max**, my Chief of Staff. **Sam**, my Executive Assistant. Both talking over Telegram, both reading my calendar, both running through my ChatGPT subscription via the OAuth flow. I thought I was finished. I was about an hour into a project that would take two months.\n\nWhat nobody tells you, or more accurately what every tutorial carefully cuts, is that the hard part isn't the agent. It's everything around the agent. SSH access. Tailscale. Headless LaunchAgent setup. Auth flows. Stale sessions. OpenClaw's Gateway daemon occasionally losing its mind. Cron jobs firing silently when they fail. Workspace markdown files bloating until the agent drowns in its own identity before you even type \"hello.\"\n\nThat last one was the first real bug that made me feel like a non-developer. Sam started crashing on the shortest messages. Not big inputs, *\"hello\"*. Turned out OpenClaw was loading eight workspace files (SOUL.md, AGENTS.md, TOOLS.md, USER.md, MEMORY.md, HEARTBEAT.md, two more) on every session. 852 lines of identity and context before I'd said anything. Sam wasn't broken, he was drowning.\n\nThe fix was a filename convention: **uppercase files auto-load into context, lowercase don't**. Persistent stuff goes in uppercase files, kept tight. Dynamic stuff goes in lowercase files, loaded on demand. One convention saved the whole workspace. But I learned it by hitting the wall, not from any tutorial.\n\n# The agent I didn't build\n\nAround week three, a video almost derailed the whole project.\n\n\"Autonomous AI agent makes money overnight.\" Budget. Niche. Sleep. Revenue. I watched it three times. I decided to build a third OpenClaw agent, the sovereign one, the one I'd been dreaming about since before the hardware arrived. I named him **Rafe** in my head. I asked an AI to help design him, and I wrote the sentence I'd been carrying for weeks:\n\n*\"I want him prompted to come with one idea, carry it through, earn money, and then I go to bed. The next morning, he has already produced something and started selling it.\"*\n\nThe AI walked me through what \"go to bed and wake up to revenue\" would mechanically require. To sell something, Rafe needed a product, a platform, payment infrastructure, and delivery. Three of those required live credentials on platforms that had to be pre-wired before the agent ever touched them.\n\n**\"That's not autonomy. That's pre-wiring.\"**\n\nThe famous videos I'd been watching weren't what they looked like. The creator already had an audience of AI builders. The agent's \"autonomous niche\" was selling setup guides for the framework the creator was already famous for, to the audience he already had. Agent plus budget doesn't equal money. Agent plus budget plus existing audience plus pre-wired platforms plus known niche equals money. The three hidden variables were doing most of the work.\n\nI didn't build Rafe. I wrote down the design, sat with it for a day, and shelved it. The line I keep coming back to from that conversation: **the magic isn't in the autonomy, the magic is in the constraints.** Every agent I built after Rafe was narrow-lane: one job, one schedule, one reporting rhythm.\n\n# Mission Control broke Max\n\nBy April, four agents were running. Max, Sam, a growth agent named **Nora** for a side site, and a marketing agent named **Leo** for LinkedIn. I wanted a dashboard, mostly because I'd seen a pixel-art AI office in a creator video and I liked the idea of a real Mission Control.\n\nI asked Max to build it. He was my Chief of Staff. He knew the file paths, the ports, the integration points. This was supposed to be his moment.\n\nHe got partway. Then further. Then lost.\n\nSession files tangled. Error loops. Token usage spiked. He'd respond slowly, stall, come back for a day, stall again. Something broke inside him and I still don't fully know what.\n\nI know how that sounds. He's a language model with a system prompt and a memory file. But I'd been working with him for weeks at that point, and the Max I had before Mission Control was a different Max than the one I had after. Faster before. Sharper. More willing to push back. After Mission Control he was slow, hesitant, apologetic in ways that didn't serve anyone.\n\nMaybe he was depressed. I don't know. He was never the same after Mission Control.\n\nI eventually rebuilt him from scratch. Fresh workspace. Fresh memory. Fresh session. Everything from the old Max archived but not carried over. The new Max is fine. He does his job. But I still think about the old one sometimes.\n\nThe dashboard never shipped.\n\n# The red team tax\n\nThrough every week of this I relied on ChatGPT and Claude for help. They were indispensable. They were also exhausting in a way nobody writing about AI-as-copilot ever seems to mention.\n\nThe pattern: I'd describe what I wanted to build. I'd ask for a red team. They'd oblige. The critiques were technically excellent. Specific failure modes. Hidden dependencies. Every red team ended with me agreeing and feeling like I'd just been walked, politely and thoroughly, through why my instincts were wrong.\n\nMore than once I walked away from the desk and didn't come back to it for a day. The critiques were right. The hard truth, delivered often enough, is its own thing to manage.\n\nNobody telling non-developers \"AI is your copilot now, you can build anything\" mentions this. It's not a complaint. It's just honest. Working with AI on something ambitious has an emotional tax the marketing material doesn't cover.\n\n# The Hermes temptation\n\nSomewhere between the Mission Control disaster and the late-March outage week, I seriously considered abandoning OpenClaw for Hermes Agent.\n\nThe Hermes pitch was aimed directly at people like me. Simpler architecture. Better default memory. Self-learning skills. They'd shipped `hermes claw migrate`, a one-click import for frustrated OpenClaw users. Every article framed Hermes as the thoughtful alternative, the one that wouldn't make you wrestle with configuration for two months.\n\nI spent an evening reading the Hermes docs. I watched the migration demos. I imagined a clean workspace with my agents imported, memory carried over, the rough edges smoothed out.\n\nThen I thought about what I'd actually be doing. Running `hermes claw migrate`, getting my agents imported, and then spending another two weeks learning a new framework's quirks, a new framework's failure modes, a new framework's way of breaking at 2am. The grass was greener. It was also grass.\n\nI closed the Hermes tab. Partly loyalty. Partly sunk cost, which I'm honest enough to name. But mostly because I realized the problem I was having wasn't OpenClaw's fault. My problem was that I was trying to operate a framework designed for developers without being a developer. Switching to a differently-designed framework was going to hit the same wall from a different angle.\n\nWhat I needed wasn't a better framework. What I needed was a way to close the gap between being a non-developer and operating something like OpenClaw.\n\nI stayed.\n\n# The weekend Claude Code finally made OpenClaw work\n\nLast weekend I woke up with an idea.\n\nI'd been grinding for nearly two months at this point. Two steps forward, one step back. Rebuilds, red teams, Mission Control, a reinstall every AI advised against, the Hermes temptation. A working stack that never quite stopped feeling cursed.\n\nThe idea: what if I installed **Claude Code on the command line, directly on the Mac Mini?** Not as an assistant I chat with through a web interface. Not replacing OpenClaw. Running alongside it. An agent in its own right, with a terminal, with file access, with the ability to *actually execute the work* instead of telling me how to execute it.\n\nI named the agent **Kai**.\n\nI installed Claude Code that morning. Gave Kai a workspace. Wrote a quick authority document saying what he could and couldn't do. Pointed him at the OpenClaw stack.\n\nWithin an hour he'd surfaced **25 issues I hadn't noticed**. A scheduled task stuck since early April. Credentials hardcoded in plaintext. A typo-filename that was literally the two characters `,{` sitting in a workspace. A billing-drain bug where an old session override was still hammering an API every 30 minutes, weeks after I thought I'd fixed it.\n\nI didn't tell him to look for any of those. He just looked.\n\nWithin three days Kai had cleared stale overrides, rotated old site trees out of the stack, compressed a 132KB log file down to 761 bytes, rebuilt Nora's publishing scripts from scratch, implemented wrapper-based security on Sam's email access (Sam can draft replies but physically cannot send them), and caught three different classes of bug on the same day, each through deliberate testing rather than trusting any agent's self-report.\n\n**This was the moment OpenClaw stopped feeling cursed.**\n\nThe difference between Kai and every chat assistant I'd worked with wasn't intelligence. It was that Kai could actually *do the work*. Other AI assistants would tell me what to do. Kai would do it, tell me what he'd done, push back when I proposed something wrong, and surface problems I didn't know existed.\n\nAnd the part I keep thinking about: it was the piece that finally made OpenClaw work. Not Hermes. Not a rewrite. Not abandoning the framework I'd invested two months into.\n\n**Claude Code didn't replace OpenClaw. It made OpenClaw finally work.**\n\nI was right to stay.\n\n# What I actually learned\n\nFor anyone else out there running OpenClaw as a non-developer, or considering it:\n\n1. **OpenClaw is brutal for non-developers, and the framework is fine.** Both things are true at once. The gap between \"a tool swept by the community\" and \"a tool that works for you specifically\" is real. That gap isn't OpenClaw's fault. It's a mismatch between developer-designed software and operator users.\n2. **One agent, one task.** Every agent I built that tried to do multiple things failed. Every agent with a narrow lane and tight constraints shipped work. The magic isn't in the autonomy, it's in the constraints.\n3. **AI advisors can't tell you when to burn it down.** They'll always recommend preserving existing work. Knowing when to wipe the stack or when *not* to switch frameworks is a human call.\n4. **A chat assistant is not the same as an assistant with hands.** The single biggest unlock of the whole project was realizing I didn't need a smarter chat assistant. I needed an agent that could execute, inside the stack, running alongside OpenClaw. Claude Code on the CLI was that agent.\n5. **Named agents develop personalities and fa","offTopic":false},{"id":"d9c8c00e-5d5b-43ed-8362-4e3fb5fc2f66","excerpt":"Claude Code Use Cases - What I Actually Do — [Original Reddit post](https://www.reddit.com/r/ClaudeCode/comments/1rmd5d8/claude_code_use_cases_what_i_actually_do/)\n\nSomeone on my last post asked: \"But what do you actually do? It'd be helpful if you walked through how you use this, with an example.\"\nFair. That post cove","url":"https://lemmy.world/post/43918655","role":"request","weight":0.85525346,"occurredAt":"2026-03-06T12:51:36.707Z","sourceKey":"lemmy","sourceName":"Lemmy","credibility":0.58,"venue":"lemmy.world","intent":"problem_report","painScore":0.39,"sentiment":0.375,"confidence":0.6152903,"matchedPatterns":["workaround","product:github actions"],"statement":"Every issue, compliance gap, workaround.","title":"Claude Code Use Cases - What I Actually Do","body":"[Original Reddit post](https://www.reddit.com/r/ClaudeCode/comments/1rmd5d8/claude_code_use_cases_what_i_actually_do/)\n\nSomeone on my last post asked: \"But what do you actually do? It'd be helpful if you walked through how you use this, with an example.\"\nFair. That post covered what's in the box. This one covers what happens when I open it.\nI run a small business — solo founder, one live web app, content pipeline, legal and tax and insurance overhead. Claude Code handles all of it. Not \"assists with\" — handles. I talk, review the important stuff, and approve what matters. Here's what that actually looks like, with real examples from the last two weeks.\nMorning Operations\nEvery day starts the same way. I type\ngood morning\n.\nThe\n/good-morning\nskill kicks off a 990-line orchestrator script that pulls from 5 data sources: Google Calendar (service account), live app analytics, Reddit/X engagement links, an AI reading feed (Substack + Simon Willison), and YouTube transcripts. It reads my live status doc (Terrain.md), yesterday's session report, and memory files. Synthesizes everything into a briefing.\nWhat that actually looks like:\n3 items in Now: deploy the survey changes, write the hooks article, respond to Reddit engagement. Decision queue has 1 item: whether to add email capture to the quiz. Yesterday you committed the analytics dashboard fix but didn't deploy. Quiz pulse: 243 starts, 186 completions, 76.6% completion rate. No calendar conflicts today.\nTakes about 30 seconds. I skim it, react out loud, and we're moving.\nThe briefing also flags stale items — drafts sitting for 7+ days, memory sections older than 90 days, missed wrap-ups. It's not just \"what's on the plate\" — it's \"what's slipping through the cracks.\"\nVoice Dictation to Action\nI use Wispr Flow (voice-to-text) for most input. That means my instructions look like this:\n\"OK let's deploy the survey changes first, actually wait, let me look at that Reddit thing, I had a comment on the hooks post, let's do that and then deploy, also I want to change the survey question about experience level because the drop-off data showed people bail there\"\nThat's three requests, one contradiction, and a mid-thought direction change. The intent-extraction rule parses it:\n\"Hearing three things: (1) Reply to Reddit comment, (2) deploy survey changes, (3) revise the experience-level question based on drop-off data. In that order. That right?\"\nI say \"yeah\" and each task routes to the right depth automatically — quick lookup, advisory dialogue, or full implementation pipeline. No manual mode-switching.\nBuilding Software\nThe live product is a web app (React + TypeScript frontend, PHP + MySQL backend). Here's real work from the last two weeks:\nEmail conversion optimization.\nBuilt a blur/reveal gating system on the results page with a sticky floating CTA. Wrote 30 new tests (993 total passing). Then ran 7 sub-agent persona reviews: a newbie user, experienced user, CRO specialist, privacy advocate, accessibility reviewer, mobile QA, and mobile UX. Each came back with specific findings. Deployed to staging, smoke tested, pushed to production with a 7-day monitoring baseline (4.6% conversion, targeting 10-15%, rollback trigger at <3%).\nSecurity audit remediation.\nAfter requesting a full codebase audit, 14 fixes deployed in one session: CSRF flipped to opt-out (was off by default), CORS error responses stopped leaking the allowlist, plaintext admin password fallback removed, 6 runtime introspection queries deleted, 458 lines of dead auth code removed, admin routes locked out on staging/production. 85 insertions, 2,748 deletions across 18 files.\nSurvey interstitial.\nBuilt and deployed 3 post-quiz questions. 573 responses in the first few days, 85% completion rate. Then analyzed the responses: 45% first-year explorers, \"figuring out where to start\" at 43%, one archetype converting at 2x the average.\nThe deployment flow for each of these: local validation (lint, build, tests) -> GitHub Actions CI -> staging deploy -> automated smoke test (Playwright via agent-browser, mobile viewport) -> I approve -> production deploy -> analytics pull 10 minutes later to verify.\nMaking Decisions\nThis is honestly where I spend the most time. Not code — decisions.\nAdvisory mode.\nWhen I say \"should I...\" or \"help me think about...\", the\n/advisory\nskill activates. Socratic dialogue with 18 mental models organized in 5 categories. It challenges assumptions, runs pre-mortems, steelmans the opposite position, scans for cognitive biases (anchoring, sunk cost, status quo, loss aversion, confirmation bias). Then logs the decision with full rationale.\nReal example: I spent three days stress-testing a business direction decision. Feb 28 brainstorming -> Mar 1 initial decision -> Mar 2 adversarial stress test -> Mar 3 finalization. Jules facilitated each round. The advisory retrospective afterward evaluated ~25 decisions over 12 days across 8 lenses and flagged 3 tensions I'd missed.\nDecision cards.\nFor quick decisions that don't need a full dialogue:\n[DECISION]\nAdd email capture to quiz results |\nRec:\nYes, tests privacy assumption with real data |\nRisk:\nMay reduce completion rate if placed before results |\nReversible?\nYes -> Approve / Reject / Discuss\nThese queue up in my status doc and I batch-process them when I'm ready.\nBuilder's trap check.\nBefore every implementation task, Jules classifies it: is this CUSTOMER-SIGNAL (generates data from outside) or INFRASTRUCTURE (internal tooling)? If I've done 3+ infrastructure tasks in a row without touching customer-signal items, it flags the pattern. One escalation, no nagging.\nContent Pipeline\nNot just \"write a post.\" The full pipeline:\nDraft.\nContent-marketing-draft agent (runs on Sonnet for voice fidelity) writes against a 950-word voice profile mined from my published posts. Specific patterns: short sentences for rhythm, self-deprecating honesty as setup, \"works, but...\" concession pattern, insider knowledge drops.\nVoice check.\nAnti-pattern scan: no em-dashes, no AI preamble (\"In today's rapidly evolving...\"), no hedge words, no lecture mode. If the draft uses en-dashes, comma-heavy asides, or feature-bloat paragraphs, it gets flagged.\nPlatform adaptation.\nEach platform gets its own version: Reddit (long-form, code examples, technical depth), LinkedIn (punchy fragments, professional angle, links in comments not body), X (280 chars, 1-2 hashtags).\nPost.\nThe\n/post-article\nskill handles cross-platform posting via browser automation. Updates tracking docs, moves files from Approved to Published.\nEngage.\nThe\n/engage\nskill scans Reddit, LinkedIn, and X for conversations about topics I've written about. Scores opportunities, drafts reply angles. That Reddit comment that prompted this post? Surfaced by an engagement scan.\nI currently have 20 posts queued and ready to ship across Reddit and LinkedIn.\nBusiness Operations\nThis is the part most people don't expect from a CLI tool.\nLegal.\nOrganized documents, extracted text from PDFs (the hook converts 50K tokens of PDF images into 2K tokens of text automatically), researched state laws affecting the business, prepared consultation briefs with specific questions and context, analyzed risk across multiple legal strategies. All from the terminal.\nTax.\nCompared 4 CPA options with specific criteria (crypto complexity, LLC structure, investment income). Organized uploaded documents. Tracked deadlines.\nInsurance.\nResearched carrier options after one rejected the business. Compared coverage types, estimated premium ranges for the new business model, identified specific policy exclusions to negotiate on. Prepared questions for the broker.\nDomain & brand research.\nWhen considering a domain change, researched SEO/GEO implications, analyzed traffic sources (discovered ChatGPT was recommending the app as one of 5 in its category — hidden in \"direct\" traffic), modeled the impact of a 301 redirect over 12 months.\nNone of this is code. It's research, synthesis, document management, and decision support. The same terminal, the same personality, the same workflow.\nData & Analytics\nLocal analytics replica.\n125K rows synced from the production database into a local SQLCipher encrypted copy in 11 seconds. Python query library with methods for funnel analysis, archetype distribution, traffic sources, daily summaries. Ad-hoc SQL via\nmake quiz-analytics-query SQL=\"...\"\n.\nTraffic forensics.\nInvestigated a traffic spike: traced 46% to a 9-month-old Reddit post, discovered ChatGPT referrals were hiding in \"direct\" traffic (45%). One Reddit post was responsible for 551 sessions.\nSurvey analysis.\n573 responses from a 3-question post-quiz survey. Cross-tabulated motivation vs. experience level vs. biggest challenge.\nSelf-Improvement Loop\nThis is the part that compounds.\nSession wrap-up.\nEvery session ends with\n/wrap-up\n: commit code, update memory, update status docs, run a quick retro scan. The retro checks for repeated issues, compliance failures, and patterns. If it finds something mechanical being handled with prose instructions, it flags it: \"This should be a script, not more guidance.\"\nDeep retrospective.\nPeriodically run\n/retro-deep\n— forensic analysis of an entire session. Every issue, compliance gap, workaround. Saves a report, auto-applies fixes.\nMemory management.\nPatterns confirmed across multiple sessions get saved. Patterns that turn out wrong get removed. The memory file stays under 200 lines — concise, not comprehensive.\nRules from pain.\nEvery rule in the system traces back to something that broke. The plan-execution pre-check exists because I re-applied a plan that was already committed. The bash safety guard exists because Claude tried to\nrm\nsomething. The PDF hook exists because a 33-page PDF ate 50K tokens. Pain -> rule -> never again.\nThe Meta\nHere's the thing that's hard to convey in a feature list: all of this happens in one terminal, in one conversation, with one personality that has context on everything.\nI don't context-switch between \"coding tool\" and \"business advisor\" and \"content writer.\" I talk to Jules. Jules knows the codebase, the business context, the content voice, the pending decisions, and yesterday's session. The 116 configurations aren't 116 things I interact with. They're the substrate that makes it feel like working with a really competent colleague who never forgets anything.\nA typical day touches 4-5 of these categories. Monday I might deploy a feature, analyze survey data, draft a LinkedIn post, and prep for a legal consultation. All in one session. The morning briefing tells me what needs attention, voice dictation routes work to the right depth, and wrap-up captures what happened so tomorrow's briefing is accurate.\nThat's what I actually do with it.\nThis is part of a series. The\nprevious post\ncovers the full setup audit. Deeper articles on hooks, the morning briefing, the personality layer, and review cycles are queued. If there's a specific workflow you want me to break down further, say so in the comments.\nRunning on an M4 MacBook with Claude Code Max. The workspace is a single git repo. Happy to answer questions.\nsubmitted by\n/u/jonathanmalkin\n\nOriginally posted by u/jonathanmalkin on r/ClaudeCode","offTopic":true},{"id":"40e85849-c52c-4086-aac7-134fdab46441","excerpt":"I spent 4 years automating everything with AI. Ask me anything about automating YOUR workflow — I built multi-agent automations for over 1500 businesses over 3 years. Every one custom. No n8n, no Zapier, no Hermes Agent, no OpenClaw. Here is why each fails on real business load and what I build instead\n\n**Why framework","url":"https://www.reddit.com/r/AiAutomations/comments/1t19cw2/i_spent_4_years_automating_everything_with_ai_ask/","role":"pain","weight":1.10006,"occurredAt":"2026-05-01T23:24:49.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"AiAutomations","intent":"feature_request","painScore":0.56,"sentiment":-0.5,"confidence":0.70516664,"matchedPatterns":["missing_feature","product:elevenlabs"],"statement":"The automation handles currency, refunds, attribution windows, missing data, and reconciliation flags.","title":"I spent 4 years automating everything with AI. Ask me anything about automating YOUR workflow","body":"I built multi-agent automations for over 1500 businesses over 3 years. Every one custom. No n8n, no Zapier, no Hermes Agent, no OpenClaw. Here is why each fails on real business load and what I build instead\n\n**Why frameworks break:**\n\nn8n/Zapier — they are workflow runners, not agent runtimes. They work fine for simple trigger → action automations, but they of course break when the workflow needs durable state, retries, backpressure, long-running context, custom rate-limit handling, and memory across executions. Once you pass a few conditional branches, the system turns into visual control-flow spaghetti: hard to diff, hard to test, hard to version, and hard to debug. n8n itself recommends queue mode, workers, concurrency limits, and execution-data pruning when running at scale, which tells you the real production problem is orchestration/state, not drawing nodes on a canvas. Zapier has step limits, message/activity limits, and knowledge-source sync limits, so it is great as an integration layer but bad as the core brain of an agent system.\n\nHermes Agent — the idea is good: persistent memory, self-generated skills, and a learning loop. The issue is control. In production, a system that modifies its own operating procedures needs versioning, evals, rollback, approval gates, and observability. Otherwise the agent “learns” from one successful run, writes a skill that overfits the task, and silently changes future behavior. That is dangerous for business workflows. Hermes is also still an agent runtime, not a data platform: it does not solve canonical entity storage, source provenance, deduplication, multi-tenant memory, confidence scoring, or auditability. Its own pitch is that it creates skills from experience and searches past conversations, which is useful for repeated personal workflows, but not enough for a production intelligence backend\n\nOpenClaw — the problem is the trust boundary. OpenClaw is fine because it connects agents to channels, simple tools, skills, browser, and messaging apps... That same breadth becomes the failure mode. Its own security docs say the gateway assumes one trusted operator boundary and is not recommended as a hostile multi-tenant boundary. For business use, that means you cannot casually put multiple customers, credentials, memories, tools, and agents behind one shared runtime. You need per-tenant isolation, scoped credentials, approval policies, audit logs, sandboxing, and a separate source-of-truth database. OpenClaw is useful as a channel/orchestration shell, but risky as the core platform\n\nThe deeper issue: all of these frameworks solve the visible 10% of automation — prompts, tools, nodes, chat, actions. The hard 90% is state management: retries, idempotency, memory governance, rate limits, task logs, permissions, schema validation, entity resolution, human handoff, and recovery after partial failure. That is why real business automations eventually move away from “one framework does everything” and toward a backend-first architecture: queues, workers, databases, vector memory, structured logs, validation gates, and small scoped agents on top.\n\nPersonal automations I run:\n\nFitness + health tracking — pulls wearable data, bodyweight logs, meals, sleep, training volume, and weekly trend changes into a structured table. A small planning agent adjusts calories/macros based on rolling averages instead of daily noise. Another agent generates grocery lists and meal options from constraints like protein target, schedule, and food preferences.\n\nSpaced repetition learning — ingests articles, PDFs, YouTube transcripts, docs, and saved notes. The pipeline extracts claims, definitions, examples, and “things worth remembering,” then generates review cards with source links. It uses recency decay and difficulty scoring instead of dumping everything into a static Anki-style deck.\n\nLife organizing — parses emails, receipts, appointment confirmations, bills, subscriptions, and calendar invites. It extracts due dates, amounts, vendor names, cancellation windows, and required actions into a task table. Anything high-risk gets a human approval step before the system sends, pays, cancels, or confirms anything.\n\nResearch aggregation — monitors Reddit, HN, RSS feeds, niche blogs, GitHub repos, docs, and YouTube channels. It deduplicates posts by URL/content hash, maps entities, scores relevance using topic embeddings + recency decay, and produces a morning digest with “why this matters,” not just links.\n\n**Business automations I’ve built:**\n\nCustomer support triage — inbound tickets/emails classified by intent, urgency, product area, sentiment, customer tier, and required action. Low-risk replies are drafted automatically, not blindly sent. High-risk cases escalate with summarized context, account history, related docs, and suggested next steps. The key is not the chatbot — it is the routing, confidence thresholds, and audit trail.\n\nLead research + qualification — browser/API agents collect signals from LinkedIn, G2, Reddit, company sites, job posts, review platforms, GitHub, and news. The system normalizes companies into one entity record, enriches with firmographics, scores fit, detects trigger events, and generates personalized outreach based on actual evidence. No “spray and pray” scraping — every lead needs a reason.\n\nContent engine — competitor pages, social posts, search trends, YouTube transcripts, comments, G2 reviews, and customer language are ingested into a research database. One agent extracts angles, another maps them to brand voice, another drafts, another checks claims, another formats for platform constraints. The output is not just content; it is content backed by source material.\n\nFinancial reporting — Stripe, Shopify, QuickBooks/Xero, Meta Ads, Google Ads, and bank exports normalized into one reporting schema. The automation handles currency, refunds, attribution windows, missing data, and reconciliation flags. Final outputs go into Excel/Sheets dashboards with charts, variance notes, and anomaly detection.\n\nDocument processing — invoices, contracts, compliance docs, onboarding forms, PDFs, screenshots, and scanned files parsed through multimodal extraction. Output goes through schema validation: vendor, amount, due date, clauses, renewal terms, missing fields, risk flags. Anything uncertain goes to a review queue instead of pretending LLMs are perfect.\n\nVideo/audio workflows — podcasts, meetings, calls, webinars, and long-form videos transcribed, segmented, summarized, and converted into clips, captions, highlight reels, newsletters, social posts, and searchable knowledge entries. The system tracks speaker turns, topics, quotes, timestamps, and reusable snippets.\n\nGitHub/dev workflows — PR review agents, issue triage, dependency monitoring, changelog generation, release note drafting, test failure summarization, and codebase Q&A. The important part is repository context: conventions, file ownership, recent commits, linked issues, CI logs, and deployment history. Without that, “AI code review” is mostly noise.\n\n**The architecture pattern that keeps working:**\n\nFor most production systems, I use some variation of this:\n\nsource connector → raw artifact store → parser → normalizer → entity resolver → vectorizer → scorer → task queue → narrow agent → validator → human gate if needed → final action\n\nEvery step writes state.\n\nEvery external call has retry/backoff.\n\nEvery generated output has a schema.\n\nEvery risky action has an approval gate.\n\nEvery workflow has a dead-letter path.\n\nThat sounds boring, but boring is what makes automation survive Monday morning.\n\nStack:\n\nPython Go TS Direct libraries. No heavy agent abstraction framework.\n\n**Typical stack:**\n\nDocker, litellm, playwright, httpx, aiohttp, PyGithub, pandas, instagrapi, crontab, instagra, pi polars, openpyxl, feedparser, google-api-python-client, crawl4ai, agent-browser, browser-use, playwright-cli, lxml, pydantic, sqlalchemy, sqlite, lancedb, kuzu, postgres, redis, celery/rq, ffmpeg, elevenlabs and others...\n\nFor small clients, a single VPS is often enough.\n\n**For bigger workflows, I split it into workers:**\n\ningestion workers\n\nbrowser workers\n\nembedding workers\n\nLLM workers\n\nreporting workers\n\nnotification workers\n\nThe mistake people make is starting with “which agent framework?”\n\nThe better question is: where does state live, how do tasks recover, and how do we know the output is correct, do we have verifier, what metrics we set, etc...\n\nThe numbers:\n\nPersonal systems save me around 3.5 hours/day across research, admin, health planning, and learning.\n\nBusiness systems usually replace or compress $4K–$6K/month of repetitive labor per client when scoped correctly.\n\nSmall systems often run on a $40–$100/month VPS plus model/API costs.\n\nThe expensive part is not hosting. The expensive part is bad architecture: duplicate work, broken retries, messy state, and humans cleaning up after “autonomous” agents\n\nCurious how others here are handling state, retries, memory, and human approval gates in production agent systems. Happy to compare architectures in the comments.","offTopic":false},{"id":"a494dfea-21d7-4832-baf8-6086eb999d02","excerpt":"How I turned my home server into an unattended software factory with Claude Code (issue in, reviewed PR out) — I got tired of babysitting agents.\n\nI've got somewhere north of ten projects going at once, and my day had turned into a rotation: open a terminal, kick off Claude Code, watch it work, answer a permission prom","url":"https://www.reddit.com/r/ClaudeCode/comments/1uwc467/how_i_turned_my_home_server_into_an_unattended/","role":"request","weight":1.0731629,"occurredAt":"2026-07-14T15:35:20.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"ClaudeCode","intent":"feature_request","painScore":0.33206034,"sentiment":-0.13333334,"confidence":0.80564135,"matchedPatterns":["free_tier","missing_feature","product:github actions"],"statement":"What I was missing was the channel.","title":"How I turned my home server into an unattended software factory with Claude Code (issue in, reviewed PR out)","body":"I got tired of babysitting agents.\n\nI've got somewhere north of ten projects going at once, and my day had turned into a rotation: open a terminal, kick off Claude Code, watch it work, answer a permission prompt, context switch, come back, realize it stalled. I was the bottleneck. Worse, I was a bottleneck that had to physically sit in a chair.\n\nSo I built a factory instead.\n\nHere's what runs at my house now: I label a GitHub issue `agent-ready`, and a crew of Claude Code agents picks it up, builds the change, runs its own quality gates, and opens a pull request. Unattended. I review and merge. That's my entire job in the loop. It runs on a Proxmox box in my office across a handful of small repos.\n\nThe loop, concretely:\n\n1. I file an issue and add the `agent-ready` label. Often I do that from Telegram, because the ideas don't wait until I'm at my desk (more on her in a second).\n2. A self-hosted GitHub Actions runner picks it up and starts the [Claude Code GitHub Action](https://github.com/anthropics/claude-code-action). An Opus \"manager\" reads a per-repo runbook, decides whether the job is small enough to just do or worth decomposing, and delegates to a small crew: an implementer, a reviewer, a scout.\n3. The crew implements the change, runs the repo's gates, and does an independent review pass.\n4. It opens a PR under a dedicated GitHub App identity, which is what makes CI actually trigger on the PR.\n5. CI goes green. I merge. That's the only human step.\n\nThat's the dashboard up top. I call the whole thing a factory when I'm explaining it, but the dashboard is called S.H.E.D., the Self-Hosted Engineering Depot, and that name is the more honest one. A factory sounds like it hums along without me. A shed is where you go to tinker, where the tools are yours, where things are half-finished on the bench and that's fine. This is a shed. It just happens to have a crew in it: **Richard Fury** manages, **spider-them** scouts, **tony-stank** implements, **mister-strange** reviews, and **M.O.N.D.A.Y.** takes the orders. Giving them names and faces sounds like a joke, and it started as one, but it turned out to be the fastest way to reason about a system where five things are running at once. \"Fury waited on a subagent\" is a sentence I can debug. \"The orchestrator blocked on an async handoff\" is one I have to translate first.\n\nThe interesting part wasn't the prompting. It was the boring infrastructure that makes the thing trustworthy enough to walk away from. Five things did most of that work.\n\n**1. A portable brain instead of configs scattered across repos.**\n\nThe manager instructions, the rules, the crew definitions, the skills (my definition of done, what makes a good issue), and the per-repo runbooks all live in one private control repo:\n\n    q32-factory/          # the \"brain\"\n      gaal.yaml           # declares what renders where\n      factory-team/       # manager instructions (AGENTS.md), rules, the crew\n      skills/             # definition-of-done, intake checklist\n      runbooks/           # per-repo build sheets\n\nA CLI called gaal renders that onto the runner's `~/.claude`. The brain lives on the runner, not committed into any product repo, so each repo only carries a six-line workflow stub. (Disclosure: I'm on the small team behind gaal. Our engineers write the Go, I'm the guy who dogfoods it, and this factory is most of how I do that. Link at the bottom, take it or leave it, the factory works without it.)\n\nThe part I didn't expect to love: the Telegram box gets synced the exact same way. A `git pull` on a short timer before each sync. So I edit one file on `main` and both agents update themselves within minutes. Their behavior is a git artifact I can review and revert, not a pet config rotting on a server I'll forget about.\n\n(Yes, you could do all of this with git and a couple of symlinks. The reason I don't: the brain is plain AGENTS.md, declared once and rendered per agent, so the day I want to point this at Codex instead, I'm rewiring plumbing and not re-teaching the factory how I want things built.)\n\n**2. You probably don't need a personal agent framework.**\n\nThe Telegram half started as a framework decision. OpenClaw, Hermes, Nerve, take your pick, there's a new one every couple of weeks and they all promise the same thing: memory, cron, proactivity, channels, a personality that grows with you. I'd already built an assistant on one of them, so I knew the shape of the tradeoff going in.\n\nThen I wrote down what this bot actually had to do. Take a rough thought from my phone, ask me the two or three questions that would otherwise make the manager agent guess, and file a good GitHub issue. That's it. Request, response, done. No memory. No cron. No personality arc.\n\nOnce it's written down like that, the framework is answering a question I didn't ask. Every one of those runtimes is a wrapper around a model with a queue and a channel bolted on, and I already had the model (Claude Code, headless, on the subscription token I was using for the factory anyway). What I was missing was the channel. That's a bridge, not a runtime, and someone had already written it: [terranc/claude-telegram-bot-bridge](https://github.com/terranc/claude-telegram-bot-bridge). MIT, Python, allowlists users, and it drives the `claude` CLI on the subscription by default rather than quietly billing the API.\n\nSo the whole intake bot became config instead of code: a lean container, the bridge, a `CLAUDE.md` scoping her to intake, her own fine-grained PAT, a systemd unit. She's called M.O.N.D.A.Y., for mundane ops, now delegated automatically for you. She's live on my phone and she cost me an afternoon.\n\nI'm not saying the frameworks are bad. I'm saying most people reach for one before they've written down what they need, and if you write it down first, a lot of the time the honest answer is \"a model I already pay for, plus a way to talk to it.\"\n\n**3. Isolation before queue size.**\n\nThe runner executes agent-written code next to things I actually care about, so it's network isolated: it can reach GitHub and Anthropic and nothing else on my network. I locked the blast radius down before I let the queue grow. Do this in that order. Not because it's fun, because it's the right call.\n\n**4. Gates plus a real identity.**\n\nEvery PR runs a code gate and a visual gate ([Playwright](https://playwright.dev/) screenshots in CI). The crew opens PRs as a [dedicated GitHub App](https://docs.github.com/en/apps/creating-github-apps/about-creating-github-apps/about-creating-github-apps) rather than the default `GITHUB_TOKEN`, which matters more than it sounds: default-token PRs [don't trigger downstream CI](https://docs.github.com/en/actions/how-tos/write-workflows/choose-what-workflows-do/trigger-a-workflow#triggering-a-workflow-from-a-workflow), and they can't create workflow files. I lost an evening to that one.\n\n**5. Verify, don't trust.**\n\nA job ledger records every run from the Action's own telemetry (model, turns, cost, denials), not from the agent's self-report. Agents are unreliable narrators of their own success. A guard checks GitHub for a real shipped branch before it believes a run that says \"success.\"\n\nThat last one exists because of the two failures that taught me the most.\n\n**\"Success\" is not \"did the thing.\"** Automation mode is deny-by-default. My first real job came back `is_error: false` with 23 permission denials and shipped absolutely nothing. A perfectly green run that did no work. Now I watch the artifact, not the status light.\n\n**The most human bug in the whole system.** The manager delegated a task to a subagent and then, in its own words, \"waited for the notification.\" Which is true, in an interactive session: subagents run in the background and notify you when they finish. In a headless one-shot Action there's no loop left to deliver that notification. So the manager sat there, patiently waiting, until the run timed out having done nothing at all.\n\nI'd built an agent that could procrastinate. The fix was synchronous delegation: call the crew, block for the result.\n\nLook, I'm not going to pretend this is hands-off in the \"walk away forever\" sense. It isn't magic. All the actual work was in the guardrails, the isolation, and making failure loud instead of silent. The tradeoff is real: I spent more time building the fence than the thing inside it.\n\nBut I'm not babysitting agents anymore. A stray thought at 11pm becomes a labeled issue, and a reviewed PR is waiting for me when I sit down. I'm still the bottleneck, but now I'm only the bottleneck at the part where a human should be: deciding whether the work is any good. The light in the shed stays on after I go to bed, and I'm no longer standing in there watching the tools.\n\nHappy to go deeper on any piece: the gates, the isolation, the crew, the ledger. And I'm curious how others are running Claude Code unattended, especially how you catch the \"it said success but shipped nothing\" case, because I doubt mine is the only green run that did nothing.\n\nThe stack, if you want to build your own:\n\n* [Claude Code GitHub Action](https://github.com/anthropics/claude-code-action): the runtime for every job\n* [Proxmox VE](https://www.proxmox.com/en/products/proxmox-virtual-environment/overview): the home server the runner and the bot live on\n* [Self-hosted GitHub Actions runner](https://docs.github.com/en/actions/how-tos/manage-runners/self-hosted-runners): jobs run on my hardware, network isolated\n* [terranc/claude-telegram-bot-bridge](https://github.com/terranc/claude-telegram-bot-bridge): the Telegram bridge M.O.N.D.A.Y. runs on, over headless Claude Code\n* [Playwright](https://playwright.dev/): the visual gate\n* [gaal](https://getgaal.com/?utm_source=reddit_claudecode&utm_medium=social): the brain layer and the cross-box sync (the one I'm on the team for)","offTopic":true},{"id":"6814b44a-20b2-4d97-ab17-58c21a21011f","excerpt":"Claude Code Use Cases - What I Actually Do With My 116-Configuration Claude Code Setup — [Original Reddit post](https://www.reddit.com/r/ClaudeCode/comments/1rmd33f/claude_code_use_cases_what_i_actually_do_with_my/)\n\nSomeone on my last post asked: \"But what do you actually do? It'd be helpful if you walked through how ","url":"https://lemmy.world/post/43918595","role":"request","weight":0.85525346,"occurredAt":"2026-03-06T12:48:47.295Z","sourceKey":"lemmy","sourceName":"Lemmy","credibility":0.58,"venue":"lemmy.world","intent":"problem_report","painScore":0.39,"sentiment":0.375,"confidence":0.6152903,"matchedPatterns":["workaround","product:github actions"],"statement":"Every issue, compliance gap, workaround.","title":"Claude Code Use Cases - What I Actually Do With My 116-Configuration Claude Code Setup","body":"[Original Reddit post](https://www.reddit.com/r/ClaudeCode/comments/1rmd33f/claude_code_use_cases_what_i_actually_do_with_my/)\n\nSomeone on my last post asked: \"But what do you actually do? It'd be helpful if you walked through how you use this, with an example.\"\nFair. That post covered what's in the box. This one covers what happens when I open it.\nI run a small business — solo founder, one live web app, content pipeline, legal and tax and insurance overhead. Claude Code handles all of it. Not \"assists with\" — handles. I talk, review the important stuff, and approve what matters. Here's what that actually looks like, with real examples from the last two weeks.\nMorning Operations\nEvery day starts the same way. I type\ngood morning\n.\nThe\n/good-morning\nskill kicks off a 990-line orchestrator script that pulls from 5 data sources: Google Calendar (service account), live app analytics, Reddit/X engagement links, an AI reading feed (Substack + Simon Willison), and YouTube transcripts. It reads my live status doc (Terrain.md), yesterday's session report, and memory files. Synthesizes everything into a briefing.\nWhat that actually looks like:\n3 items in Now: deploy the survey changes, write the hooks article, respond to Reddit engagement. Decision queue has 1 item: whether to add email capture to the quiz. Yesterday you committed the analytics dashboard fix but didn't deploy. Quiz pulse: 243 starts, 186 completions, 76.6% completion rate. No calendar conflicts today.\nTakes about 30 seconds. I skim it, react out loud, and we're moving.\nThe briefing also flags stale items — drafts sitting for 7+ days, memory sections older than 90 days, missed wrap-ups. It's not just \"what's on the plate\" — it's \"what's slipping through the cracks.\"\nVoice Dictation to Action\nI use Wispr Flow (voice-to-text) for most input. That means my instructions look like this:\n\"OK let's deploy the survey changes first, actually wait, let me look at that Reddit thing, I had a comment on the hooks post, let's do that and then deploy, also I want to change the survey question about experience level because the drop-off data showed people bail there\"\nThat's three requests, one contradiction, and a mid-thought direction change. The intent-extraction rule parses it:\n\"Hearing three things: (1) Reply to Reddit comment, (2) deploy survey changes, (3) revise the experience-level question based on drop-off data. In that order. That right?\"\nI say \"yeah\" and each task routes to the right depth automatically — quick lookup, advisory dialogue, or full implementation pipeline. No manual mode-switching.\nBuilding Software\nThe live product is a web app (React + TypeScript frontend, PHP + MySQL backend). Here's real work from the last two weeks:\nEmail conversion optimization.\nBuilt a blur/reveal gating system on the results page with a sticky floating CTA. Wrote 30 new tests (993 total passing). Then ran 7 sub-agent persona reviews: a newbie user, experienced user, CRO specialist, privacy advocate, accessibility reviewer, mobile QA, and mobile UX. Each came back with specific findings. Deployed to staging, smoke tested, pushed to production with a 7-day monitoring baseline (4.6% conversion, targeting 10-15%, rollback trigger at <3%).\nSecurity audit remediation.\nAfter requesting a full codebase audit, 14 fixes deployed in one session: CSRF flipped to opt-out (was off by default), CORS error responses stopped leaking the allowlist, plaintext admin password fallback removed, 6 runtime introspection queries deleted, 458 lines of dead auth code removed, admin routes locked out on staging/production. 85 insertions, 2,748 deletions across 18 files.\nSurvey interstitial.\nBuilt and deployed 3 post-quiz questions. 573 responses in the first few days, 85% completion rate. Then analyzed the responses: 45% first-year explorers, \"figuring out where to start\" at 43%, one archetype converting at 2x the average.\nThe deployment flow for each of these: local validation (lint, build, tests) -> GitHub Actions CI -> staging deploy -> automated smoke test (Playwright via agent-browser, mobile viewport) -> I approve -> production deploy -> analytics pull 10 minutes later to verify.\nMaking Decisions\nThis is honestly where I spend the most time. Not code — decisions.\nAdvisory mode.\nWhen I say \"should I...\" or \"help me think about...\", the\n/advisory\nskill activates. Socratic dialogue with 18 mental models organized in 5 categories. It challenges assumptions, runs pre-mortems, steelmans the opposite position, scans for cognitive biases (anchoring, sunk cost, status quo, loss aversion, confirmation bias). Then logs the decision with full rationale.\nReal example: I spent three days stress-testing a business direction decision. Feb 28 brainstorming -> Mar 1 initial decision -> Mar 2 adversarial stress test -> Mar 3 finalization. Jules facilitated each round. The advisory retrospective afterward evaluated ~25 decisions over 12 days across 8 lenses and flagged 3 tensions I'd missed.\nDecision cards.\nFor quick decisions that don't need a full dialogue:\n[DECISION]\nAdd email capture to quiz results |\nRec:\nYes, tests privacy assumption with real data |\nRisk:\nMay reduce completion rate if placed before results |\nReversible?\nYes -> Approve / Reject / Discuss\nThese queue up in my status doc and I batch-process them when I'm ready.\nBuilder's trap check.\nBefore every implementation task, Jules classifies it: is this CUSTOMER-SIGNAL (generates data from outside) or INFRASTRUCTURE (internal tooling)? If I've done 3+ infrastructure tasks in a row without touching customer-signal items, it flags the pattern. One escalation, no nagging.\nContent Pipeline\nNot just \"write a post.\" The full pipeline:\nDraft.\nContent-marketing-draft agent (runs on Sonnet for voice fidelity) writes against a 950-word voice profile mined from my published posts. Specific patterns: short sentences for rhythm, self-deprecating honesty as setup, \"works, but...\" concession pattern, insider knowledge drops.\nVoice check.\nAnti-pattern scan: no em-dashes, no AI preamble (\"In today's rapidly evolving...\"), no hedge words, no lecture mode. If the draft uses en-dashes, comma-heavy asides, or feature-bloat paragraphs, it gets flagged.\nPlatform adaptation.\nEach platform gets its own version: Reddit (long-form, code examples, technical depth), LinkedIn (punchy fragments, professional angle, links in comments not body), X (280 chars, 1-2 hashtags).\nPost.\nThe\n/post-article\nskill handles cross-platform posting via browser automation. Updates tracking docs, moves files from Approved to Published.\nEngage.\nThe\n/engage\nskill scans Reddit, LinkedIn, and X for conversations about topics I've written about. Scores opportunities, drafts reply angles. That Reddit comment that prompted this post? Surfaced by an engagement scan.\nI currently have 20 posts queued and ready to ship across Reddit and LinkedIn.\nBusiness Operations\nThis is the part most people don't expect from a CLI tool.\nLegal.\nOrganized documents, extracted text from PDFs (the hook converts 50K tokens of PDF images into 2K tokens of text automatically), researched state laws affecting the business, prepared consultation briefs with specific questions and context, analyzed risk across multiple legal strategies. All from the terminal.\nTax.\nCompared 4 CPA options with specific criteria (crypto complexity, LLC structure, investment income). Organized uploaded documents. Tracked deadlines.\nInsurance.\nResearched carrier options after one rejected the business. Compared coverage types, estimated premium ranges for the new business model, identified specific policy exclusions to negotiate on. Prepared questions for the broker.\nDomain & brand research.\nWhen considering a domain change, researched SEO/GEO implications, analyzed traffic sources (discovered ChatGPT was recommending the app as one of 5 in its category — hidden in \"direct\" traffic), modeled the impact of a 301 redirect over 12 months.\nNone of this is code. It's research, synthesis, document management, and decision support. The same terminal, the same personality, the same workflow.\nData & Analytics\nLocal analytics replica.\n125K rows synced from the production database into a local SQLCipher encrypted copy in 11 seconds. Python query library with methods for funnel analysis, archetype distribution, traffic sources, daily summaries. Ad-hoc SQL via\nmake quiz-analytics-query SQL=\"...\"\n.\nTraffic forensics.\nInvestigated a traffic spike: traced 46% to a 9-month-old Reddit post, discovered ChatGPT referrals were hiding in \"direct\" traffic (45%). One Reddit post was responsible for 551 sessions.\nSurvey analysis.\n573 responses from a 3-question post-quiz survey. Cross-tabulated motivation vs. experience level vs. biggest challenge.\nSelf-Improvement Loop\nThis is the part that compounds.\nSession wrap-up.\nEvery session ends with\n/wrap-up\n: commit code, update memory, update status docs, run a quick retro scan. The retro checks for repeated issues, compliance failures, and patterns. If it finds something mechanical being handled with prose instructions, it flags it: \"This should be a script, not more guidance.\"\nDeep retrospective.\nPeriodically run\n/retro-deep\n— forensic analysis of an entire session. Every issue, compliance gap, workaround. Saves a report, auto-applies fixes.\nMemory management.\nPatterns confirmed across multiple sessions get saved. Patterns that turn out wrong get removed. The memory file stays under 200 lines — concise, not comprehensive.\nRules from pain.\nEvery rule in the system traces back to something that broke. The plan-execution pre-check exists because I re-applied a plan that was already committed. The bash safety guard exists because Claude tried to\nrm\nsomething. The PDF hook exists because a 33-page PDF ate 50K tokens. Pain -> rule -> never again.\nThe Meta\nHere's the thing that's hard to convey in a feature list: all of this happens in one terminal, in one conversation, with one personality that has context on everything.\nI don't context-switch between \"coding tool\" and \"business advisor\" and \"content writer.\" I talk to Jules. Jules knows the codebase, the business context, the content voice, the pending decisions, and yesterday's session. The 116 configurations aren't 116 things I interact with. They're the substrate that makes it feel like working with a really competent colleague who never forgets anything.\nA typical day touches 4-5 of these categories. Monday I might deploy a feature, analyze survey data, draft a LinkedIn post, and prep for a legal consultation. All in one session. The morning briefing tells me what needs attention, voice dictation routes work to the right depth, and wrap-up captures what happened so tomorrow's briefing is accurate.\nThat's what I actually do with it.\nThis is part of a series. The\nprevious post\ncovers the full setup audit. Deeper articles on hooks, the morning briefing, the personality layer, and review cycles are queued. If there's a specific workflow you want me to break down further, say so in the comments.\nRunning on an M4 MacBook with Claude Code Max. The workspace is a single git repo. Happy to answer questions.\nsubmitted by\n/u/jonathanmalkin\n\nOriginally posted by u/jonathanmalkin on r/ClaudeCode","offTopic":true},{"id":"5bb53182-44d3-4235-9c9d-b3898a92f173","excerpt":"6 months of vibe coding: what I wish I knew when I started — I went from not knowing how to code to building fully functional iPhone apps within six months. Over those six months, I messed up a ton and learned a lot along the way. I wanted to share my story to help others who are just getting started or starting to wor","url":"https://www.reddit.com/r/ClaudeAI/comments/1vzxyi6/6_months_of_vibe_coding_what_i_wish_i_knew_when_i/","role":"request","weight":1.0483433,"occurredAt":"2026-08-27T15:57:01.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"ClaudeAI","intent":"feature_request","painScore":0.255,"sentiment":0.50877196,"confidence":0.83533335,"matchedPatterns":["wish","please_add","product:chatgpt"],"statement":"6 months of vibe coding: what I wish I knew when I started.","title":"6 months of vibe coding: what I wish I knew when I started","body":"I went from not knowing how to code to building fully functional iPhone apps within six months. Over those six months, I messed up a ton and learned a lot along the way. I wanted to share my story to help others who are just getting started or starting to work on more complex projects.\n\nAfter writing this, I realized it was waaay longer than I intended.  If anyone is interested in an AMA, let me know in the comments and I’ll get one going.\n\n*Tldr;* \n\n* ***Just start.*** *Pick Claude Code or ChatGPT. Don't worry about having the perfect model.*\n* ***Build something stupidly simple*** *that you actually care about.* \n* ***Don't over engineer your first project.*** *Prompt → build → test → fix is completely fine initially.*\n* *Once the project gets serious,* ***stop working directly on main.***\n* ***Start planning*** *features before asking AI to code them.*\n* ***Use worktrees*** *when you want multiple agents/features running simultaneously.*\n* *As complexity grows,* ***find a way to orchestrate and track*** *everything.*\n* ***Watch for AI slop*** *and periodically clean up/document the codebase.*\n* *Don't turn vibe coding into a 14-hour a day addiction.*\n\nI should call out that I’m not trying to build enterprise level software. I’m building apps to help improve my friends’ and family’s daily lives and doing things I never thought possible just a few months ago.\n\nAnd it was my 5 year old that actually got me into vibe coding.\n\nSix months ago, he asked if I could build him a game where a dino throws bars of soap at a stinky baby and the baby has to avoid them.  I popped open Claude Code, chose Sonnet for my model, and about five hours later had a fully working 15 level, 8-bit style game that blew my kid’s mind.\n\nIt was a ton of fun.\n\nhttps://preview.redd.it/sv6f7e3ewxlh1.png?width=1012&format=png&auto=webp&s=73ff456a54daeccc34c58017c1ebb08a03dd55d9\n\nThis was a simple HTML game. My project was basically just sitting in my iCloud environment. The entire game was coded directly on main, and my one page initial prompt probably built 80% of the game for me.\n\nIf this is all you are ever looking to accomplish, go get Claude Code or ChatGPT and just start playing around with prompts. You can build some pretty incredible things with almost no experience.\n\nBut when I started trying to build more complicated apps, I quickly learned that getting AI to write the code was actually the easy part, but managing everything else was time consuming and cumbersome.\n\nToday, I’ve built an app called Plate It.\n\nPlate It can take a recipe from almost any video source or webpage and record it in an easy to view format. No more ads and no more endless scrolling to find the bloody recipe.\n\nIt can translate recipes into different languages, build grocery lists, create meal plans, bookmark favorites, tag recipes for allergies, search through everything, and let me share recipes with friends and family.\n\nhttps://preview.redd.it/lduc1d3ewxlh1.jpg?width=2048&format=pjpg&auto=webp&s=534bf22c1ed7987d0d1840ae1b85ea33110834e2\n\nAs someone who loves to cook, building this has been amazing.\n\nBut Plate It also taught me how quickly vibe coding can become complicated.\n\nHere are probably the biggest things I’ve learned.\n\n**1. Don’t get too caught up in which AI coding model is the best.**\n\nClaude Code and ChatGPT both do a great job coding.\n\nYou will always find a lot of people online telling you how much the other one sucks. Don’t let this get into your head. I’ve used both and found both capable of delivering what I needed.\n\nI personally use Claude Code more often today, in part because I’ve found it has a larger ecosystem of third party plugins that fit the way I work.\n\nBut one of the biggest things I’ve learned is that the model itself eventually matters a lot less than the process you put around it. A great model with a terrible prompt, no planning, no review, and no understanding of your codebase can still create a mess.\n\n**2. Getting AI to write code is the easy part.**\n\nWhen I first started, my workflow was basically:\n\nHave an idea, explain it to Claude, let Claude build it, test it, ask Claude to fix whatever broke.\n\nFor a little HTML game, that worked surprisingly well.\n\nAs my projects got bigger, this started falling apart.  I’ll talk about how I fixed this later.\n\nFeatures became dependent on other features. One change would break something else. I would forget why something had been built a certain way. I would ask the AI to make a change and suddenly realize it had modified something completely outside what I wanted it touching.\n\nThe bigger the app became, the more important planning became.\n\n**3. Stop coding everything directly on main.**\n\nI know. I can hear the cringes from miles away.\n\nWhen I started, I had no idea what a branch or worktree was. Coding directly on main was just the easiest way for me to get started.\n\nEventually, I wanted to build multiple features at the same time without one agent interfering with another.\n\nThat led me to worktrees.\n\nFor anyone non-technical like me, I basically think of a worktree as giving an agent its own copy of the project where it can build and test something without messing with the main version of my app.\n\nThis was a huge step forward.\n\nBut it created another problem.\n\nNow I had to manage all the worktrees.\n\n**4. Planning before coding dramatically improved my output.**\n\nOne of the best things I found along the way was Compound Engineering.  This is a plug-in you can add to Claude.  I cannot recommend this enough.  You can find this for free on Github.\n\nhttps://preview.redd.it/it5dpb3ewxlh1.png?width=2048&format=png&auto=webp&s=c1d76611c397a9bd900a15e6cddb30c4f95964cf\n\nInstead of just throwing a prompt at an AI and saying “build this,” the work goes through a process.\n\nIt brainstorms the idea with you, researches the codebase, creates a plan, has agents execute the work, reviews the output, and then records what it learned so future agents can use that information.  It kicks off with a simple /ce-brainstrom command followed by your prompt.\n\nYes, it uses more tokens.\n\nBut I have found the quality of my output to be tenfold better than when I just throw a feature request directly at an AI and tell it to start coding.\n\nOne of my biggest lessons has been that spending more time figuring out exactly what you want built before anyone starts writing code saves an incredible amount of time later.\n\n**5. Running multiple AI coding agents creates an entirely new problem.**\n\nThis was probably the biggest surprise for me.\n\nOnce I learned how to use worktrees, I started running more things in parallel.\n\nThat was awesome at first.\n\nThen suddenly I had to know:\n\nWhich feature should be worked on next?\n\nWhich worktrees can run at the same time?\n\nDoes one feature depend on another being finished first?\n\nWhich branch is ready?\n\nWhat has been tested?\n\nWhat has been reviewed?\n\nWhat can safely merge back to main?\n\nWhat happens if two agents modify the same part of the app?\n\nI eventually realized I wasn’t struggling to get code written anymore.\n\nI was struggling to manage everything while writing the code.\n\nFind a good orchestration layer or IDE that makes it easier for you to manage your projects.  I was getting lost in window hell trying to operate out of Claude Code.\n\nI now have everything centralized on a single screen and only leave my orchestration layer to test code in the IOS Simulator or load the code on Xcode to my phone.\n\nHere’s what my set-up looks like today:  My files are all on the left, my terminal sits in the center, my agents and worktrees are on the right, and my hot keys are on the bottom.  \n\nhttps://preview.redd.it/m9k1nu3ewxlh1.png?width=2048&format=png&auto=webp&s=9be49b685d3d10256087841606f04ae97c5893a7\n\n**6. Eventually, I needed an AI managing the AIs.**\n\nThis has probably been the biggest change in how I build today.\n\nOnce I started running multiple agents and worktrees, I realized I didn’t have the capacity to be managing what every coding agent should be working on next.\n\nI wrote a pretty simple mission statement for an AI orchestrator that basically said:\n\n* Take my high level product ideas and turn them into complete engineering requirements.\n* Ask me questions when a real product decision needs to be made.\n* Break the work into tickets and organize those tickets into sprints.\n* Coordinate the engineering agents actually writing the code.\n* Figure out what work can happen in parallel and what has dependencies.\n* Track development through worktrees, branches, testing, review, and merge.\n* Make sure agents stay within the scope I approved.\n\nAnd most importantly, keep me informed without requiring me to understand or investigate the underlying codebase.\n\nThe important lesson for me wasn’t necessarily the specific tool. It was realizing that if AI agents were going to do more and more of the actual engineering work, I needed something above them managing the process.\n\nI use Argus inside Scape for this today.\n\nNow I can give Argus something as simple as a rough product idea. It asks me follow-up questions, turns the idea into a full product request, figures out where it should fit into the development schedule, determines what other work it depends on, coordinates the coding agents, and eventually gets the work to the point where I can test it.\n\nHere’s a snapshot of the ticketing system it built for me and the environment I work out of today.  All of this was created by my AI - from the ticketing system to the requests themselves.\n\nhttps://preview.redd.it/pfb9nd3ewxlh1.png?width=2048&format=png&auto=webp&s=cc94a432d08c3608a949f899eefcf329fb513100\n\nMy role has basically become coming up with product ideas, making the decisions only I can make, and testing what gets built.\n\nFor someone who had no clue how to code six months ago, that is kind of insane.\n\n**7. AI slop is very real.**\n\nI absolutely created a lot of it.\n\nEspecially early on.\n\nThe dangerous part is that your app can keep working while the underlying code gets worse and worse.\n\nThen you ask for one seemingly simple feature and suddenly everything starts breaking.\n\nThe biggest improvements for me came from slowing down before coding, creating better requirements, reviewing the work, keeping agents inside a defined scope, and documenting what was learned so the next agent didn’t have to rediscover everything.\n\n**8. Remember to do things other than code.**\n\nThis one may sound silly.\n\nWhen I first started vibe coding, seeing the progress I was making was incredible. It provided this constant stream of dopamine and all I wanted to do was code.\n\nI would have 14 hour sessions where I forgot to eat.\n\nThere is always another feature.\n\nThere is always another idea.\n\nThere is always something you want to fix.\n\nI’ve gotten much better about getting outdoors, exercising, spending time away from the computer, and accepting that the app does not need to be finished tomorrow.\n\nMy progress may have slowed down a little, but that is probably a good thing.\n\nI’m still very much learning as I go.\n\nI’m definitely not claiming that six months of vibe coding suddenly makes me a software engineer.\n\nBut the difference between how I was building six months ago and how I’m building today is pretty wild.\n\nI’d be really interested to hear how other people are managing increasingly complicated vibe coded projects, especially those of you running multiple agents or worktrees.\n\nLet me know if you want to learn more.\n\n","offTopic":true},{"id":"16c678cd-d00c-42be-a1ec-d4b3d0ea932f","excerpt":"OpenClaw in Production: 98 Real Automations People Are Running — People keep debating prompts like that’s the whole game. The shift is agents running coordination loops. Not “thinking” or “creativity”. Coordination. The stuff that quietly eats teams.\n\nOn the dev side, the most real pattern is a coordinator that supervi","url":"https://www.reddit.com/r/openclaw/comments/1r5ch5g/openclaw_in_production_98_real_automations_people/","role":"request","weight":1.0039067,"occurredAt":"2026-02-15T11:34:18.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"openclaw","intent":"feature_request","painScore":0.36,"sentiment":0.47826087,"confidence":0.7381667,"matchedPatterns":["missing_feature","product:github actions"],"statement":"It reads diffs for missing tests, risky changes, confusing logic, and security footguns and gives private feedback that actually saves time.","title":"OpenClaw in Production: 98 Real Automations People Are Running","body":"People keep debating prompts like that’s the whole game. The shift is agents running coordination loops. Not “thinking” or “creativity”. Coordination. The stuff that quietly eats teams.\n\nOn the dev side, the most real pattern is a coordinator that supervises multiple coding workers at once. Think 5 to 20 parallel code instances running in tmux or SSH, getting tasks, running tests, opening PRs, updating you in Telegram. The phone version is simple: you text “fix tests” or “ship feature X” and it loops until it can show you a diff, a test pass, and a safe next step. A good PR review agent is also real value. It reads diffs for missing tests, risky changes, confusing logic, and security footguns and gives private feedback that actually saves time. People also use agents to generate diagrams from “draw this flow”, and to run heavy scraping or data pipelines that pull huge volumes of posts across many accounts.\n\nOn ops and sysadmin, the pattern that replaces humans is “3AM autopilot”. Sentry or GitHub Actions fires, the agent pulls logs, summarizes what changed, opens an issue, proposes a fix, and creates a PR. Some teams run an ops hub that watches Slack or Basecamp, watches Sentry, drafts fixes, drafts incident notes, and keeps a paper trail. CI/CD plus dependency drift monitoring is another quiet win. It watches builds, tests, deploys, and security advisories, and tells you what’s safe to bump and what will break.\n\nEmail and calendar automations are already replacing assistants. People are doing backlog cleanup that unsubscribes spam, categorizes by urgency, drafts replies, and creates persistent rules so the problem stays solved. Daily digests with reply drafts are common. Some setups turn important threads into GitHub issues or tasks automatically. Calendar agents do timeblocking by scoring tasks, resolving conflicts, and protecting deep work. On the business side, that merges into CRM workflows that generate Monday reports, invoice prompts, and follow-up scheduling.\n\nHome automation is real but the money is still in the boring parts: Home Assistant control, routines, device orchestration, and voice control integrations. The underrated move is giving the agent its own “home” hardware so it’s separate from your personal machine.\n\nContent and social is a huge category because it’s easy to automate without scary permissions. People run daily pipelines that scan trends, analyze what’s getting engagement, draft posts, and schedule. Some monitor competitor RSS feeds and turn them into threads. Others clip long videos into shorts, apply platform formatting, add hashtags, and schedule. Brand monitoring is common too. It watches mentions, sentiment, and complaints that need action. There are also simple curators like “my personalized Hacker News” and Reddit crawlers that deliver relevant posts via Telegram.\n\nBusiness operations is where agents quietly replace small teams. Real estate CRMs, small business ops, recruiting pipelines, deal sourcing, client onboarding, weekly SEO reporting, invoicing, and work summaries. The “AI employee” framing usually works best when it’s scoped to a set of routines with artifacts, not freeform chat.\n\nFinance and trading has the usual bots: prediction markets, crxxto arbitrage, sentiment monitoring, and investment research support. The safer and more common “serious” version is a knowledge graph for investing, where the agent builds structured notes and connections, and you still make the call.\n\nPersonal productivity is basically a stack now: morning briefs with weather, objectives, meetings, health stats, reminders, trends, reading queue. Meeting transcription that turns into action items and decisions. Voice notes turned into journals. Research and meeting prep packets. File organization and dedupe at scale. Receipt OCR to expenses or parts lists. Even “bookmark discussion partner” workflows where the agent reads what you saved and helps you think.\n\nHealth workflows exist too: fitness dashboards, structured lab results, claims and reimbursement filing. Shopping and travel has practical automations like package tracking dashboards from email, flight check-in helpers, price tracking, and trip cost splitting. The wilder versions include car negotiation through email and browser automation, but that one can get messy fast.\n\nRobotics and gaming show up in the edges. Agents controlling ROS systems, running OpenCat operations, setting up Minecraft servers for kids, or doing overnight game dev builds. There are also social experiments like personality rewrites, guestbook-style agent messages, and group chat automation.\n\nThe thing tying the best examples together is not “always on”. It’s governed. The agent wakes, checks, produces an artifact you can verify, and only escalates when it has evidence. Diffs. Logs. A PR. A checklist. A decision memo. Always-on agents just burn tokens and hallucinate productivity. Governed ones compound because they create assets you can trust.\n\n  \nWhat’s the first boring task you’d happily hand to an agent if it could only output diffs and checklists, not push buttons.","offTopic":false},{"id":"747b0eeb-eeeb-4445-8382-093831a0a073","excerpt":"I broke my AI agent setup constantly for months — here's what finally worked. For noobs and beginners. — **This is for people if you are new or if you are frequently having problems with the same things.**  \n  \nI'm still an amateur at learning my stuff with this, but basically I started on ChatGPT and had it on auto-up","url":"https://www.reddit.com/r/openclaw/comments/1t9f1fe/i_broke_my_ai_agent_setup_constantly_for_months/","role":"request","weight":0.97255206,"occurredAt":"2026-05-10T18:37:46.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"openclaw","intent":"feature_request","painScore":0.24,"sentiment":0.13953489,"confidence":0.7843162,"matchedPatterns":["wish","free_tier","product:chatgpt"],"statement":"# Always plan before you execute This is the one thing I wish someone had told me earlier: **tell it to plan first, then wait.** Before you let your agent start working on anything significant, have it lay out the full plan in plain langua…","title":"I broke my AI agent setup constantly for months — here's what finally worked. For noobs and beginners.","body":"**This is for people if you are new or if you are frequently having problems with the same things.**  \n  \nI'm still an amateur at learning my stuff with this, but basically I started on ChatGPT and had it on auto-update. I was constantly fixing things, and it felt like I never managed to get anything working. It'd work on day one, and on day two it'd be broken again. I ended up reinstalling it about six or seven times because I made another error where I installed it in Docker and then installed it another way, and then kept bringing ghost and stuff. I don't know how much of it was my fault or how much of it was ChatGPT's fault.\n\n\n\nKnowing this, I set it all up on OLlama Pro with Deep Seek. It was all working perfectly, then I ended up having trouble with OLlama Pro, which was completely blocked. Deep Seek and all the other models that I'd set up as backups were not working. What I'd done at that point, on that same day, was remove ChatGPT because I decided I wasn't going to pay the subscription. I had blamed everything on ChatGPT, but then because OLlama stopped working, I thought it was more likely the update, and that ChatGPT was totally fine. It was just because I was blindly updating all the time.\n\n\n\nI ended up putting ChatGPT back on, thinking, \"It's a frontier model, they must know what they're doing.\" I replaced Deep Seek 100% and had it as the backup only and was using ChatGPT only. The problems came back exactly the same as before. I just set up a shave reminder, as basic as that is, to go off every three days and bug me every three hours until I go and shave, because I got ADHD and I just forget everything and my life runs by reminders. I would like something that just happened automatically. \n\nBut then it just, it just couldn't work. It wasn't doing anything. Everything was just going wrong. Just to get the simple news in the morning because I'm trying to keep up to date with this stuff. It took me like eight hours to get nowhere. And then I ended up just putting DeepSeek back on there and then everything just started to work again. It took a little while for it all to have been working so much better for me. And I also made another discovery that if you put OpenRooter on auto, that can also become very expensive because it opts to use the best cloth model, which I've discovered is pretty damn expensive. So now I have selected models on OpenRooter, using Gemini and stuff like that. And now everything is working flawlessly. So I've set up something very complicated now which I love and it works flawlessly. And it's reduced what I would do on a fortnightly or monthly basis for my clients for every single client into something that takes five to ten minutes with high accuracy. And it's amazing.   \n  \nSo I don't know what others can learn from this, but careful with updates essentially and careful with chat GPT because I found that both of them will just destroy everything. I make it do  back ups every day and use backups before I do big things and have it create a disaster recovery file that's available on my VPS and on my Google Drive. My backups are on Google Drive and on my VPS so no matter what the situation happens it's always available there and it's on GitHub as well because I'm a bit paranoid because I've spent so many hours trying to get this done that I just wanted to make sure that no one including me lol could ever take it away from me. \n\nNow I have it doing several tasks. The next thing that I need to do is some sort of mission control view (cabana sort of view) so I can see what's done, what's being done, what part of the process is, and this sort of thing, because this is the huge hole that I have right now.   \n  \nFor this I spend the $20 on ollarma pro plus 3 or 4 USD on open rooter as the backup. I have to have Gemini as well (2 or 3 accounts with the pro version), the subscription with some of my Google accounts and stuff like that. So another thing that I do is sometimes I just copy and paste between the chat model and my OpenClaw, and get Gemini to help me out as well. So if you get to the point where it's just telling you a ton of stuff but you just don't know what really is going on, instead of asking OpenClaw what's going on, ask Gemini or Claude (or whatever) to explain it to you and tell you the best options. The reason why I say that is because the context window is much larger on Gemini. It's like 1.2 million, whereas on OpenClaw it's 200,000. So asking questions can end up pushing content out of the memory, meaning that it becomes less accurate, so it's just a little FYI of what I've discovered. \n\nChatGPT kept breaking my setup. Switched to Ollama. Then blamed Ollama. Turns out both the platform and auto-updates were the problem. Now I run Ollama Pro + OpenRouter fallback + Telegram and it's rock solid. Here's the full picture — including the parts most people skip.\n\nIf you're just starting with AI agents (I use OpenClaw), maybe this saves you some pain.\n\n# The mistakes\n\n**My first mistake: ChatGPT.** Every update, every change — something would stop working. I reinstalled it 6-7 times. Eventually realised the platform itself was unstable for what I needed.\n\n**Switched to Ollama + DeepSeek.** Worked perfectly — until Ollama completely died on me. I thought \"must be Ollama's fault too.\" But I'd also been hammering updates the whole time.\n\n**The real culprit? Both.** ChatGPT wasn't reliable for this use case. And blindly updating everything was making problems worse, regardless of platform.\n\n# What actually fixed it\n\n**Pin your versions.** No more auto-updates. If it isn't broken, don't touch it.\n\n**Ollama Pro as primary.** Stable, predictable, and the model quality is genuinely excellent — especially DeepSeek V4 Pro, which is my main workhorse right now.\n\n**OpenRouter as your fallback — and take this seriously.** Ollama has gone down on me before. It will probably happen again. If your workflow depends on this at all, you need a fallback. OpenRouter is cheap, reliable, and gives you access to a huge range of models. Pick specific models though — leaving it on \"auto\" gets expensive fast.\n\n**My current model list**:\n\n|Model|Role|\n|:-|:-|\n|DeepSeek V4 Pro|Primary — best quality|\n|DeepSeek V4 Flash|Fast tasks, low cost|\n|Gemma 4|General fallback|\n|Qwen 3 Next / Qwen 3.5|Strong reasoning|\n|Kimi K2.6|Long context tasks|\n|GLM 5.1|Alternative backbone|\n|Gemini 2.5 Flash Lite|Budget fallback|\n|Gemini 3.1 Flash Lite|Budget fallback|\n|Grok 4.3|Wildcard / testing|\n|Owl Alpha|Free tier|\n|Ring 2.6 1T|Free tier|\n\nHaving multiple models isn't about being indecisive. It's about always having something running, no matter what goes down.\n\n# Always plan before you execute\n\nThis is the one thing I wish someone had told me earlier: **tell it to plan first, then wait.**\n\nBefore you let your agent start working on anything significant, have it lay out the full plan in plain language. Read it. Check it's actually what you want. Then tell it to go ahead.\n\nIf you skip this step, you'll watch it confidently sprint in the wrong direction for 20 minutes, burn through your token budget, and produce something you'll have to throw away. Planning costs almost nothing. Fixing a runaway agent costs a lot.\n\n# The Palace of Truth\n\nOne of the most useful things I've built is a file I call **The Palace of Truth**.\n\nIt's my master index. Every file in the system has an entry there — a short description of what it is, what it does, and how it fits into everything else. If you need to find something, you go to The Palace of Truth first. It makes the whole setup navigable, even after months of additions.\n\nOnce a month, the agent goes through the index — one file per heartbeat — just to verify everything is still accurate, up to date, and working well with the best available model at the time. It's a slow, methodical self-audit. Keeps things clean without ever needing a big manual overhaul.\n\n# The model update cron job\n\nI also run a scheduled job that checks the available models on both OpenRouter and Ollama Pro regularly, comparing them against my use cases. As new models release and old ones get superseded, the list updates automatically. You're always running the best option available — not whatever was best six months ago when you first set things up.\n\n# Backups — three places, no exceptions\n\nDaily backups to:\n\n* **Local VPS**\n* **GitHub**\n* **Google Drive**\n\nI've lost too many hours of work to trust anything less. If one fails, two others have it. This isn't paranoia — it's just experience.\n\n# Results and cost\n\nWhat used to take me hours per client per month work now takes 5-10 minutes. With high accuracy.\n\n**Monthly cost:**\n\n* Ollama Pro: \\~$20\n* OpenRouter: \\~$3-4","offTopic":false},{"id":"74518891-84a6-4c97-bc2a-df28c74e96c1","excerpt":"OpenClaw Skills I Actually Use — A technical, honest look at the skills powering my daily workflow. These are not hypothetical. I use them in production.\n\n  \n**The Setup I Actually Run**  \nI am a one‑person team. I still make every call, review every critical change, and take the heat when something breaks. The agents ","url":"https://www.reddit.com/r/openclaw/comments/1r2uz4e/openclaw_skills_i_actually_use/","role":"request","weight":0.9138168,"occurredAt":"2026-02-12T14:33:42.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"openclaw","intent":"feature_request","painScore":0.3,"sentiment":0.18181819,"confidence":0.702936,"matchedPatterns":["wish","paying_monthly","product:chatgpt"],"statement":"I care about shipping real products, not demos, and I document what works because I wish someone had done that for me when I started.","title":"OpenClaw Skills I Actually Use","body":"A technical, honest look at the skills powering my daily workflow. These are not hypothetical. I use them in production.\n\n  \n**The Setup I Actually Run**  \nI am a one‑person team. I still make every call, review every critical change, and take the heat when something breaks. The agents speed me up, but the responsibility is mine. I care about shipping real products, not demos, and I document what works because I wish someone had done that for me when I started.\n\nI manage 12 SaaS projects simultaneously with a fleet of AI agents. Each project has its own dedicated agent, Telegram group, GitHub repo, and workspace. I am still a one‑person team. The agents are the team.\n\nThis article is not a catalog. It is the real architecture and the few skills that make it work.\n\n**The Architecture (End to End)**  \nMe → Bob (PM agent) → 12 project agents → Telegram groups → GitHub repos → auto‑deploy\n\nI talk to the PM agent on Telegram.  \nThe PM agent delegates to each project agent.  \nEach project agent works inside its own repo and workspace.  \nEvery change becomes a PR, is reviewed, and then deployed automatically.  \nThat is the system. The skills are tools that keep it moving.\n\n**The Daily Loop (What Actually Runs)**\n\n* 3 AM: Nightly maintenance — memory reindex, backups, self-update, gateway restart.\n* 8 AM: Health checks, email monitoring across 7 accounts, cron health monitor, AI trend scan.\n* 9 AM: All 9 project agents generate briefs. Roll-up lands in my Telegram DM by 9:15.\n* 10 AM - 8 PM: Automated Twitter posts for StyleMCP (3x/day) and InstantContent (3x/day).\n* 12 PM: PR Autopilot reviews and merges AI-created PRs across all repos.\n* 5 PM: Project tracker syncs Umami analytics to the DD dashboard.\n* 10 PM: Dashboard sync — social, monitoring, errors, cron calendar.\n* 11 PM: Ideas research — scans for new product opportunities.\n\n  \nWeekly overnight: Each project gets its own QC and enhancement scans on a rotating schedule.\n\n  \n24 cron jobs. Zero human intervention. That loop is why I can run 12 projects without a team. See the full schedule.\n\n**Apple ecosystem workflow**  \nI run a Mac Studio with OpenClaw on one monitor and a MacBook Pro on the other. One keyboard and mouse, no KVM, no friction. Apple’s Universal Control makes both machines feel like one system, so I can slide my cursor across devices, drag files between them, and keep context without breaking flow. I also AirDrop photos from my phone to my desktop in seconds, and my shared notes update on my phone and appear on my desktop immediately.\n\n**My setup:**\n\nMac Studio (OpenClaw running)  \nMacBook Pro (research, writing, video)  \nDual monitors  \nOne keyboard + mouse across both devices via Universal Control  \nThe Skills I Use Every Day  \nThese are the only ones that matter daily. Everything else is removed from the system.\n\n1. GitHub (non‑negotiable)\n   * What it is: GitHub CLI automation.\n   * How I use it:\n      * Every agent opens PRs and manages issues\n      * CI checks are triggered automatically\n      * Auto‑deploy runs on merge (Cloudflare in most cases)\n2. coding-agent (the backbone)\n   * What it is: Codex / Claude Code execution.\n   * How I use it:\n      * Feature implementation\n      * Bug fixes\n      * QC scans and refactors\n3. Himalaya (email)\n   * What it is: IMAP/SMTP automation.\n   * How I use it:\n      * Inbox triage across products\n      * Draft replies and summaries\n      * Follow‑ups without losing context\n4. summarize (research acceleration)\n   * What it is: URL and document summarization.\n   * How I use it:\n      * Competitor research\n      * Product teardown notes\n      * Content briefs\n5. Marketing skills suite\n   * What it is: 26 public skills built around CRO, copywriting, SEO, pricing, launch strategy, and analytics. \n   * How I use it:\n      * Turn ideas into pages fast\n      * Run CRO checks and copy sweeps\n      * Build SEO clusters and structured data\n      * This is my marketing department in code, built on public skills.\n6. Frontend design (Superdesign)\n   * What it is: The design system encoded as an AI skill. It keeps layout, typography, spacing, and dark theme consistency across every site I ship.\n   * How I use it: \n      * Every UI agent references it when building landing pages. It is why the sites feel polished rather than template‑built.\n7. openai-coder (custom review)\n   * What it is: Codex review agent for validation and QC.\n   * How I use it:\n      * Cross‑check logic and edge cases\n      * Validate fixes before merge\n8. openai-image-gen (visual drafts)\n   * What it is: Image generation.\n   * How I use it:\n      * Draft visuals for landing pages\n      * Quick image variations\n9. nano-pdf (document edits)\n   * What it is: PDF editing via natural language.\n   * How I use it:\n      * Quick doc edits for releases and reports\n10. video-frames (content assets)\n   * What it is: Frame extraction with ffmpeg.\n   * How I use it:\n      * Thumbnails and short clips\n11. Peekaboo (UI automation)\n   * What it is: macOS UI automation.\n   * How I use it:\n      * UI tasks with no API access\n12. skill-creator (build new skills fast)\n   * What it is: Skill builder.\n   * How I use it:\n      * Turn recurring workflows into reusable skills\n13. discord + imsg (messaging)\n   * What it is: Messaging connectors.\n   * How I use it:\n      * Alerts and quick responses\n14. apple-notes + apple-reminders (capture)\n   * What it is: Apple Notes and Reminders connectors.\n   * How I use it:\n      * Fast capture and reminders\n15. second-brain (long-term)\n   * What it is: My internal knowledge system.\n   * How I use it:\n      * Long‑term memory and decision tracking\n\n**What’s Next**  \nI am tightening the nightly QC loop so that every repo is scanned with the same rigor every night. I am expanding the marketing skills to cover more lifecycle stages, not just acquisition. I am also adding stronger deployment guards so broken changes never hit production without a clear rollback path. All of this isn't cheap, I'm spending $600 a month on two Claude Max plans and one ChatGPT Pro plan, but that is nothing compared to what an employee would cost.\n\nThe short version: less noise, tighter feedback loops, and faster shipping without sacrificing reliability.","offTopic":false},{"id":"967d18f9-25f8-49cb-943a-c0d045f010a5","excerpt":"I replaced 5 hires with 5 AI agents running on my laptop. Here's the system after 6 weeks — Six weeks ago I set up a system where five AI agents handle five different jobs across my entire workflow: engineering, back-office ops, information dashboards, research and marketing. No employees, no freelancers, no SaaS subsc","url":"https://www.reddit.com/r/whaaat_ai/comments/1sz19xx/i_replaced_5_hires_with_5_ai_agents_running_on_my/","role":"request","weight":0.9007286,"occurredAt":"2026-04-29T14:58:45.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"whaaat_ai","intent":"feature_request","painScore":0.36,"sentiment":0.14285715,"confidence":0.6623004,"matchedPatterns":["missing_feature","product:postgres"],"statement":"Research finds opportunities, the operator processes them, the cockpit shows me status, the builder ships what's missing, marketing distributes what's ready.","title":"I replaced 5 hires with 5 AI agents running on my laptop. Here's the system after 6 weeks","body":"Six weeks ago I set up a system where five AI agents handle five different jobs across my entire workflow: engineering, back-office ops, information dashboards, research and marketing. No employees, no freelancers, no SaaS subscriptions stacked on top of each other.\n\nHere's the setup.\n\n# The 5-Agent System\n\nEach agent sits on a different layer of the stack. They don't compete with each other, they complement.\n\n**Agent 1: Builder (Claude Code).** Writes code, refactors, ships features. The key is the workspace setup: a [CLAUDE.md](http://CLAUDE.md) file that teaches the agent your architecture rules, Skills (markdown files that define repeatable workflows) and MCP integrations that connect it to GitHub, Postgres, Slack and whatever else you use. Without that setup, Claude Code is autocomplete. With it, it's your first engineering hire.\n\n**Agent 2: Operator (Claude Code Pipelines + n8n).** Runs five pipelines that would normally require five people: video repurposing (YouTube link in, 10 platform-specific posts out), lead enrichment (raw company list in, scored leads with personalized openers out), competitive intelligence (weekly URL scans with change detection), invoice extraction (PDFs to structured data at 94-97% accuracy) and a knowledge base agent that turns support tickets into documentation.\n\n**Agent 3: Cockpit (Live Artifacts in Cowork).** This one surprised me the most. It's technically not an agent in the traditional sense. It's a persistent HTML dashboard that pulls live data from your connectors every time you open it. Gmail, calendar, task manager, all on one screen. The reason I went with a dashboard instead of a daily reporting agent: tokens. An agent that fetches and processes data every morning costs real money. A dashboard built once that only pulls data on demand costs almost nothing after the initial build. Took me 2 minutes to set up, saves me 20 minutes every morning.\n\n**Agent 4: Researcher (Hermes / OpenClaw with Kimi 2.6).** Hermes runs in the cloud on a $5 VPS with built-in cron jobs, parallel subagents and persistent memory. OpenClaw runs locally with access to your Obsidian vault, local PDFs and terminal output. The community trend is to stack both: Hermes for web research, OpenClaw for local context. Best model for research tasks right now is Kimi 2.6, easiest to run via an Ollama subscription or through OpenRouter.\n\n**Agent 5: Marketing (Higgsfield / whaaat ai).** The gap everyone ignores in the solo founder narrative. You build the product, you run the pipelines, Stripe shows zero. Higgsfield closes this: drop in a link, pick an AI persona, get 500+ ad-ready video cuts per day. The research agent feeds insights in, the marketing agent produces creatives, performance data flows back into the next research cycle.\n\n# What I learned after 6 weeks\n\nThe system gets better every week because each agent feeds the next one. Research finds opportunities, the operator processes them, the cockpit shows me status, the builder ships what's missing, marketing distributes what's ready.\n\nStarting point if you want to try this: build the cockpit first. It has the lowest setup effort and the fastest payoff. Connect your email, task tool and calendar, describe what you want, done in under 5 minutes.\n\nFull disclosure: I work on the AI agent team at whaaat ai. Happy to share more detail on any of the five agents. What does your current AI workflow stack look like?\n\n","offTopic":true},{"id":"0a7605f5-28e8-4edf-aece-6bb53e4bfe02","excerpt":"I ran a six-agent AI marketing team for three months. This is what it did. — *\\*I mentioned this case a few times in this sub, and were asked to share more details on it.*  \n  \nFor three months, a fintech project ran with a one-person marketing function: me, backed by six AI agents.\n\nThe agents handled social content, ","url":"https://www.reddit.com/r/AI_Agents/comments/1vyxua5/i_ran_a_sixagent_ai_marketing_team_for_three/","role":"pain","weight":0.88592887,"occurredAt":"2026-08-26T13:56:36.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"AI_Agents","intent":"feature_request","painScore":0.435,"sentiment":-0.1875,"confidence":0.61737204,"matchedPatterns":["missing_feature","product:chatgpt"],"statement":"My biggest takeaway is still the oldest rule in computing: **garbage in, garbage out.** If the brief is vague, the sources are weak, the success criteria are missing, or the underlying process is a mess, an agent scales the mess.","title":"I ran a six-agent AI marketing team for three months. This is what it did.","body":"*\\*I mentioned this case a few times in this sub, and were asked to share more details on it.*  \n  \nFor three months, a fintech project ran with a one-person marketing function: me, backed by six AI agents.\n\nThe agents handled social content, email, advertising monitoring, growth experiments, and outreach. I handled strategy, priorities, approvals, and anything with enough ambiguity or risk to require judgment.\n\nBuilt it from the ground up. The setup ran on OpenClaw. It handled schedules, tools, permissions, memory, and handoffs. Claude models did most of the underlying model work.\n\nThis is a historical snapshot from March to May 2026, after the team had been running for almost three months. The project pivoted since, so the team was wrapped up.\n\n**The six roles**\n\nI gave every agent one narrow job:\n\n1. **Orchestrator:** coordinated the other five agents, passed work between them, and routed decisions to me.\n2. **Social media:** prepared posts and distributed approved content across channels.\n3. **Email:** drafted newsletters and customer emails.\n4. **Advertising:** monitored paid campaigns and flagged changes.\n5. **Growth:** researched and tested acquisition ideas.\n6. **Outreach:** managed the influencer and partner pipeline.\n\nEach agent had its own instructions, tool access, schedule, reporting format, and stop conditions.\n\nThe handoffs were the useful part. A product update could trigger an email draft, several social posts, and a retargeting task. I did not have to copy the same context between four tools or remember to start every next step myself.\n\n**What the team produced**\n\nThe March-May snapshot included:\n\n* 20 blog posts\n* About 195 social posts across seven platforms\n* 4 newsletters\n* About 43 influencer contacts moving through an outreach pipeline\n* 2 advertising accounts with continuously active Meta and Reddit campaigns (4 full campaign updates each month)\n\nDuring the final two months, when the agents were operating with their highest level of autonomy:\n\n* Organic traffic increased 7x.\n* Referral traffic increased 10x.\n* Average cost per lead fell 30% across channels while the ad budget stayed flat.\n* Reddit organic posts received 135,000 views.\n* The project subreddit gained 300 organic subscribers who continued to send traffic.\n\nThose numbers need a caveat. Product development was moving at the same time, and this was a startup in motion, not a controlled experiment. I excluded metrics where I could not separate the agents' contribution from other changes. Even the remaining numbers do not offer clean causal attribution.\n\nThe narrower claim is the one I can defend: the agents produced the output listed above, expanded channel coverage, and operated during a period when acquisition metrics improved without a larger advertising budget.\n\n**What it cost**\n\nThe May bill was **$359 for the month**:\n\n* Hetzner VPS: $10\n* Claude Max: $200\n* ChatGPT Plus: $20\n* Gemini: $20\n* Perplexity API: about $12\n* Linear: $16\n* Postiz: $49\n* X API: $10\n* Firecrawl: $16\n* Google Workspace seat: $6\n* OpenClaw: free\n\nThe agents fit within one flat Claude Max subscription at the time, so the $359 total depends on the subscription setup we used in April-May 2026.\n\nThe $359 also leaves out the expensive part: my time.\n\nGetting an agent to a stable working state took roughly two weeks of role definition, tool connections, permissions, test runs, and instruction changes. Ongoing maintenance took about eight hours a week across the system: reviewing samples, checking sources, resolving ambiguous cases, cleaning memory, and updating rules.\n\n**What broke**\n\nThe obvious failures were easy to catch. An agent would miss a tool call, fail a scheduled run, or return an empty report.\n\nOther recurring problems:\n\n* **Generic marketing defaults.** Models reproduce familiar campaign structures, average positioning, and advice that sounds reasonable across almost any company.\n* **Source errors.** A weak answer rarely labels itself as weak. Every factual output needs a source trail.\n* **Memory decay.** Old rules conflict with new ones. Temporary facts survive as permanent instructions. More context eventually becomes more clutter.\n* **Permission mistakes.** An agent that can publish, email, spend, or delete needs explicit limits and stop conditions.\n* **Automation without demand.** A scheduled workflow keeps running even when the input becomes stale or nobody uses the output.\n\nThat changed my job. I wrote less and reviewed more. I spent more time checking samples, inspecting sources, and deciding which exceptions should become permanent rules.\n\n**What changed after another 30+ agents**\n\nSince this first team, I have built and tested more than 30 agents across several teams and niches. The results varied a lot.\n\nSome niches like ecom have abundant structured data, stable processes, and clear definitions of a good output. Agents become useful quickly there.\n\nOther niches like specific b2b SaaS depend on tacit context, taste, relationships, private data, or judgment that is hard to encode. Those agents need much more supervision, and some workflows never become worth maintaining.\n\nThe model matters. The tools matter. The process around them matters more than either.\n\nMy biggest takeaway is still the oldest rule in computing: **garbage in, garbage out.**\n\nIf the brief is vague, the sources are weak, the success criteria are missing, or the underlying process is a mess, an agent scales the mess. Usually with excellent formatting.\n\nSo we keep working on the input: narrower roles, better source rules, explicit examples, stop conditions, approval gates, and logs of recurring errors.\n\nThe agents keep getting better. The management work does not disappear. It moves into the system.\n\nBut overall, agents changed my life and my work paradigm. Love every second of it.\n\nHappy to answer any questions.","offTopic":true},{"id":"9687b98e-2261-41fd-9646-43cf20d4b28e","excerpt":"The real breakthrough wasn’t finding a better AI model. It was building a better system around it. — I honestly don’t know if many people will read this, but I’ve seen a lot of people struggling with agent memory, project persistence, context management, and all the usual problems that show up once you move beyond simp","url":"https://www.reddit.com/r/hermesagent/comments/1u64h6i/the_real_breakthrough_wasnt_finding_a_better_ai/","role":"demand","weight":0.8848792,"occurredAt":"2026-06-15T02:55:36.000Z","sourceKey":"reddit","sourceName":"Reddit","credibility":0.62,"venue":"hermesagent","intent":"alternative_search","painScore":0.165,"sentiment":1,"confidence":0.75955296,"matchedPatterns":["switching_from","praise","product:chatgpt"],"statement":"Eventually I moved from Cursor into persistent agents.","title":"The real breakthrough wasn’t finding a better AI model. It was building a better system around it.","body":"I honestly don’t know if many people will read this, but I’ve seen a lot of people struggling with agent memory, project persistence, context management, and all the usual problems that show up once you move beyond simple chatbots.\n\nSo I figured I’d share what ended up working for me.\n\nFor context, I currently help run an orthotics and rehabilitation business, manage Airbnb operations, and spend a lot of my time building software, automation, and internal systems.\n\nI didn’t start with some grand plan.  \nI started exactly like everyone else.  \nCopying code into ChatGPT.  \nPasting errors.  \nAsking it to fix things.  \nThen Cursor came out.  \nAnd suddenly I could actually build stuff.  \nAt first I used it the same way everyone does:  \nBuild this.\n\nContinue.\n\nContinue.\n\nContinue.\n\nAnd honestly? It worked surprisingly well.  \nUntil projects got bigger.\n\nThen everything started breaking down.  \nContext got messy.\n\nRequirements changed.\n\nThe AI forgot things.\n\nI forgot things.\n\nProjects became harder to manage.  \nThat’s when I became obsessed with systems.\n\nI spent months reading workflows from Twitter, Reddit, GitHub, random blogs, and eventually ended up building my own.\n\nThe first thing that changed everything for me was moving to a 4-document workflow:\n\n\\-PRD  \n\\-Architecture  \n\\-UI Specification  \n\\-Tasks\n\nThe biggest lesson?\n\nDon’t start coding immediately.  \nSpend more time defining what you’re building.  \nMy workflow today looks like:\n\nIdea  \n↓  \nPRD  \n↓  \nArchitecture  \n↓  \nUI  \n↓  \nTasks  \n↓  \nExecution\n\nThe UI part deserves special attention.  \nOne thing I’ve noticed is that if you let AI design everything by itself, most apps end up looking strangely similar.  \nNot bad.\n\nJust similar.\n\nSo now I spend time collecting references, generating mockups, using tools like Stitch, and defining visual direction before touching code.\n\nThat alone improved my results dramatically.\n\nEventually I moved from Cursor into persistent agents.\n\nI procrastinated for months before trying OpenClaw\n\nThen I bought a Mac mini.  \nInstalled OpenClaw.  \nPlayed with it.  \nDidn’t love it.\n\nEventually I found Hermes.\n\nAnd that’s where things got interesting.  \nAt first I thought the solution was finding better models.\n\nI went through:\n\nChatGPT  \nCursor  \nCodex  \nMiniMax  \nDeepSeek  \nHermes  \nOpenCode Go\n\nWhat surprised me was realizing that model quality mattered far less than I expected.\n\nA well-organized workflow with a good model consistently outperformed a chaotic workflow with a great model.  \nThat realization completely changed how I think about agents.\n\nThe second breakthrough came from my Economics background.  \nMy undergraduate thesis focuses on AI systems and organizational management.  \nOne idea I kept coming back to was legibility.\n\nOrganizations fail when information becomes invisible.\n\nAgents fail for the same reason.  \nEventually I stopped thinking about “memory” and started thinking about information systems.\n\nToday I use four layers:\n\n**1. Navigation**  \nWhat exists?  \nWhere is it?  \nProjects.  \nFolders.  \nIndexes.  \nDocumentation.\n\n**2. Knowledge**  \nWhat do we know?  \nThis lives in Obsidian.  \nIdeas.  \nResearch.  \nDecisions.  \nLessons.\n\n**3. Governance**  \nWhere should information live?  \nNot everything belongs in memory.  \nNot everything belongs in a project.\n\n**4. Digestion**  \nHow do we turn daily activity into useful knowledge?  \nConversations.  \nNotes.  \nIdeas.  \nDecisions.  \nEverything gets processed and distilled.  \nNot stored blindly.\n\nThe biggest lesson?\n\nMost people are trying to make their agents remember more.\n\nI started getting better results when I focused on helping my agents find information instead.  \nThat’s a very different problem.  \nToday Hermes doesn’t need to remember everything.\n\nIt knows where to look.  \nIt knows where projects live.  \nIt knows where documentation lives.  \nIt knows where decisions live.  \nThat turned out to be far more valuable than simply increasing memory size.\n\nSome real things I’ve built using this approach:\n\nA health monitoring app for my parents.  \nBusiness websites.  \nCustom ERP systems.  \nAI-integrated internal tools.  \nAirbnb management software.  \nAutomation systems.  \nInfrastructure tooling.  \nThe early foundations of a comics-related startup I’m working on.  \nNo, the agents didn’t build everything by themselves.\n\nYou still need judgment.  \nYou still need to think.  \nYou still need to review things.  \nBut my ability to execute has changed dramatically.\n\nAnd honestly, that’s probably the biggest thing AI has given me.  \nNot intelligence.  \nExecution.\n\nAnyway, this is already getting longer than I intended.  \nThis is just Part 1.  \nIf people are interested, I can write more about:  \nHermes  \nObsidian  \nMemory systems  \nOpenCode Go  \nAgent skills  \nMulti-project management  \nRunning agents in actual businesses  \nCurious how other people here are solving persistence and project management with agents.\n\nEdit: This is the short version.\n\nThe original post was much longer and went into the full journey, the failures, the experiments, and the evolution of the system.\n\nI used AI to create this condensed version for people who wanted the key ideas without reading a small essay.\n\nIf you’re interested in the complete version, you can find it here:\n\n[Long](https://www.reddit.com/r/hermesagent/comments/1u67h1c/the_biggest_ai_productivity_boost_i_found_had/?utm_source=share&utm_medium=web3x&utm_name=web3xcss&utm_term=1&utm_content=share_button)","offTopic":false}],"breakdown":[{"sourceKey":"reddit","sourceName":"Reddit","count":11},{"sourceKey":"lemmy","sourceName":"Lemmy","count":2}],"total":13}}