Running a one-person business in Korea means relying heavily on scheduled scripts and AI agents to handle daily tasks like drafting posts, publishing, and pulling metrics. The most expensive failures in this stack weren't crashes or loud errors, but silent successes that looked perfect in the logs while doing nothing useful. A recent post on DEV.to by ssapable breaks down five specific automation failures where the system reported success at a step before the actual outcome mattered.
The Approval Loop and Environment Mismatches
One major failure involved a feedback loop where a bot sent drafts to a phone for approval. The bot logged the tap, but the row never made it to the Postgres database. The agent continued to 'learn' only from edits and rejections, ignoring approvals entirely. The fix was a daily SQL query comparing tap counts to database rows. Another failure occurred with a systemd timer running a Node script on a VPS. The service file pointed to /usr/bin/node (v18), while the shell used /usr/local/bin/node (v22). Playwright refused to run on the older binary, causing every scheduled job to fail silently for hours because the timer technically 'fired'.
UI Deception and Metric Blindness
UI-based automations are equally treacherous. An upload of a PDF product file stopped at 99% progress, leaving the download button pointing to the old file while the page looked fine. Without manually downloading the file as a logged-out buyer, the update would have been announced to customers who never received it. Similarly, social media replies often don't appear immediately on the post page, leading to duplicate posts if agents retry based on visible absence. The author now verifies actions on the account's own activity tab rather than the public post page to avoid spam.
The Illusion of Engagement
Perhaps the most insidious failure was a scheduled posting automation that returned valid IDs for every post, confirming they were 'live'. However, impression metrics revealed 2 and 0 views for the first two posts. The automation worked technically but failed strategically. This highlights a critical gap in AI agent design: status codes are not outcomes. The system needs to measure the end resultβviews, clicks, or actual file downloadsβrather than trusting the intermediate step that returns a success flag.
Key Takeaways
- Verify the end result where it actually lives, not the step before it (e.g., check the DB, not just the log).
- Always use full paths for binaries in systemd services to avoid environment mismatches.
- Never retry an action based solely on UI visibility delays; check reliable activity tabs instead.
- Measure output metrics (views, clicks) rather than just API success responses.
The Bottom Line
AI agents are terrible at recognizing their own uselessness. If you aren't validating the final user-facing outcome, your automation is just generating expensive noise.
Practical Implementation
The author shares a self-learning agent setup using Postgres and pgvector in the browser, along with a paid playbook titled 'Just Say Do It'. The core lesson for builders is to treat every 'success' response from an LLM or API as a hypothesis that requires verification against the real-world state. Whether it's a file download, a database row, or a view count, the truth is always at the edge, not in the middle.