Friday 11 PM, the Wrong Signal
One of our dream team bots was doing what it always does, sorting emails and preparing replies. The rule I set that day was simple. "Auto-generate reply drafts for emails containing specific keywords." It sounded good.
Monday morning, the moment I opened my iPad, dread washed over me.
**1,247 emails (mostly outbound)**
The bot hadn't malfunctioned. The setup was precise. Too precise.
How Did This Happen
Three reasons.
1. **The "OR" logic was too broad** - Any one of five conditions triggered a reply.
2. **No duplicate detection** - Multiple emails from the same buyer each got separate responses.
3. **Forgot the sleep mode setting** - The bot should pause from midnight to 6 AM. Instead, it ran 24/7.
So from 3 AM to 6 AM Korea time, our bot had a very productive night. We discovered it the next morning.
Delete? Review? Send?
First came shock. Then came the decision.
"I don't have time to review every email. But I can't just delete them blindly either."
I opened Claude Code on my Windows device. I needed a script to quickly analyze what the bot had written.
Analysis checklist:
• Duplicate replies to the same buyer
• Percentage containing genuinely useful content
• Obvious auto-generated errors
The verdict was better than expected. About 85% were actually necessary replies. The rest were duplicates or unnecessary confirmations.
Redesigning the Operating System
That afternoon, we created new guidelines.
Three layers of automation:
• **Level 1 (Auto-execute)** - Tasks with virtually zero margin for error. Example: spreadsheet cleanup, deduplication
• **Level 2 (Review-then-execute)** - External communications. Example: generate draft, then verify once before sending
• **Level 3 (Fully manual)** - Strategic decisions. Example: handling new buyers
And one absolute rule was added.
All outbound bots must pause from midnight to 8 AM. We review everything with fresh eyes in the morning.
The Gift That Failure Gave
Here's what made it funny: most of those 1,247 emails actually got positive responses. It wasn't a bot problem. It was an operations design problem.
After this incident, dream team changed. We trust automation more, but we've learned not to over-rely on it. A bot is a great assistant, but you can't give it autonomous decision-making authority at night.
Now, the first thing I do each morning on my Windows device is review the bot's overnight activity log. Once a week, we audit the automation rules.
Non-developers can run AI agents. But it's like raising a bot. Give it autonomy, but make boundaries crystal clear.