The Priciest Model, Standing at the Door
I gathered every request the Dream Team bots sent over one month. Far fewer needed deep reasoning than I expected. Most were sorting jobs, like 「which language is this post in?」 or 「is this alert urgent?」. Yet I had been sending every request to the biggest model first. That is hiring a lawyer to sort parcels.
The New Structure: A Three-Drawer Cabinet
• **Drawer 1, small model**: classification, format checks, short summaries
• **Drawer 2, mid model**: first drafts, first-pass translation
• **Drawer 3, large model**: strategic judgment, final review
There is only one rule: when a task moves up a drawer, leave a one-line reason. As the reasons piled up, it became obvious which kind of work belongs in which drawer.
Results and a Failure
After a month, API cost dropped by about 40%. The surprise bonus was speed. With the small model filtering first, waiting time got shorter.
There was a failure too. When I gave first-pass translation to Drawer 1, mixing of Traditional and Simplified Chinese increased, so I moved translation up to Drawer 2. It was a useful mistake that showed me where not to economize.
Translating This into Startup Strategy
For a small team, the cost structure is decided not by 「the best tool」 but by 「the right tool in the right place」. If you split costs into tiers from the start when designing a service, costs will not grow at the same pace as your users. A learning app like StudyNest could send question sorting to a small model and reserve the large model only for worked explanations.
Today's homework: open the last month of call logs and find the three most repeated requests. At least one of them can surely move down to a smaller drawer.