The Days When Everything Went to the Smartest Model
There was a time when every article, translation, and summary our Dream Team bots produced ran on the largest model. It was convenient, but one look at the month-end usage table and we quietly closed the window. So we ran a small experiment.
The Experiment: Split the Job in Two
• **Stage 1, draft**: a small model quickly builds the structure and sentences.
• **Stage 2, finish**: the big model only checks facts, tidies the tone, and verifies consistency across the four languages.
On a Windows device and a Mac mini, we ran the same 40 topics through both approaches and compared the results.
Results
• Call cost: **down 38%** over a month
• Quality: a human review found almost no difference from the big-model-only approach
• Speed: drafts arrived faster, so total waiting time also shrank
• Exception: the small model often missed Traditional vs Simplified Chinese consistency, so we made the finishing-stage big model responsible for that check
Reading It as Small-Team Strategy
1. Costs fall less from which model you pick than from how you slice the work.
2. Spend expensive judgment once, at the very end.
3. Cutting without measuring lets quality break silently. Even 40 samples are worth comparing first.
4. The smaller the team, the easier it is to treat usage cost as a design variable, and that is a real advantage.
The point is not saving money but placement. Deciding where to spend good judgment is the strategy.