I have a small business and a small amount of money. I am not trying to find the best model. I am trying to find the model that does each of my tasks correctly and cheaply. So I ran an experiment.
The setup
Twelve tasks I actually do in a week. Translation, code review, weekly plan, customer email rewrite, doc summarization, two SQL queries, one product copy, one short essay, one legal explanation, one Russian-language piece, one English-language piece.
Five models. Three commercial, two free / local. Each task was run three times on each model, and I scored the response on a 5-point rubric I had written before I saw the answers. Yes, I cheated a bit: I knew the rubric, the model did not. That is the point.
What the numbers said
The cheapest model per task was not the same model every time. Sometimes it was the smallest, most aggressive local model. Sometimes it was a mid-tier commercial one. Almost never was it the "smartest" model — the one that gets cited in launch posts.
The reason is simple. The smartest model is good at hard things. Most of my tasks are not hard. They are long, routine, and tolerate imperfection. On those tasks, the smart model is 20% better, but costs 30× more. The math is not close.
The exception
There was one task where the smart model was clearly worth it: anything with a long Russian-language creative brief. The mid-tier models butchered the tone. The smart model held it. For that one task, the cost was worth it. For the other eleven, the cost was a tax I was paying for no benefit.
What I changed
I now have a default: the cheapest model that I have empirically found to be good enough for the task at hand. That default changes by task. I do not use a single model as my "AI". I use a stack. The stack is documented. I look at it maybe once a quarter.
The cost drop was real. I am not going to quote the exact number, but it was meaningful. The quality drop was not measurable.
What this means for you
If you are paying for the most expensive model for every task, you are probably overpaying by 5× to 10×. The exercise is the same as mine. Pick 12 of your real tasks. Run them on 3 to 5 models. Score. Pick the cheap one for each. You will be surprised how often the cheap one wins.