Three of my AI apps, each run against four Claude models at several effort levels, and scored on the app's own evaluation. The point is the trade-off: how much accuracy each extra dollar and second actually buys, so the choice of model is a measured decision rather than a default.