Put down the firehose
An exploration of building with local small models and why we probably don't need the firehose approach of frontier models for most things. Frugal AI can be done, if we choose to, and it doesn't need to cost us the quality.
Despite telling everyone recently that I'm slowing down and building less, I spent a chunk of last week on side projects. The main one is a rebuild of The Question Bank, a thing I've made and used several times over the years, very much a Data for Action concept. What's different this time is that I'm rebuilding it to be local-first: local deployment, local storage, small local AI models.
Why local and why small? Well everything I've ever written about Why we created TechFreedom, and why we think it's important and Responsible AI? are a good starting point. But I also wanted to show not just tell and this was a nice way to do it with something I know well.
Now one thing I immediately notice was that it's been slower to build because there has been much more figuring-things-out, many more constraints.
Open source models have got genuinely good. GLM 5.2 is an open source model that can compete at the very top of tasks, comparable or even beating Claude Opus and GPT 5.5. But although GLM 5.2 is open source, it's not something I can run locally. I don't have a rack of GPU's sitting there ready to roll. No, I have some limited tech (an Apple Mac Mini with 16gb), so I need to work with open source small models. And they need much more attention, and precision in how you use them. Context windows, timeouts, tooling, the shape of the prompt it all has to be considered.
With the flagship hosted models you can mostly use it like a fire hose. Throw some stuff at it and it'll probably figure out what you meant, even if that means burning tokens to get there. (There are risks in that, and I'm not pretending precision doesn't matter at the top end too.) But you don't have that luxury with a small local model. You have to be precise. You have to think the process through in much more detail.
And I've come to think that makes it better. Not just for privacy, not just for energy use, though both of those are real and they matter to me, but for precision and quality.
Just enough AI?
When I first built Bearing I built it to test a hypothesis: that much of what we actually want to use AI for can be done in a much more frugal way, through choosing models that fit the need, and in some cases using local small models. I made the case that we don't really need frontier models for a lot of things. But was this hypothesis right?
Well there's a Stanford and Together AI paper called Intelligence per Watt. They ran a million real-world queries across twenty-odd small local models (the kind with 20 billion parameters or fewer, the kind that run on an Apple M4) and asked a simple question: how much of this could actually be done locally?
The answer is a lot. As of late 2025, a single best small model could match a frontier model's quality on around 71% of queries, up from 23% in 2023 and 49% in 2024. Pool the small models together and at least one of them handles nearly 89% of single-turn chat and reasoning. Obviously there is a difference in what you want AI to do, better at creative work (writing etc) which scored 90%, down to around 68% for the genuinely technical stuff. But really the work that needs a frontier model with massive data centres is a smaller slice than the way we use these tools would suggest.
They also make an interesting point about routing, that you don't need a perfect system to create savings. A router that's right 80% of the time about whether to keep a query local or send it to the cloud captures roughly a 64% cut in energy and a 59% cut in cost with no drop in answer quality. There is so much talk of AI creating efficiency yet we seem to be missing that we should be thinking about efficiency of the AI, we just need to be a bit more deliberate in how we approach things maybe?
And what of my own data? Well as bearing has an open dataset of what people want to use AI for I took a look. Yes it's only about 300 entries, but roughly two thirds of all requests could be ran on consumer hardware. And about 1/3 could be ran on my mac mini easily.
So why is our default to just fire everything through the biggest models? The ones that cost the most, both financially, environmentally and in many case ethically?
I think possibly it's about maturity of both the market and our own skills and thinking.
For the market, well open source models are improving all the time, and how they utilise consumer hardware is also improving. Google Gemma, liquid AI are all newer models designed for running on phones, standard laptops etc. Other infrastructure improves all the time, ollama just today announcing that Gemma 4 is now up to 90% faster on Apple hardware. And when these two things come together, like they have in the last 12 months, you get ingredients that can be used, if we can figure out how to use them effectively.
And this learning about how and where to use them effectively is where we still need to mature. And it's where some of my focus is at the moment. Learning how to make best use of the ingredients available.
Which is something I'm constantly doing with food. Yes I'm making a food analogy, live with it. Cooking and baking is pretty simple right? Get some ingredients, throw them together, generally apply heat and we have something. As you progress you look for recipes that tell you certain ingredient types and quantities, and that gives you new options and ideas. But at a certain point, if we you ever want better food you need a combination of better ingredients and better skills. You need to learn different ways of apply heat, frying, oven, grill, boiling, and you need to apply different levels of heat. One of the chef tests is eggs in multiple ways, learning to control the methods and the skills.
At the minute I think we're in a place with AI where we are gathering a bunch of ingredients and throwing them in the biggest oven we can find and turning it to 11 and hoping for the best. Sometimes that will make something edible, sometimes even nice. But it's wasteful.
Taking the food analogy further to food production and growing crops we're in the spot where we know we need to water our plants, but we're just spraying everywhere, flooding irrigation. But again, we've matured here, learning that this approach is wasteful, water runs off, evaporates and only a portion of our water is used how we want it to be. We learned to deliver the water in a more directed, precise way through drip irrigation.
And so we need to develop those skills with AI. We need to experiment and learn what is possible and how. Rebuilding the Open Question Bank is one small experiment designed to explore constraints. I can only work with the models (nomic embedding and a version of qwen 3.5) that can run on that hardware...so how do I design it to run as effectively as possible within those constraints?
Careful usage, efficient calls all help. And I now have a version of the tool that is locally deployed and locally ran. I have control of cost, of privacy. In bearing i set 7 aspects to consider with AI models: costs, speed, capability, transparency, privacy, sustainability, quality. I can choose those. I'm making less compromises on the things important to me because I thought about the design, and I don't lose out on the quality. And now I have the option, should I choose to open the tap. If I have more capable hardware, I can run bigger models. Or even I want I can plug into cloud versions for frontier models if I NEED to. But I probably won't to, because we rarely do.
Yes, there will be things I can't do on a single mac mini, but I want to find the limit of what I can. And once I've found that limit, well I've already got ideas for a more mutual approach with an experiment in the works.
So, put down the firehose, be more frugal, precise, controlled, maybe pick up a water pistol and still have fun.