First of all, download is not a user.
Every so often, the local AI world gets a number that feels bigger than the scene itself.
There are roughly 1 million downloads for a Qwen 3.8 27B model and while my co-worker was talking about it, I asked him if he runs that model on his local machine and he responded that he did download it, but his machine wasn’t as powerful as he thought it is.
The question is not really about one model.
And so I did a quick survey at my workplace.
3/40 in the engineering team.
1/18 in the marketing team.
Most people said: It was about the gap between local AI as a visible internet movement and local AI as a thing people actually use on their own hardware.
That gap is real tho.
A download count can include curiosity clicks, duplicate downloads, quantized versions, failed experiments, mirrors, forks, scripts, cloud machines, and people who downloaded a model once and never used it again.
It can also miss people who use a model through a hosted interface, a community quant, a torrent, a cache, or a private setup. A download count is a sign of motion.
It is not a census.
I asked and
Most people at home still have 4 to 8GB machines and are mostly lurking and reading.
while
Some have dual 3090s.
Some have Mac Studios with large unified memory.
Some rent GPUs from services like Vast.ai.
Some even did have enough VRAM to make my original question look timid.
But not most people.
A few were not even interested in the new model yet because they had work, benchmarks, or better things to do.
That mix says more than the download number.
Local AI is not one market.
It is several overlapping habits:
hobby,
infrastructure project,
privacy choice,
productivity setup,
hardware obsession,
research workflow,
and weekend entertainment.
In fact, the question “how many people are really using this?” only gets useful after we ask what “using” means.
But, don’t people run models in ways that do not fit a clean GPU-size bucket?
They do. That’s right.
Some people split memory across multiple cards.
Some run quantized versions on 16GB cards.
Some use Apple silicon with unified memory, where the old desktop GPU language does not map cleanly.
Some run on CPU or system RAM at painful speeds because they only want to test the model.
And as I pointed above, some rent hardware when they need it.
And this is where the most download count came from in the first place.
The problem is that local AI asks the user to care about quantization formats, memory pressure, inference backends, GPUs, Macs, CUDA, Vulkan, context windows, and whether their case can physically hold another card.
That makes the crowd smaller.
It also makes the crowd unusually committed.
Okay, now let me tell you one more thing:
Local does not always mean sitting ON your desk all the time.
I understand model fatigue where you spend weeks testing a model and it turns out to be shit.
Claude or Chat GPT could have done that thing in minutes.
But, actually,
The idea of being “pseudo-local” is the best.
Or even so as renting from Vast.ai so you could scale up or down instead of owning every GPU you might need.
I know it also raises an awkward but useful question:
what counts as local AI?
For some people, local means the model runs on hardware they own, in their room, disconnected from a provider.
For others, it means open weights, controllable inference, and a workflow that does not depend on a closed chatbot.
A rented GPU does not satisfy the strict privacy definition, but it does satisfy some of the practical and philosophical ones.
You choose the model.
You choose the runtime.
You can change the stack.
You choose money.
That middle ground will probably matter more over time.
The fully local dream is powerful, but it is expensive.
A good GPU costs real money. Multi-GPU setups add complexity. Large unified-memory Macs are convenient but not cheap.
Older server cards can be clever bargains, but they bring their own noise, power, and setup issues.
Renting hardware lets people experiment without buying into the whole lifestyle.
The better question is not “how many people have the hardware?”
It is “how many people have a repeatable workflow where local or self-directed AI is worth the friction?”
That group is smaller than the hype suggests.
It is also bigger than a strict hardware census would imply.
In case we are meeting for the first time, come over here, it’ll be worth the roller coaster of articles that are gonna come up in the next few weeks.
I run a bunch of apps at AIBucket.
I do not use AI in my writings and you shouldn’t either. So, How did I go from 0 to 1000 here?