The limiting factor for top models is primarily hardware.
You're right that you're not going to get Kimi K3 performance, requiring roughly 600GB of VRAM to run, from a small number of 24GB consumer GPUs. You're also not going to get closed-top-model performance, because you simply don't have any way to get the weights.
If we're talking about K2.7, Deepseek R1 Distilled, Qwen 2.5, Llama 3.2 though - All of which can be run locally -The question really becomes what you expect from them. They're all strong models that know more than we (individually) do about 99% of topics. A "high-school level education" is more than enough for most humans to get by in the world. Now imagine a "high-school level" understanding across every topic. It's not doing calculus, but it can tell you what patterns to use. It's not writing Shakespeare, but it can likely recite most of his works to you from memory. It's not going to come up with a clean-room reimplementation of Windows 11, but it can check your syntax and tell you the function prototypes of every publicly-available library in every programming language without breaking a sweat.
More importantly, to the GP's point - Is a model "smart" enough to understand you're trying to abuse it necessarily better than one that just carries out your orders successfully? Case in point, lately I've found Gemini's new "watcher" blocks a good 90% of my totally legit questions before the question ever makes it to the real model. Umm... No thanks? The best AI is the one that works, not the one that gives the deepest philosophical reason for saying no.