Notes
Marketing Effect: The AGI
The more I use and compare these models, the harder it is for me to separate real progress from hype and marketing. At this point, I care less about AGI claims and more about whether AI can actually deliver reliable results without burning through time, tokens, and money.
If you ever heard about AGI (Artificial General Intelligence), it might have make you feel anxious about the future - or the present - we're living in. If you haven't, well, get outta here (nah, just kidding).
You know, we've heard it all multiple times before. Every new model is one step closer to AGI. Every new model is the "best one for coding ever" — disruptive, breathtaking, more capable than anything that came before it.
What I see instead, at least in my personal projects, is that newer models are increasingly becoming token consumers.
I tried GPT 6 Astra, and it was incredibly frustrating. I made a simple request using Medium effort and, in about an hour, it consumed 50% of my weekly quota — and it didn't even finish the work.
Then there are the benchmarks. The numbers they publish always show the new model way above its predecessor and — without making it too obvious — slightly better than their rival's latest and most capable model.
I don't see how that adds up.
Most of us don't have enough metrics, resources, or money to properly evaluate these models ourselves. That leaves us with only a few options, one of the most common being to give different models the same prompt and compare the outcomes.
Now, that doesn't mean these comparisons are useless. The problem, to me, is treating a single run as if it tells us which model is better. If you ran the exact same task 20, 30, or 50 times, kept the harness and settings consistent, and then compared things like success rate, cost, latency, token usage, and how often the model actually completed the task, that would tell us something much more meaningful.
But that's not usually what we see.
Instead, someone runs one prompt through Model A, gets an amazing result, runs it through Model B, gets something worse, and suddenly we have a post saying one model "destroys" the other. Run the same experiment again and the conclusion might be completely different.
And that's part of what makes all of this so hard to evaluate from the outside. The result depends not only on the model, but also on the prompt, the harness, the tools available to it, the reasoning effort, the context, and sometimes just the luck of that particular run.
So, to me, it's really hard to reconcile that with all the posts we see on social media about AGI.
At the end of the day, I'm increasingly left with the feeling that a lot of this is just marketing. AI is expensive. It requires enormous amounts of money, and investors eventually want to see what they're getting in return.
If the word associated with AI used to be "hype," I think it's slowly becoming "ROI".