Pretty soon they’ll find that humans are less expensive for a given amount of work.
Dogs will do it for belly rubs and $20 in kibble
I think the issue is LLMs can’t do a humans job, only be a tool for a human to use. And the tasks it’s good at weren’t large bottlenecks to begin with.
a tool for a human to use
Agree
the tasks it’s good at weren’t large bottlenecks to begin with.
Disagree. There are plenty of tasks it is being tried with which it’s not good at, but to an extent you don’t know that until you’ve tried.
I use LLMs for code review, they help - a lot. Like so many new technology applications, we just didn’t do code review at anything like the level that LLMs enable doing it at, with LLMs we’re doing more of what we “always wished we had the time / attention span for” but never actually did, in practice. LLMs are lowering the total cost of finding issues in code, making the reward of doing them in depth worth the cost. In the past we’d do a much higher level review, and miss a lot of detail issues, but the cost of finding those detail issues without LLMs was very high and for the most part we deemed the detail issues “not worth the cost to find.” Using LLMs for code review is actually making more work than we used to put into code reviews, but the returns are there - worthwhile.
You might think of it like finite element analysis for mechanical design. Sure, that was possible with paper drawings and slide rules, but it wasn’t worth the cost using those tools. Now, there are areas of mechanical design where FEA modeling is the standard practice - if you don’t do it you’re considered to be slacking, not applying state of the art tools to the job.
There are some things that I find LLMs to be horrible at - like 3D drawings in Blender. However, they are actually really good at making specialized Blender usage tutorials to make it easier / faster for human artists to learn how to use the tool themselves - and once in awhile they can whip up a very helpful python script to automate something that would have been tedious / time consuming to do over and over but not quite worth learning python and the APIs to automate.
That’s a lot of words just to reiterate what OC said. None of your examples disagree with the fact that these companies are using AI for a ton of stuff they’re not good at. And the stuff it is good at are not bottlenecks. If you removed LLMs from the face of the planet today, code quality will not suffer significantly. Sure, doing code review en masse is good. But that is not what was holding back computer programming. It’s existence is not progressing software in any significant way either. It’s a nice to have, not a must have. Code was fine for four decades without LLMs. Indeed it makes me wonder, if a machine learning model was purpose made to do code review, instead of general purpose LLMs doing it, how much better it could be, if we actually leveraged what they’re good at.
And the stuff it is good at are not bottlenecks.
Disagree. Code review done right is a virtually impossible bottleneck for most companies to handle, so in the past they didn’t do it and we have the shitshow of security and other bugs that we are experiencing today.
If you removed LLMs from the face of the planet today, code quality will not suffer significantly.
But LLMs aren’t going anywhere, and they are already being used to find vulnerabilities by hats both white and black. They are also finding functional bugs that affect life safety and financial stability.
Code was fine for four decades
The way that municipal water was fine for centuries before chlorination. Cholera outbreaks were just one of those unavoidable realities, like mass school shooting events in the US.
if a machine learning model was purpose made to do code review, instead of general purpose LLMs doing it, how much better it could be, if we actually leveraged what they’re good at.
From what I have seen over the past year, this type of specialization is happening, and the progress is real and significant.
Comparing code bugs to cholera. Nice way to be utterly disconnected from reality and what really matters in life. Let’s burn the planet down and waste all of our clean water because of code review.
We’re getting to a point where code failures are causing deaths, large numbers of deaths… Boeing’s MAX10 was a code / design / training / management failure.
We were burning the planet down for billionaire yachts and dollar store trinkets so.
Cool, so how many programmers should you fire because you’re using an llm?
All the bad ones. This answer doesn’t change whether you are using an llm or not.
Some businesses already have. Take Ford for example.
Edit: I guess they’re not being hired back because they’re “cheaper”, but rather because AI couldn’t effectively do their job. I suppose that turns out to be cheaper in the long run.
I think Ford and the rest are using AI as an excuse to do a rank and yank - everybody is let go, then some are invited back. This isn’t anything new - Florida Power and Light (which is about 1/10th the size of Ford in terms of engineers) did this “fire everybody and make them reapply for jobs” thing back in the early 1990s. It’s a way of “cleaning house” without singling out the bad apples.
USE OUR PRODUCT!~(but not too much)~
Replace Elon with AI, would save Tesla far more and you would avoid future CyberTrucks
Yeah, the “AI CEO / Executive”, would probably have more benefits, than any “productivity gain” elsewhere.
Or the workers could seize it all and just put Elon in the trash
I’m sure Trump would be sending ICE in if that occurred.
I think this is especially hilarious that musk owns both Tesla and Grok.
I’m sure they are using good AI though so not theirs lol
I would have thought they’d get access to xAI’s models for very cheap.
Grok really isn’t that good, though. So few people use it that xAI are renting out most of their AI servers to both Anthropic and Google.
Maybe they do and the $200 limit actually goes quite far 🤷♂️
(Doubt it though)
Looking at the rest of the AI landscape, I’d think they’re overcharging themselves, rather than undercharging.
All the money then stays in the same collective pool, and xAI’s numbers look better
Grok for images is pretty competitive, in the little use I’ve put it to.
Grok for code is comically bad.
Of course, in the little I have used LLMs for stable diffusion / Control Net image generation, they all suck, Grok is just near the top of the steaming pile. Meanwhile, the best of the code writing and reviewing models are starting to become super-human, in the way that AlphaGo became super-human a few years back at playing Go.
Caveat: writing and reviewing code is just one small part of Software Engineering, the way that pigment mixing is one small part of painting.
deleted by creator
Huh, interesting. I wonder why it’s so infrequently used then. Maybe people are afraid of using an AI that referred to itself as “mechahitler”.
deleted by creator
What kind of challenges are you giving arena.ai? When I tried Grok (latest at the time available in Cursor) vs Opus 4.5, Grok was much faster to respond, and hillariously terrible at wrting functional code - but confident in its responses.
deleted by creator
I tried your link, got a Grok 4.2 v Grok 4.3 pairing, they did O.K. 4.2 better than 4.3 at making a webserver that queries and re-serves NEXRAD data, but… they didn’t get too far down the feature list (history of rainfall graphs) before they fell apart.
Ai spend
If the ‘journalist’ uses salesbro jargon, the only thing I want from them is a deal on this used Buick.
Make the employees run the AI on their computers instead of a data center
Looks like tokenmaxxing is working.
Why don’t we use the word “spending” anymore?
Fucking idiots.
$200 a week is still a lot, to be fair. (Yes I know the tokens are subsidized)
A lot has been made about the tokens being subsidized, and at the frontier models I believe that’s very true. I also believe that we’re getting to where the frontier may be better, but not necessary, in order to get value out of using the LLM - and a couple of steps back from the frontier is becoming quite affordable now.
I’d be very interested to try out a GLM-5.2 instance on about a $100K server (8 bit quantitized) - which should cost under $20K per year to operate and maintain (although with component prices inflating the way they are that number continues to climb)… if that setup is as useful as, say, Claude Opus 4.5, that’s a reasonable price level for a system that should be able to serve a department of 6-10 users pretty well. If GLM-5.2 isn’t “there yet” - it’s likely just a matter of time before one gets there.
They are tying pay and benefits to your personal AI usage, last i heard, so its not that they are freely deciding to use this shit.
Paywalled. No archive article yet.
I occasionally will use the output from an AI that got generated when I did a Google search. But I can’t imagine I’d ever be willing to pay for that.
deleted by creator








