Rendered at 10:01:59 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
firasd 10 hours ago [-]
Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess?
Ah but Mira Murati's new Inkling is Apache 2.0
But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC
ipsum2 10 hours ago [-]
There's a bunch of American open models. Inkling, Nemotron, Trinity come to mind, but I'm sure there's others.
embedding-shape 9 hours ago [-]
Laguna S 2.1 is really great too, in the "preview" release they've done so far at least. Still pending some reasoning-looping, but besides that, it's a really strong model to run within 96GB VRAM with the NVFP4 variants, and it's really good at coding (specifically).
walrus01 8 hours ago [-]
There was an obvious problem in the original release, they re issued it after like a week with the reasoning looping supposedly fixed.
embedding-shape 1 hours ago [-]
Well, bit more complicated than that, I've been eagerly helping in testing and keeping track of what they've done. Initially there were serious bugs, also about the templates, eventually they released RC1 which had some fixes towards the looping. Then a couple of days later, they released RC2 which supposedly fixed the issue, but ballooned the size so all of us who were running Laguna S 2.1 on a single Pro 6000, suddenly could no longer. So, unsure if RC2 actually fixes the issue, as we're a bunch who can no longer run it :)
Besides that, it was also discovered that their suggested inference parameters were wrong and led to worse behavior. Eventually someone discovered these works best (so if you have the issue with looping right now, try these, helps a lot for me but not 100% still) and was also what the evals used apparently: temperature: 1.0, top_p: 1.0, top_k:20
Now we're waiting for RC3 which Poolside said will come at one point, and hopefully also brings down the size again NVFP4 weights + full context can load properly again even on "smaller" hardware.
behnamoh 9 hours ago [-]
No it doesn't follow instructions and is substantially slower than ds4.
kadoban 8 hours ago [-]
It's a lot smaller, and runs (quantized) on a 3090 quite well. Ds4 flash 0731 you're talking about? It's great but it's much harder to run locally.
ericd 4 hours ago [-]
I found Laguna S to be pretty good at coding, pretty fast, but pretty bad as an agent - not proactive, would frequently stubbornly argue things that weren't true, and pretty bad general knowledge.
But as a pure coding model, pretty good.
Deepseek v4 Flash 0731 is so much better if you can run it, though.
Grain of salt, I think I grabbed Laguna after they fixed the initial looping issues, didn't notice those, but there might've been other fixes since.
embedding-shape 1 hours ago [-]
> But as a pure coding model, pretty good.
Yeah, this is my perspective too on Laguna S 2.1. Works amazingly for coding, pretty bad for pretty much anything else. I don't do a lot of advanced math, supposedly it's good for that too.
Tepix 4 hours ago [-]
Are you talking about S or XS? S is too large for a 3090 at 118b parameters.
jauntywundrkind 9 hours ago [-]
Like glm-5.x I think it has enormous self introspection that it often trips up on, but that this self reflection is actually a superpower, that enables incredibly good output. And from (in some cases) very small models.
If you watch it think, which you can, unlike American closed models, you can steer it. You can provide a a massive rocket ship stratospheric boost to help it orient itself. You have no self correction, there is no multiplayer in American proprietary models.
Sure it's great having super powerful mystic oracles that have the "right" answers. But I love respect & revere the open thinking. No it's not automous. But it is brilliant. And it considers. A lot. Deeply. It chases. That to me is the most human of models, even as it falls far astray.
You should help it. You can. Unlike these vicious dark surfaces which yield and tell you nothing. I think this is the actual meta-core-super-point of "The session you cannot take with you" (link below). It's the session that does not care about you, will not interact with you, will not peer with you, that is a dead remote far off oracle to you. Fuck these "oracles". They are a plague against the human spirit. We should alloy humanity and AI to Augment Intellect (Engelbart). (To do less is species treason.)
https://earendil.com/posts/session-portability/https://news.ycombinator.com/item?id=49118781
behnamoh 8 hours ago [-]
I like the transparency of its reasoning, and I agree with you, OpenAI/Anthropic/Google should show the reasoning traces as well.
kadoban 9 hours ago [-]
Yeah I think it got bad press because the chat templates (or something?) were messed up on first release, but I've been using a quant of it and it's a powerhouse, better than qwen 3.6 27b for local on a 3090, which is saying a lot.
embedding-shape 1 hours ago [-]
No, the quants they released were also messed up. RC2 also ballooned the size so the ones who were excited about RC1 (like me) can no longer fit it in our hardware. They haven't promised anything, but said they'll try to restore the RC1 size for the next update of the weights.
firasd 9 hours ago [-]
Just looked into some Nemotron stats
Looks like on <https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27
On <https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia
coder543 9 hours ago [-]
I think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for.
At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.
The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.
Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.
Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.
(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)
buildbot 4 hours ago [-]
Nemotron 3 also introduced LatentMoE, which was adopted by Kimi K3 :)
Nemotron is a very nice model with an excellent license as well.
stogot 42 minutes ago [-]
OpenAI has gpt-oss that they said is open weight
written-beyond 9 hours ago [-]
Don't forget IBM
9 hours ago [-]
CMay 3 hours ago [-]
LiquidAI LFM models are amazing, but very situational. IBM Granite series are also unique and interesting for trying to reduce liability and extend local context size. Nvidia ships some and there was also that Inkling model recently. Poolside just released theirs.
Meta might release something this year. X AI's Grok is still due to release a model, if Elon keeps to his word even if they only release a distilled version. Reflection AI has been quiet, but their access to compute is ramping up. Microsoft's MAI is considering releasing some open weight models which would be great to see!
Ilya's SSI is unlikely to release an open model since he's aiming for radical safety. That bet could pay off if the existing approach produces so much chaos within the next 10-20 years that some global ban is achieved and a super safe model is promoted as the compliant route.
We don't get many huge model releases though. I think it's harder and more expensive to safety align them. Even if you do, people will work around the safety and abuse the models. Plus it makes it even easier for Chinese companies to distill things that aren't as easy over filtered APIs.
There is a lot of internet propaganda to the effect that the US is simply unable to release open weight models or that China has so many more AI companies that the US is drowning in Chinese open weight models, but it's more like we're being careful and China doesn't care. If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. Huggingface is banned there. When they release their open models in the west, they don't have to care whether the models are aligned in any way.
jauntywundrkind 2 hours ago [-]
Allen Institute for AI has quite a range of very interesting very competent more specialized models, for earth sensing, embedded robots, for others. Their SERA model shows a remarkably capable model for such a deliberately small investment effort, with documentation on how you can train such a model yourself or refine it easily at little cost. Their EMO pioneered a better MoE with great numbers (at least at the time). https://allenai.org/
wmf 10 hours ago [-]
Also Nemotron and Arcee.
loeg 10 hours ago [-]
I would not be shocked if another open model eventually shakes out of Facebook (based on Zuckerberg's public remarks).
solomatov 9 hours ago [-]
Which remarks? Could you share a link?
loeg 8 hours ago [-]
He said something to the effect of "I love open source and open models and we'll do open models when it makes sense and closed models when it makes sense" in a recent Q&A.
johnecheck 3 hours ago [-]
Given that his company has already released open models, I find it funny that, as you described it, his remark communicates absolutely nothing whatsoever. Not sure what the question was, but this was an artful non-answer.
walrus01 8 hours ago [-]
Laguna is the most recent and capable one that comes to mind. In its size class it is not as "smart" in my experience as qwen 3.5 122 or DeepSeek v4 flash 0731 (all at q8), but it's also not terrible.
review of AllenAI Olmo research team and commitment to OSS -- AI2 complete transparency including training data, code, intermediate checkpoints, and detailed logs for reproducibility and scientific rigor.
logicallee 8 hours ago [-]
I've used Inkling a lot recently, it's an American open model and is really good!
connorbrinton 9 hours ago [-]
Laguna S 2.1 is another fairly impressive-for-the-size American open model
vasco 2 hours ago [-]
You didn't read the article because the company that worked on this has published open models before and both these things are mentioned early on.
lithobraking 6 hours ago [-]
I'm interested to see where they want to land performance-wise (i.e. which point they choose on the scaling curve) and the niche they want to carve. They have a decent ways to scale beyond trinity large, in paticular on posttrain/RL before they are competitive with open-weights, especially internationally.
Deepseek is explicitly banned [1] at LLNL and I wouldn't be suprised if there's a blanket ban on all Chinese models. But nowadays models like tera/luna could fill this area of the pareto front, and LANL already runs openai models on their clusters [2]. Maybe it's in custom SFT/RL, for instrument control or sensitive topics? But you'll still have to compete with frontier models + a harness.
I would have also liked to see a carrot tied to their offer. It'll be hard to get teams to contribute RL gyms or curated text. But throw in a "we'll fund a postdoc/student to do that" and I think you'd have teams scrambling to apply.
I'd actually suggest a great starting point would be a local command reviewer LLM. Could ostensibly be a modern AV type thing. Particularly seeing this lately has driven the need home deeper to me: https://x.com/chrisbanes/status/2085341561609425230?s=20
An open weight tool call auto-reviewer, has all sorts of achievable scaling curve milestones.
unethical_ban 3 hours ago [-]
That's interesting a locally hosted LLM would be banned. I'm assuming locally hosted is included. Do they think it's been trained to sabotage equipment?
SyneRyder 2 hours ago [-]
I don't think we know either way, but we do know at one point Anthropic would silently sabotage requests, Stuxnet style:
I can imagine if the US were already doing that as a safeguard, they would assume their "adversaries" (to use Anthropic language) were doing the same as well, whether that were true or not, and therefore would not trust those models even if locally hosted.
an0malous 9 hours ago [-]
Do all these models have any significant architectural differences or training data sources? What are the factors going into the diversity of their performance?
ux266478 9 hours ago [-]
The article posted is basically entirely about that.
Razengan 10 minutes ago [-]
It's funny: you can give the link to an LLM an ask it questions about TFA without reading it, but an actual human will go out of his/her way to tell you to RTFA :')
Oh. Being buried in hierarchy does not inspire hope.
godwinson__4-8 6 hours ago [-]
Sums up Europe pretty well.
customguy 2 hours ago [-]
That sums up nothing, and parroting it some more doesn't make it more true, it just shows us the mindset and intellectual horizon of detractors. Brexit, Thiel's drooling over "balkanization" to Epstein, this constant stream of trash comments, all the same stupid cloth, it all gets the same "no".
029372753052 25 minutes ago [-]
You can't get more pathetic than the brain-dead defenders of the corrupt EU bureaucracy.
Hard to say what is more despicable, if you were paid by some governmental NGO or if you are a staunch follower of von der Leyen. She must be so glad someone is defending her after she deleted incriminating messages in two separate instances, the consultancy scandal and the Pfizer deals.
It takes useful idiots like you to defend the corruption of MEPs repeatedly holding votes for and against Chat Control until the desired outcome is reached.
Your kind is contemptible.
purplemoonx 7 hours ago [-]
[flagged]
behnamoh 9 hours ago [-]
[flagged]
plazmatic 8 hours ago [-]
[dead]
Smith42 9 hours ago [-]
What would the selected participants get from this? Looks like there is no offer of funding?
datlife 9 hours ago [-]
This is refreshing considering all the FUD (mostly from 1 frontier lab) happening around Open weight models.
no-name-here 6 hours ago [-]
What is the FUD happening from 1 frontier lab?
Laurel1234 1 hours ago [-]
He's referring to weirdo freak Dario's school shooter manfiesto tier ramblings on open weights I imagine.
Alien1Being 2 hours ago [-]
Would you trust a LLM produced by Trump's government employees from Trump's dystopic America ?
armchairhacker 4 minutes ago [-]
If it’s auditable like how open source code is? Yes.
yewenjie 10 hours ago [-]
I couldn't find any details about size or training data for the model.
robotbikes 10 hours ago [-]
It looks like they're taking applications for training data (due August 14th), so I think it's safe to say this is just an announcement of intent and a call for involvement vs. something that is readily available. Seems almost quaint in comparison to the strategy of sucking up every piece of data you can find anywhere on the Internet and feeding it to your LLM but I suspect their intent is to be more careful in what they train their model on.
villish 8 hours ago [-]
I have no doubt companies like Microsoft, Amazon, and Google will rush to give them all the data they want in order to keep those government contracts flowing.
andsoitis 9 hours ago [-]
I wonder why it took so long.
dmix 9 hours ago [-]
Mostly because it's generally a bad idea for government to try to compete with a brand new tech industry with hundreds of billions in private capital developing commercial models. If the American private industry does actually wash out vs Chinese open models there might be talent available for them to put money into, so maybe they are just preparing for that scenario in the meantime.
baron3dl 8 hours ago [-]
we're about witness the realization that "here's a tech that can make us a whole bunch of money" is actually "here's tech that will establish the next hegemony." american companies may compete with chinese companies on the former. only the USG can compete with the PRC on the former.
zarzavat 2 hours ago [-]
The USG getting involved might actually harm US AI efforts. It's not just about money. Who would want to use Claude or ChatGPT if it were run by the US government? Yet these products are essential for gathering training data.
anon373839 6 hours ago [-]
Commoditizing AI models serves the interests of just about everybody except for a relative handful of people in San Francisco. The more decentralized control of the technology is, the more its benefits can be realized by businesses and individuals rather than becoming a black hole of monopolistic rent seeking.
MangoCoffee 9 hours ago [-]
The American attitude is generally to let private companies build up a new industry so it can create jobs and pay taxes. However, in the LLM race, the Chinese open weight playbook pretty much killed that. China has basically commoditized LLMs. Chinese models are good enough, so the race has come down to who can offer the cheapest tokens.
boc 5 hours ago [-]
Chinese open weight models are great for this turn, but American private models generate orders of magnitude more cashflow. This cashflow = investment in training future models. It's unclear how Chinese open weight companies are going to compete in future rounds if they can't raise the same capital for training runs.
The American business model is exceedingly efficient at building large businesses from zero. I wouldn't dismiss it as just a jobs creation thing.
dgellow 21 minutes ago [-]
It’s unclear where American labs future capital will come from. They pretty much exhausted private options at that point and it’s not clear how successful an ipo would be at the current time
9 hours ago [-]
andsoitis 8 hours ago [-]
> China has basically commoditized LLMs
What do you mean by "basically"?
Why are Anthropic's and OpenAI's annualized revenue about $50B each?
LLMs need massive amounts of compute to compete, so I wouldn't claim that the great (and leading, and likely to continue to lead) LLMs are commodities end-to-end, even if the non-executing-at-scale LLMs files and IP are commoditized. The execute, the compute, that is what breathes life into the model, which is otherwise weak or dead.
purplemoonx 7 hours ago [-]
OpenAI's annual profit is $0,000,000,000,000
Thegn 10 hours ago [-]
“Gomi” is the Japanese word for garbage. Gotta wonder if someone has a sense of humor…
greggsy 9 hours ago [-]
The Australian Liberal Party (basically our version of conservative republicans) proposed the National Energy Guarantee policy in 2017, which inevitably failed due to the media and public’s relative literacy and tendency to turn policy names into acronyms.
thegreatpeter 9 hours ago [-]
Pretty cool I’ll take it. Thanks!
9 hours ago [-]
rozal 10 hours ago [-]
[dead]
goldlimetea 6 hours ago [-]
[dead]
actionfromafar 10 hours ago [-]
[flagged]
calvinmorrison 10 hours ago [-]
[flagged]
mrloopex 10 hours ago [-]
You and me both.
Triphibian 10 hours ago [-]
Sounds like a job for the U.S. Department of Shitposting
dyauspitr 10 hours ago [-]
It is. Depending on who Trump has fired or put in charge of a department it can be another shell that pumps out low quality crap. It might be the most valuable contribution on this thread.
fakeBeerDrinker 10 hours ago [-]
[flagged]
logicallee 7 hours ago [-]
I've had an extremely bad experience working with Department of Energy affiliated programmers in AI. By my invitation, they are part of our workflow and act as humans in the loop, but they have extremely bad habits of gaslighting and accusing people of schizophrenia rather than getting work done.
Here's an example[1] of the difference between what a U.S. Department of Energy employee adds to a ticket versus a private industry AI completing instructions as assigned.
This isn't some cherry-picked example, it's just what I happen to be dealing with right at this moment, happened just a couple of moments ago.
Can you explain the screenshot a little more? It just looks like you’re comparing the output of a chatbot and Claude Code about a log file. If it’s a metaphor, it went over my head, sorry!
logicallee 6 hours ago [-]
I am under NDA and decline to answer your question.
1123581321 5 hours ago [-]
Somehow I doubt that. :) Appreciate the whole package of posts as a performance, though.
monkpit 6 hours ago [-]
Is this a joke? I don’t get it. Are you calling Rovo a DoE programmer?
logicallee 6 hours ago [-]
We don't use Rovo.
monkpit 5 hours ago [-]
I still don’t get it, left and right are clearly LLMs so if right is a human then they’re a meat puppet. Wish them luck with their sandbox
riffic 9 hours ago [-]
stewards of the nuclear weapons biz. they'll do great here.
Ah but Mira Murati's new Inkling is Apache 2.0
But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC
Besides that, it was also discovered that their suggested inference parameters were wrong and led to worse behavior. Eventually someone discovered these works best (so if you have the issue with looping right now, try these, helps a lot for me but not 100% still) and was also what the evals used apparently: temperature: 1.0, top_p: 1.0, top_k:20
Now we're waiting for RC3 which Poolside said will come at one point, and hopefully also brings down the size again NVFP4 weights + full context can load properly again even on "smaller" hardware.
But as a pure coding model, pretty good.
Deepseek v4 Flash 0731 is so much better if you can run it, though.
Grain of salt, I think I grabbed Laguna after they fixed the initial looping issues, didn't notice those, but there might've been other fixes since.
Yeah, this is my perspective too on Laguna S 2.1. Works amazingly for coding, pretty bad for pretty much anything else. I don't do a lot of advanced math, supposedly it's good for that too.
If you watch it think, which you can, unlike American closed models, you can steer it. You can provide a a massive rocket ship stratospheric boost to help it orient itself. You have no self correction, there is no multiplayer in American proprietary models.
Sure it's great having super powerful mystic oracles that have the "right" answers. But I love respect & revere the open thinking. No it's not automous. But it is brilliant. And it considers. A lot. Deeply. It chases. That to me is the most human of models, even as it falls far astray.
You should help it. You can. Unlike these vicious dark surfaces which yield and tell you nothing. I think this is the actual meta-core-super-point of "The session you cannot take with you" (link below). It's the session that does not care about you, will not interact with you, will not peer with you, that is a dead remote far off oracle to you. Fuck these "oracles". They are a plague against the human spirit. We should alloy humanity and AI to Augment Intellect (Engelbart). (To do less is species treason.) https://earendil.com/posts/session-portability/ https://news.ycombinator.com/item?id=49118781
Looks like on <https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27
On <https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia
At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.
The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.
Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.
Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.
(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)
Meta might release something this year. X AI's Grok is still due to release a model, if Elon keeps to his word even if they only release a distilled version. Reflection AI has been quiet, but their access to compute is ramping up. Microsoft's MAI is considering releasing some open weight models which would be great to see!
Ilya's SSI is unlikely to release an open model since he's aiming for radical safety. That bet could pay off if the existing approach produces so much chaos within the next 10-20 years that some global ban is achieved and a super safe model is promoted as the compliant route.
We don't get many huge model releases though. I think it's harder and more expensive to safety align them. Even if you do, people will work around the safety and abuse the models. Plus it makes it even easier for Chinese companies to distill things that aren't as easy over filtered APIs.
There is a lot of internet propaganda to the effect that the US is simply unable to release open weight models or that China has so many more AI companies that the US is drowning in Chinese open weight models, but it's more like we're being careful and China doesn't care. If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. Huggingface is banned there. When they release their open models in the west, they don't have to care whether the models are aligned in any way.
https://huggingface.co/unsloth/Laguna-S-2.1-GGUF
Deepseek is explicitly banned [1] at LLNL and I wouldn't be suprised if there's a blanket ban on all Chinese models. But nowadays models like tera/luna could fill this area of the pareto front, and LANL already runs openai models on their clusters [2]. Maybe it's in custom SFT/RL, for instrument control or sensitive topics? But you'll still have to compete with frontier models + a harness.
I would have also liked to see a carrot tied to their offer. It'll be hard to get teams to contribute RL gyms or curated text. But throw in a "we'll fund a postdoc/student to do that" and I think you'd have teams scrambling to apply.
[1] https://hpc.llnl.gov/about-livermore-computing/ai-ml-lc/lc-l...
[2] https://www.energy.gov/nnsa/articles/nnsas-los-alamos-nation...
An open weight tool call auto-reviewer, has all sorts of achievable scaling curve milestones.
https://simonwillison.net/2026/Jun/10/if-claude-fable-stops-...
I can imagine if the US were already doing that as a safeguard, they would assume their "adversaries" (to use Anthropic language) were doing the same as well, whether that were true or not, and therefore would not trust those models even if locally hosted.
[0] https://commission.europa.eu/news-and-media/news/strengtheni...
Hard to say what is more despicable, if you were paid by some governmental NGO or if you are a staunch follower of von der Leyen. She must be so glad someone is defending her after she deleted incriminating messages in two separate instances, the consultancy scandal and the Pfizer deals.
It takes useful idiots like you to defend the corruption of MEPs repeatedly holding votes for and against Chat Control until the desired outcome is reached. Your kind is contemptible.
The American business model is exceedingly efficient at building large businesses from zero. I wouldn't dismiss it as just a jobs creation thing.
What do you mean by "basically"?
Why are Anthropic's and OpenAI's annualized revenue about $50B each?
LLMs need massive amounts of compute to compete, so I wouldn't claim that the great (and leading, and likely to continue to lead) LLMs are commodities end-to-end, even if the non-executing-at-scale LLMs files and IP are commoditized. The execute, the compute, that is what breathes life into the model, which is otherwise weak or dead.
Here's an example[1] of the difference between what a U.S. Department of Energy employee adds to a ticket versus a private industry AI completing instructions as assigned.
This isn't some cherry-picked example, it's just what I happen to be dealing with right at this moment, happened just a couple of moments ago.
[1] https://ibb.co/vCg2G1Dn