This is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.
wmf · 2026-08-25 21:52:05 UTC
Existing CPUs have been extremely optimized by ~6 competing, well-funded teams. I expect AI to accelerate things somewhat but it's not clear that there is any low-hanging fruit available for AI to find.
manquer · 2026-08-26 00:40:42 UTC
ASICs always do better than general purpose chips. General purpose chips is turtles and turtles of virtualization and have to consider 4+ decades of backward compatible instructions set support.
ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.
General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.
New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.
anthonypasq · 2026-08-25 18:38:23 UTC
Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.
datakan · 2026-08-25 18:45:11 UTC
Token prices coming down means nothing if the models keep wasting them
gwerbin · 2026-08-25 18:54:41 UTC
Hopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy.
Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
tmp10423288442 · 2026-08-25 19:07:15 UTC
Nah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
gwerbin · 2026-08-27 20:22:35 UTC
We already had plenty of data centers in the US before the AI boom that weren't severely harmful to their neighbors. Cutting red tape is not the same as eliminating meaningful regulation. There are plenty of old industrial sites that could be repurposed as data centers. It turns out it's cheaper to bribe some small town government to give you a tax cut and discounted electricity and water rate.
vlyan · 2026-08-25 19:28:34 UTC
>polluting on-site generators
how much pollution do you believe modern gas-turbine engines to produce?
>Not to mention the water use controversy.
what percentage of US water usage do you believe is by AI data centers?
gwerbin · 2026-08-27 20:19:00 UTC
Reducing everything to national aggregates provides no insight into the strong negative externalities, imposed on the immediate surrounding communities, of unregulated industrial facilities. That's literally the reason we have zoning laws in the first place,
dgellow · 2026-08-25 18:59:07 UTC
There is just so much downward pressure on token price, from every direction. We would need a completely new understanding of economics to explain why the price shouldn’t go down. Or market collusion/regulatory manipulation.
simianwords · 2026-08-25 19:04:16 UTC
The price has been going down for ages, its not clear what you are pointing at
phoghed · 2026-08-25 20:29:54 UTC
Pointing at the nay sayers who say tokens are heavily subsidized and it’s all going to come crashing down soon, surely any moment now
dgellow · 2026-08-25 21:27:50 UTC
I mean, it will obviously crash at some point. With so much pressure on token price to go down that means way less opportunity for margin for AI providers. OpenAI is in a pretty bad situation
simianwords · 2026-08-26 00:47:21 UTC
What does this have to do with margins? It can remain the same once prices go down
dgellow · 2026-08-25 21:25:27 UTC
At the price going down? And that it will continue to go down, even if the hardware improvements stop. Not sure what isn’t clear
dumberquestions · 2026-08-25 19:23:56 UTC
The demand for them is growing _per person_, not just across the wider economy, if tokens cost half as much but you want to use 3 times as much you're going to have to pay more.
jazzyjackson · 2026-08-25 19:30:48 UTC
Maybe 1000s of tokens per second unlocks realtime robotic decision making, and now every robot needs to continuously stream tokens to and from the cloud to operate. That could 1000x demand overnight, just to speculate :)
hypfer · 2026-08-25 19:33:48 UTC
Think about the agents buying computers for their agents. /s
jacquesm · 2026-08-25 19:35:08 UTC
I would very much like it if anything that moves with appreciable mass is governed locally just in case the link drops and/or latency suddenly goes up. Motion is very unforgiving and accidents will happen if that's not taken into account.
HDThoreaun · 2026-08-25 20:15:22 UTC
Seems unsafe to make locomotive decisions remotely
dgellow · 2026-08-25 21:30:40 UTC
I think you just found what we will see in the S-1 prospectus of OpenAI
mathisfun123 · 2026-08-25 19:02:59 UTC
this is a story about a proprietary accelerator being built/designed by a token provider. and you think they're going to return the efficiency gains to the customer instead of capture the value for themselves? interesting take.
simianwords · 2026-08-25 19:09:42 UTC
Yes, I can bet on this happening. If anything, this is a net gain for consumers as it is a competitive market.
mathisfun123 · 2026-08-25 19:25:41 UTC
go ahead and bet: alibaba is a publicly traded company
spacephysics · 2026-08-25 19:09:49 UTC
We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss
So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.
anthonypasq · 2026-08-25 19:22:07 UTC
> We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss.
what makes you think this?
RealityVoid · 2026-08-25 19:42:23 UTC
Because everyone keeps saying this so it must be true. Real "it is known" kind of vibe with these statements.
polski-g · 2026-08-25 23:10:15 UTC
He's a subscription truther. There's loads of them. OpenAI's profit increases with each subscription that is cancelled. Pretty soon they'll have more profit than God.
anthonypasq · 2026-08-25 19:21:47 UTC
OpenAI just dropped the price of Luna by 80% and Sol by 20-30%
mathisfun123 · 2026-08-25 19:24:36 UTC
and amazon shipping used to be free without prime, and uber used to be cheaper than taxis, and airbnb used to be cheaper than hotels.
you really don't get it?
simianwords · 2026-08-25 19:33:20 UTC
almost every pure tech commodity has gone down in price
- gpus
- retail computers
- laptops
- ~gpu~ appliances like washing machines
- cloud computing
i think you don't get how economy usually works in tech
zirkonit · 2026-08-25 19:38:33 UTC
I'm especially enjoying how RAM and SSDs are going down in price.
fl4regun · 2026-08-25 19:49:05 UTC
GPUs and laptops and memory and storage are all crazy expensive
thefreeman · 2026-08-25 20:01:11 UTC
listing gpu's here is crazy considering the current prices
fer · 2026-08-25 22:56:02 UTC
I was checking laptops today for an upgrade from the model I bought back in 2019 and it's not gonna happen from how cheap they are.
asveikau · 2026-08-26 01:07:43 UTC
GPUs and memory have gone up in price. It's more expensive to buy a 1-2 year old video card than it was at launch, sometimes by a shockingly large factor. Laptop vendors have recently shipped flagship models with less memory than the previous model, because they can't match price expectations for a laptop.
mathisfun123 · 2026-08-26 02:28:40 UTC
i figured out why this comment is so confusing: this is actually a message from the past, around 2020. either that or simianwords is a time traveler that arrived today and hasn't read the news yet.
ilaksh · 2026-08-25 19:12:18 UTC
Yeah but is it really even as good as Rubin? Seems just competitive.
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
kilroy123 · 2026-08-25 19:15:29 UTC
This is exactly what I see happening now.
Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.
CapsAdmin · 2026-08-26 05:20:21 UTC
Is this a normal thing now? I remember seeing this talked about as a surprising thing, but now I'm seeing posts about this as if it's normal.
(I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)
anthonypasq · 2026-08-25 19:20:50 UTC
the total cost spent on tokens may go up, but i just cant imagine per token costs going up
jrflo · 2026-08-25 19:27:39 UTC
Depends on compute capacity. If we become supply constrained on tokens, then prices will necessarily go up.
anthonypasq · 2026-08-25 19:47:26 UTC
no they dont because inference stacks are getting more efficient and models are getting more intelligent per parameter.
vlovich123 · 2026-08-26 02:47:13 UTC
I would posit there’s no way in hell they’re getting sufficiently cheaper on a short enough time frame vs how much demand is sky rocketing. AI companies are seeing quarterly doubling of revenue if not more.
sobellian · 2026-08-25 19:41:02 UTC
If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to the renaissance, mechanical work was much cheaper throughout the industrial revolution and remains cheaper to this day. We can still definitely say that the easier it is to produce tokens, the cheaper they will be.
vlovich123 · 2026-08-26 02:46:08 UTC
All Jevon’s paradox says is that as a resource becomes cheaper total consumption of that resource increases. It applies equally well to the inputs of token production as it does to the tokens themselves. The former would describe the effect the sellers into AI companies see (energy, GPU chips, RAM etc - if they lower their prices they’ll have more overall consumption) while the latter describes what the AI companies see with their customers (if they lower token prices consumers will use more tokens overall).
sobellian · 2026-08-26 20:35:43 UTC
Jevons paradox states it might increase. It is not an ironclad law and there are many many cases where increasing efficiency wrt. a certain resource really will decrease the total consumption of that resource. Yes tokens are an input themselves, but this thread is discussing hardware that is more efficient at generating tokens. To increase the efficiency by which tokens are converted into some other product would require innovation in some other area - harnesses, the models themselves, better skill from the prompters, etc.
altmanaltman · 2026-08-25 19:45:48 UTC
I think you're reducing a very complex thing (the global economy) into a very simplistic model (Jevons' paradox) and thinking both are the same thing. This has no predictive power or rigor. You're just wishing things would happen as they did before, without considering that conditions and situations change significantly, and instead of Jevon's paradox, we look back at today 50 years from now and talk about Jensen's paradox.
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
holoduke · 2026-08-25 21:09:54 UTC
That's when demand is higher than capacity. Now imagine places like Gigalab and Chinese labs are online and able to produce significant percentage of chips. That could cause real surge in prices.
cactusplant7374 · 2026-08-25 23:57:11 UTC
It is incredibly cheap now. What sectors are you thinking of?
m101 · 2026-08-25 21:59:23 UTC
With the corollary that old hardware valuations will plummet with them.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
stymaar · 2026-08-26 07:55:53 UTC
> continue to plummet.
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.
fg137 · 2026-08-26 10:20:29 UTC
Counter point: Many AWS services barely decreased their prices (if at all) in the past decade despite advancement in hardware
varispeed · 2026-08-25 18:42:35 UTC
Why they don't research how to make their own RAM and they have to buy it from the common market?
They should GTFO with this crap.
Create barriers to computing for ordinary people while milking businesses for tokens.
petcat · 2026-08-25 18:47:08 UTC
Building a custom-designed ASIC is much easier than producing state of the art memory chips.
There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.
brcmthrowaway · 2026-08-25 18:49:48 UTC
NVIDIA produces memory?
fc417fc802 · 2026-08-25 18:56:49 UTC
Fabless AFAIK. And that's the actual problem - drawing up CAD diagrams doesn't help if the factories are fully booked out.
Cyph0n · 2026-08-25 19:08:59 UTC
A state of the art GPU is much harder to design & produce at scale and than an internal ASIC.
JV00 · 2026-08-25 18:51:41 UTC
Nvidia does not make RAM
varispeed · 2026-08-25 19:33:04 UTC
That doesn't excuse them from wrecking the market for ordinary person.
chris_money202 · 2026-08-25 19:34:08 UTC
Nvidia buys the memory it uses on its GPUs, same as all other ASICs.
To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.
datakan · 2026-08-25 18:48:31 UTC
People keep saying stuff like this without understanding what it takes to make RAM. It's one of, if not the most, heavily patented things in the world. The second you dip your toes into those waters the lawsuits begin.
If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.
Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.
chris_money202 · 2026-08-25 19:24:49 UTC
RAM chips are not hard to produce compared to many other types of semiconductors; Intel started in the memory game and left because the margins weren't great and they were going to fold. The failure rates on these chips are actually very tolerable; you can have a very bad yield and still have a viable chip due to things like ECC.
datakan · 2026-08-25 19:36:44 UTC
Intel entered the memory space because they partnered with Micron. They left the memory space when Micron pulled out of the partnership.
chris_money202 · 2026-08-25 19:38:51 UTC
Intel started making DRAM in 1970, Micron was founded in 1978.
varispeed · 2026-08-25 19:34:24 UTC
Yes, it is difficult, but shafting working class is easy, therefor it is okay.
If the rich decided to buy all drinking water, you would probably be saying that's okay, making water is difficult, shortly before dying.
imtringued · 2026-08-26 08:33:22 UTC
Nah, DRAM is easy and very regular, it's a transistor and capacitor plus a massive decoder/encoder for addressing. The hard part is that it's a crushingly low margin business and without the added AI demand there were constant boom and bust cycles wiping out the manufacturers.
epistasis · 2026-08-25 18:47:53 UTC
It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
nxtfari · 2026-08-25 18:56:18 UTC
Agree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.
jacquesm · 2026-08-25 19:37:37 UTC
Ternary?
jeffbee · 2026-08-25 21:16:56 UTC
Knuth's base-e proposal enters the chat.
They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.
jacquesm · 2026-08-25 21:51:28 UTC
I can totally see how ternary would work from a physical implementation perspective but I have a really hard time visualizing anything using base-e, can you explain how such a thing would work in practice?
jeffbee · 2026-08-25 22:06:12 UTC
No it's impossible. But it would be optimal!
jacquesm · 2026-08-25 23:10:23 UTC
Ah, the spherical cow of number bases :) Thanks for the response, that saved me a sleepless night.
Razengan · 2026-08-26 11:26:23 UTC
Maybe we have to do what quaternions did for complex numbers and jump straight from 2 to 4
kroaton · 2026-08-26 08:30:16 UTC
But Bonsai is garbage.
jimmySixDOF · 2026-08-25 18:52:42 UTC
I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
This would have been far more effective with 1/10 as many words.
latchkey · 2026-08-26 15:06:26 UTC
Not really. It is a long story.
verall · 2026-08-26 19:57:08 UTC
I really really disagree. Most of the prose is devoid of information. I find it hard to believe the author even read the whole thing one time.
verall · 2026-08-26 12:29:39 UTC
I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate.
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
latchkey · 2026-08-26 15:17:01 UTC
It isn't about bunk takes or not. It is about the motivation behind doing something.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
verall · 2026-08-26 20:12:19 UTC
I dunno, it's a blog, so I'm not so worried about the motivation behind it besides how it biases their takes. I think being close to people that feed you information might be prerequisite to the kind of information he sends out.
I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.
SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?
Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.
latchkey · 2026-08-26 20:46:33 UTC
> "I'm not so worried about the motivation behind it besides how it biases their takes."
Troi oi. Read what you wrote again.
Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.
It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.
verall · 2026-08-27 11:30:07 UTC
> Troi oi. Read what you wrote again.
I can get info from and even enjoy reading a biased take. It's not so hard to see where they are coming from. In the "semicon trash talk and rumors blogosphere", you take everything with a grain of salt, etc.
> Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
I've taken this to mean that it's your post so I went through it again. Some of the points are interesting but I don't think it's a very good case that he's a grifting fraud.
I enjoy the AI images done in their style though, it's pretty funny.
> AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast.
NV has enough "best engineers" to have a couple people tuning for one of the most popular public tok/$ benchmarks without sweating. IDK about amd.
tmp10423288442 · 2026-08-25 19:05:19 UTC
SemiAnalysis’ founder was roommates with Anthropic people, not OpenAI, so he may be slightly (very slightly) more objective here.
Comments
ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.
General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.
New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.
Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
how much pollution do you believe modern gas-turbine engines to produce?
>Not to mention the water use controversy.
what percentage of US water usage do you believe is by AI data centers?
So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.
what makes you think this?
you really don't get it?
- gpus
- retail computers
- laptops
- ~gpu~ appliances like washing machines
- cloud computing
i think you don't get how economy usually works in tech
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.
(I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.
They should GTFO with this crap.
Create barriers to computing for ordinary people while milking businesses for tokens.
There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.
To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.
If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.
Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.
If the rich decided to buy all drinking water, you would probably be saying that's okay, making water is difficult, shortly before dying.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.
SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?
Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.
Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.
It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.
I can get info from and even enjoy reading a biased take. It's not so hard to see where they are coming from. In the "semicon trash talk and rumors blogosphere", you take everything with a grain of salt, etc.
> Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
I've taken this to mean that it's your post so I went through it again. Some of the points are interesting but I don't think it's a very good case that he's a grifting fraud.
I enjoy the AI images done in their style though, it's pretty funny.
> AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast.
NV has enough "best engineers" to have a couple people tuning for one of the most popular public tok/$ benchmarks without sweating. IDK about amd.