• 2 Posts
  • 339 Comments
Joined 3 years ago
cake
Cake day: July 5th, 2023

help-circle
  • But it isn’t encoding knowledge, it’s encoding word correlations.

    I’m saying that humans do this a lot, too. Qualitatively, it’s different, in that this particular batch of frontier LLMs will get things wrong in ways that most human brains wouldn’t, but as a category of error it’s not unique to LLMs.

    I know a ton of facts that I learned only through reading, and have no actual firsthand knowledge/experience or ability to test it: Jupiter is larger than Saturn, the atmosphere during the Carboniferous period was high in oxygen, cigarettes cause cancer, Thomas Jefferson owned slaves, the capital of Norway is Oslo. At best, I can cross reference other sources and see that things are consistent with each other. Is my belief in those facts “knowledge,” or is it merely recognizing from my training data that those particular words can validly be presented in that order?

    If you ask average people on the street whether FAT32 is a good filesystem for a 64GB removable drive, most of them won’t know, but there are a handful of bullshitters who might confidently parrot back things they can Google but not understand. That’s part of the human condition, too.

    I’m by no means an AI booster/enthusiast. I suspect LLMs/transformers are actually a dead end, and expect the upcoming crash to be economically and financially devastating to the tech and financial sectors. But I also have a pretty dim view of human intelligence, too, and see way too many parallels in LLMs as bullshit artists to humans as bullshit artists, too.


  • It’s that they are trying to use statistics to encode entire thought processes into hidden variables from conversation snippets. They want to use statistics to go from many individual interactions to a large model, and then use that model to predict individual interactions again.

    Has it been shown that the human brain doesn’t model the world in a similar way, though? A huge portion of human knowledge is both stored and transmitted in the form of language. Lots of human knowledge also follows the garbage in, garbage out theory, where you can have entire areas of knowledge that aren’t actually true but might be internally consistent, at least within certain scopes: conspiracy theories, belief in the supernatural, entire academic disciplines built on a religion or theology that not everyone believes, etc. Or even world building in fiction, the words on a page can be enough to convey ideas such that it “tricks” human brains into filling in the gaps so that they internally see a rich, fleshed out world that is entirely fictional and where specific details might not find strong direct support in the underlying text.

    it has no concept of correctness

    But statistical weight on what is more or less likely to be correct still makes a difference to objective quality of the outputs. If the model weights are trained on the reality that high quality university texts describe something and reflect some sort of underlying model of what is described using language, then can’t the model itself learn as much as a human could from those words on a page?

    All models are wrong, but some can be useful. And different models have different quality in different domains. So although I don’t believe LLMs will overtake the hump of getting ahead of human knowledge, I also don’t believe that any given LLM can be evaluated on quality, and that Facebook’s LLMs are significantly behind other LLMs we see.

    And that maybe a huge part of it is its internal process of preparing the model to evaluate the quality of its inputs, such that the output it produces can also score high on quality.







  • Yes. But major differences:

    The dot com buildout of physical communications infrastructure involved basically 3 things:

    1. Switches/routers at the nodes for sending signals down the right route.
    2. Fiber optic cables connecting the nodes.
    3. Legal rights of way and easements for the legal right to keep the physical assets in that physical place, and to maintain/replace the stuff as needed.

    Category number 1? That stuff went obsolete quickly, and wasn’t really reused after the crash.

    Category number 2 was better. Turns out, fiber optics can carry signals on a lot more channels than those fibers were originally designed for. And they’re designed for useful lives measured in decades. So even if they sat dark from being unused for 5-10 years, eventually they could be used again.

    Category 3 is super important. That legal right is basically permanent, and so long as communications equipment needs to physically go from one place to another, having that legal right can be built on and profited on (including the ability to sell or lease those rights).

    What’s that gonna look like for the AI infrastructure? The servers full of GPUs are the bulk of the cost, and the GPUs are replaced with a new generation every 1-2 years, seem to require all new power and cooling infrastructure every 1-2 generations or so.

    Plus the AI buildout looks to be several trillion dollars. Even adjusting for inflation, that’s so much more than the tens of billions that each telecom company built out that infrastructure.

    And it’s hard to see how the servers themselves will be useful for regular businesses, much less consumers. A Blackwell 72-GPU server is $3 million and takes 130 kW to run. A residential electrical line maxes out at about 48kW. The newest Vera Rubin servers are projected to be up to 600kW, with all the power and cooling management that comes with that, plus all the ultra high end networking stuff built into that rack. Even deep pocketed businesses will have trouble finding a use for that server rack worth millions, requiring a ton of supporting infrastructure that not even normal pre-2025 data centers have.










  • AI has an interesting economic trait in that it’s very, very expensive to deploy, and made very fast progress from 2022 to 2024. That caused investors with money to believe that:

    • Pushing the frontier was going to cost a lot of money. More than any other purported revolutionary tech.
    • Extrapolation of past improvement meant that whoever was on the cutting edge may end up with a product with a huge paying market.
    • So whoever wins this race would be rich, and the investment would have been worth it for them.

    But since 2024, we’ve seen that the cutting edge got even more expensive much faster than expected, and much of the improvements in performance now come from inference rather than training, which represents a high ongoing cost.

    Now, if we extrapolate from that trend line, we’ll see that the market will be much smaller for AI services at the cost it takes to provide that service, and the question then becomes whether the industry can make its operations cheaper, fast enough to profitably provide a service people will pay for.

    I have my doubts they’ll succeed, and we might just be looking at the industry like supersonic flight: conceptually interesting, technically feasible, but just a commercial dead end because it’s too expensive.



  • Counterpoint: sometimes the best still shot requires a particular moment captured with a particular, consciously arranged setup.

    This interview of a veteran NBA photographer breaks it down of how he only has a single shot per shot because of how he necessarily relies on strobes set up to not distract the players or interfere with the broadcast. As a result, he scouts/studies each player and team so that he knows when the right moment is to actually capture the shot, because he can’t exactly ask players to do it again.

    If you read interviews of Pulitzer photography winners, they’ll often say a lot of the same things: being prepared and being lucky and having that convergence of having incredibly high skill/expertise/understanding of the setting, while being able to capture in every opportunity presented.

    You should capture a lot of photos and examine them to understand how to make them better, and increase your skill level and understand your subject so that you can still optimize for the very best shot possible.


  • Transformers are like blockchain: an interesting use of mathematical principles to solve certain problems in a novel way, where the hype around that core attracts charlatans and scammers and combinations of the two traits who claim that it will go on to solve totally different problems in such a way as to revolutionize the world we live in.

    NFTs were the end of that line for blockchain where the machine started to eat itself. I can see a future, stable use of blockchain in some limited contexts, but cryptobros have always overstated the contexts in which that particular type of digital ledger can be more useful than other types of digital ledgers.

    We’ll see where the end of the road is for transformers, and what’s left at the end. I believe that computer inference will always be useful in some contexts, and that the advances in huge models with absurdly large numbers of parameters have unlocked some previously impractical tasks, but I could also see that settling into a general background existence as just another technological tool for doing things in a world that still looks pretty similar to the world today.