Personal Learnings← Grasping Reality  Library

Grasping Reality · Economics & Policy

Spreadsheet, Not Skynet: Microdoses, Not Microprocessors

TIER 4   Tue, 24 Jun 2025 23:52:17 +0000

The modest power relative to the size of the economy of GPT LLM MAMLMs as linguistic artifacts. Of course, since the economy is super huge, a relatively modest effect on it is still huge one. But more an Excel-class innovation than an existential threat to humanity. Or am I wrong?

I am pretty confident that “AI” will, in the short- and medium-run, have a minimal impact on measured GDP and a positive but limited impact on human welfare (provided we master our attention uses of ChatGPT and company rather than finding others who wish us ill using them to enslave our attention). I just said so, earlier today: J. Bradford DeLong: MAMLM as a General Purpose Technology: The Ghost in the GDP Machine <https://braddelong.substack.com/p/mamlm-as-a-general-purpose-technology>:

But in my meatspace circles here in Berkeley and in the cyberspace circles I frequent, I find myself distinctly in the minority. Here, for example, we have a very interesting but I must regard as weird piece from the smart and thoughtful Ben Thompson:

Ben Thompson: Checking In on AI and the Big Five <https://stratechery.com/2025/checking-in-on-ai-and-the-big-five/>: ‘There is a case to be made that Meta is simply wasting money on AI: the company doesn’t have a hyperscaler business, and benefits from AI all the same. Lots of ChatGPT-generated Studio Ghibli pictures, for example, were posted on Meta properties, to Meta’s benefit. The problem… is that the question of LLM-based AI’s ultimate capabilities is still subject to such fierce debate. Zuckerberg needs to hold out the promise of superIntelligence not only to attract talent, but because if such a goal is attainable then whoever can build it won’t want to share; if it turns out that LLM-based AIs are more along the lines of the microprocessor—essential empowering technology, but not a self-contained destroyer of worlds—then that would both be better for Meta’s business and also mean that they wouldn’t need to invest in building their own…

For the sharp Ben Thompson, he has inhaled so much of the bong fumes that for him the idea that the impacts of LLM-based AIs… along the lines of the microprocessor is the low-impact case. The high-impact case is that they will become superintelligent destroyers of worlds.

This seems to me crazy.

Most important, I have seen nothing from LLM-based models so far that would lead me to classify the portals to structured and unstructured datastores they provide as more impactful than the spreadsheet, but for natural-language interface rather than for structured calculation. And I see nothing to indicate that they are or will become complex enough to be more than that. Back in the day, you fed Google keywords, and it told you what the typical internet s***poster writing about the keywords had said. Poke and tweak Google, and it might show you what an informed analyst had said. That is what ChatGPT and its cousins do, but using natural language, and thus meshing with all your mental affordances in a way that is overwhelmingly useful (because it makes it so easy and frictionless) and overwhelmingly dangerous (because it has been tuned to be so persuasive).

(Parenthetically, Thompson’s conclusion that in the low-impact case Facebook “wouldn’t need to invest in building their own” models is simply wrong. Modern advanced machine-learning models—MAMLMs—are much more than just GPT LLM ChatBots. The general category is very big-data, very high-dimension, very flexible-function classification, estimation, and prediction analysis. 80% of the current anthology-intelligence hivemind share right now is about ChatBots, and only 20% about other stuff. But the other stuff is very important to FaceBook: Very big-data, very high-dimension, very flexible-function classification, estimation, and prediction analysis is key for ad-targeting. And FaceBook needs to keep Llama competitive and give it away for free to cap the profits a small-numbers cartel of foundation-model providers can transfer from FaceBook’s to their pockets.)

Now I see the GPT LLM category of MAMLMs as limited to roughly spreadsheet—far below microprocessor—levels of impact because I see them as what I was taught in third-grade math to call “function machines”: you give them something as input, and they burble away, and then some output emerges. start with the set of all word-sequences—well, token sequences—that can be said or written.

There is a subset of that that is our training data. Our training data consists of some of those word-sequences, and then the next word that the speaker or writer sets down.

We can visualize it sort of like this:

with the reddish blob being the set of all word-sequences, and the points of light being lur training data

But now we get a prompt for the LLM. And almost all the time the exact prompt is not in the training data. The number of coherent fifteen-word English sentences is more than a thousand billion billion. And the combinatorics explode from there with longer context lengths.

The LLM needs to come up with a next word for this prompt:

So it cannot simply look in its training data and say: that is the next word.

So what does it do instead?

Well, what can it do? What it can do is look at the word-sequences in its training data that are “close” to the word-sequence that is the prompt. And then it can choose a next word (or, rather, can choose a probability distribution from which, after consulting a random number generator, it picks the next word) that is some kind of “average” of the next words from the training-data word-sequences “close” to the word-sequence that is the prompt. It “interpolates” from the training data:

Or, rather, because it is much too small to actually store the training data, from a lossy regularization, smoothing, and compression of the training data. (See Chiang (2023).)

As Shalizi (2023, 2025) puts it:

Cosma Shalizi: ‘Attention’, ‘Transformers’, in Neural Network ‘Large Language Models’ <http://bactra.org/notebooks/nn-attention-and-transformers.html>: ‘What contexts in the training data looked similar to the present context? What was the distribution of the next word in those contexts?… We can hope that similar contexts will result in similar next-word distributions…. Smoothing… will add bias to our estimate of the distribution, but reduce variance, and we can come out ahead if we do it right…. Sample from that distribution…. Drop the oldest word from the remote end of context and add the generated word to the newest end…

Then repeat.

It is in this sense that it is “autocomplete on steroids”. (It is still the case that the best comprehensible introduction to this (for me at least) that I have been pointed to is still Alessandrini & al. (2023).)

Note that from this perspective all the talk of “neural networks”, “electric brains”, “AGI—artificial general intelligence”, and so forth is (largely) a confusing distraction. Google Search with its pagerank was (initially) giving you its estimate of which webpage a typical internet user thinking about these keywords would link to. It was not “AGI”, or even the path to “AGI”. A MAMLM GPT LLM is giving you its estimate of what the typical internet s***poster who had written this particular word-sequence would write next.

Of course, it is much more than that.

For one thing, there are sets of webpages on the internet that are very "close" together and relatively "dense" and in which the view of the typical Internet s***poster is actually quite smart. Thus the machine can make and provide you with a very good estimate of what a reliable authority would write next. And it can be guided into states in which it relies upon sections of the training data that are in fact reliable, and provide substantial insight:

For another thing, the results are very, very impressive indeed across a number of dimensions. And every time I read this point, I am again unable to avoid quoting Shalizi (2023, 2025):

Cosma Shalizi: ‘Attention’, ‘Transformers’, in Neural Network ‘Large Language Models’ <http://bactra.org/notebooks/nn-attention-and-transformers.html>: ‘"It's Just Kernel Smoothing" vs. "You Can Do That with Just Kernel Smoothing!?!": [That] takes nothing away from the incredibly impressive engineering accomplishment of making the blessed thing work…. nobody [before] achieved anything like the[se] feats…. [We] put effort into understanding… precisely because the results are impressive!…

The neural network architecture here is doing some sort of complicated implicit smoothing across contexts… [that] has evolved (under the selection pressures of benchmark data sets and beating the previous state-of-the-art) to work well for text as currently found online.… Markov models for language are really old…. Nobody, so far as I know, has achieved results anywhere close to what contemporary LLMs can do. This is impressive...

How can they be so impressive?

This is a problem for me. Because they are doing things that my intuition strongly leads me to think that they should not be able to do half as well as they do.

Now don’t get me wrong. They are good summarization engines. And they have the extraordinary advantage that natural-language interfaces match my linguistic-human affordances perfectly. Still, they cannot reliably pull out author-date references from an article and construct a complete, correct reference list. By now we know, broadly, from experience, what these thing are useful for, and where they fall down. These things can be immensely useful for:

Remember:

And always, always, always: check their work. If you cannot immediately tell whether they are hallucinating or not, you have no business using them for the tasks you are attempting to do.


References:

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

#spreadsheet-not-skynet
#subturingbradbot
#dia-browser-testbed