Tag Archives: AI

Thinking in Five Dimensions using AI

Eric Betzig shared the 2014 Nobel Prize in Chemistry for super-resolved fluorescence microscopy. I’ve seen countless images resulting from his microscopes over the years, of course, but I had no idea who he was or anything about the significance of the technology he had been working on. After watching his recent interview on the 632nm Podcast, though, I’m totally hooked:

Live-Cell Imaging and the Limits of Structural Biology, Eric Betzig on Super-Resolution Microscopy, 9/2026

Much of the conversation, which runs almost three hours, covers in great detail his career in physics and engineering. He loves building things, that’s for sure. And he’s always looking around hoping to find and use new technology to create his stuff. He did near-field microscopy at Bell Labs, but then he quit science in utter frustration in 1995 and spent six years building advanced machine tools at his father’s company in Michigan. Then during two stretches of unemployment in the early 2000s, he and his physicist friend Harald Hess built the first PALM (Photoactivated Localization Microscopy) in Hess’s living room using their own cash. That work later earned Betzig the Nobel Prize. That itself is wild. I guess it pays to take some time off to hack around on some side projects, eh?

“By far the best periods of my life were my two periods of unemployment,” Betzig says.

When Betzig talks about machine tools, though, he’s not just poking around the factory floor banging on sheet metal. Instead, he was designing and building totally new machines and developing concepts to solve specific problems others hadn’t even attempted to crack. It’s also clear from his language that he sees obvious connections between disciplines that others can’t imagine. It just makes sense to him to draw from as many disparate pools of expertise as possible. That characteristic of deep thinking is common among people like Betzig.

What he did with the method he implemented with his friend matters much more to him than the Nobel Prize. In fact, he hates the culture that’s pervasive throughout science where researchers chase prestige via awards and prizes. He says it’s crazy that one person gets the credit for some bit that resulted from the work of many other people over decades of time. And it also probably helps explain why a physicist who says he never wanted anything to do with AI is now trying to build the largest vision model in the history of biology using AI.

“It’s my firm belief that you cannot understand life without looking at it live,” Betzig says. “And it drives me insane because I keep making this point over and over again that we need to go back to a holistic understanding of this complexity that exists in the cell instead of being focused on just little parts.” He’s certainly bold in his assertions. But if he succeeds, we may soon discover that our current understanding of the cell has been rudimentary at best. I guess the science is never settled. There are always new discoveries. You can always go further. To not recognize that is to not understand science itself.

The Cell we See in Textbooks

Betzig and his colleagues used PALM-style visualization technologies to watch transcription factors in live cells. These are the proteins that gather at the start of a gene before the cell reads it to initiate some process. The accepted scientific model at present says that transcription factors form stable complexes that last minutes or hours. But that’s not what Betzig saw when he looked with his new kit. “None of them were binding to the DNA for more than a second or two,” Betzig said. “And so it was like, holy fuck, our whole model of how transcription works is completely wrong.”

That’s another thing about Betzig. He curses. A lot. I like it. We need more Nobel laureates like that. It’s refreshing to hear deeply technical language mixed in with talk you’d hear in a garage. Anyway, the results he observed shaped the rest of his career as he went on to build more microscopes with even greater resolution. All of that experience has led him to a pretty wild claim:

“This is the take-home story of this whole talk. Almost everything you learn in biology textbooks is a hallucination. It is. Or at least cell biology because what they’re doing is they’re taking three reductionist tools — biochemistry, molecular biology, and structural biology — and then hypothesizing how those little pieces, tiny tiny little bits, come together both structurally, stoichiometrically, and dynamically to create the cell. They have no direct knowledge of the stoichiometry or the spatial arrangements or the dynamics. All of that is hallucination. I’m exaggerating, but largely speaking.”

For that example he was pointing to the animations that show up online and in classrooms. “You guys have probably seen on the web those beautiful things of all these molecules coming together. Here’s a cargo on a kinesin walking along a microtubule like this. And it’s all in this vast empty space. I don’t know any cell that’s a bunch of vast empty space. I’m sorry. It’s crowded as fuck in there.”

The fix for this, he says, is to look at cells while they’re alive and moving. “When you start to fucking look at the dynamics, not just the structure, you realize that you had it all wrong. And you realize that so many of the things that they thought they knew, you can’t be sure that they know. We have to reinvestigate all of it. So, that’s why I pivoted to live imaging.”

More Complex than a Neutron Star

Betzig keeps coming back to the scale of what’s inside a single cell. And he gets animated real quick. “There’s 100 trillion water molecules in every cell. There’s 10 billion protein molecules. There’s 10 billion carbohydrates. There’s 10 billion lipids. There’s metabolites. It’s by far the most complex matter in the known universe. We understand the interiors of neutron stars far better than we understand the interior of cells. It’s crazy complex.”

Betzig says that most of the order in a cell comes from what we perceive as random motion. Molecules bounce off one another trillions of times around the cell until they happen to land somewhere they stick. Entire structures build up from there, and he says the same pattern repeats at every level of life. “Living matter involves emergence from many different levels, starting from the stochastic motion of single molecules to macromolecular assemblies to membrane-bound organelles to cells to tissues to populations to the whole fucking biosphere. Everything about life is emergence from single molecules to that level.”

He offered a practical comparison from his experience in the automotive industry to illustrate the point that we just don’t know enough about how the cell really functions. “Imagine you try to reverse engineer an internal combustion engine. If all you have is the reductionist tools that rule biology today and ruled in the 20th century … trying to reverse engineer life from that would be many, many orders of magnitude harder than having a random pile of internal combustion engine parts and trying to figure out how internal combustion engines work. That’s where we are today.” If you’ve ever stripped down and rebuilt an engine, you get a small sense of what he’s talking about.

Why so Many Drugs Fail

Betzig ties this gap in observable knowledge directly to the cost of drug development today. He works with Eikon Therapeutics, which he helped found. Roger Perlmutter, the CEO, is the former head of research at Merck and one of the most prolific and successful drug discoverers in the 20th century. Betzig quoted Perlmutter: “It’s a miracle if we ever find a drug that works because we have no idea what we’re doing. And that is accurate because of this focus on such a crazy level of reductionism instead of understanding the system holistically.”

Betzig concludes: “There’s a reason why only 9% of the drugs that enter phase one come out of phase three. Because we don’t know what the fuck we’re doing. We don’t know the real mechanisms that are going on.”

A drug might bind its target protein perfectly but still fail. The protein may never reach the right place in the cell, or the drug may affect some of the other proteins in the cell. Those problems often don’t show up until clinical trials or after the drug is on the market. He’s especially skeptical of startups that design drugs from protein structures predicted by AlphaFold. “Those proteins and that stuff is such an infinitesimal part of the whole dynamic complex system that creates life. They’re burning money for nothing.”

How we engineer cells today “is to put fluorescent tags on specific proteins so they’ll light up. And that’s one of the most limiting parts of optical microscopy still because in visible wavelengths there’s a limited number of colors that you can do. There’s 20,000 different types of proteins in the cell. It would be great if we could see them all at once, but that technology does not exist. And that is a Nobel waiting to happen if somebody can directly interrogate proteins in some way and deduce them without having to put those tags on.”

His answer is to watch living systems directly at every scale from single molecules to whole organisms. His lab now has microscopes and the technology that can capture that data. The challenge is what to do with the data afterward. That’s the missing opportunity he sees.

Too Much Data and no Way to Read it

The instruments Betzig has cover time scales from milliseconds to days and sizes from nanometers to centimeters. And they work for as far as they go. But the analysis of the data can’t keep up with the collection of the data. “We evolved to see in 2D plus time,” he says. “Life happens in five dimensions. XYZT and molecular species, the 20,000 proteins, all the lipids, the carbohydrates, all the rest. Even with the microscopes that take our petabytes of data at all of those length scales, it’s fucking bits on a drive and it does nothing. It’s so fucking frustrating to have petabytes of data on drives that are completely worthless because there is no scalable way to look and understand that data.”

By five dimensions of life, he means the three dimensions of space, plus time, plus the identity of each molecule, which the microscope records as different fluorescent colors. So, he needs a new kind of model to dive into all of his data and express it as reality.

“What we need is to build a five-dimensional mind that can look and see in five dimensions. A five-dimensional vision transformer. That’s what we need to be able to crack this nut.” He needs AI. But he didn’t come to this conclusion willingly. “I was the last human on earth who ever wanted to have anything to do with AI because I hate doing what everybody else is doing. But I am forced in this direction because I think it is the only possible scalable way to really extract meaning at scale from the data that we can take.”

Segmentation Comes First

The hosts asked what the model would actually be trained to do. Betzig says the first job is basic. The model has to clearly see where things are and where things begin and end. It has to segment. It has to see the edges of everything in the cell.

“The first task that you have to do is robust 4D segmentation. The fundamental unit of life is the cell. You would like to be able to see the cells individually. Beyond that, you would like to see the organelles inside of the cell.”

He points out that even flat images are still a problem. “In this modern age with everything that we have, even 2D segmentation is imperfect.” Tools like Meta’s Segment Anything work on 2D images, and researchers often apply them one slice at a time to 3D data. He thinks that’s a mistake. “That’s not the way to do it because you’re missing a prior. The prior is there’s reasonable continuity between successive planes. And then likewise, there’s reasonable continuity in time. The cell doesn’t jump to this right away. It moves continuously. So you need to build an inherently native 4D model.”

Once an LLM is taught and understands where cells and organelles begin and end, he wants to talk to it directly. “I want to be able to ask through an LLM interface what happens when a T cell is, through immuno-oncology, engaging with the tumor. What particular proteins are expressed at the surface? What is the course of its motility in order to get to the tumor? I have all of that data on fucking drives, but I can’t access it because I can’t find out exactly where it is. I need something to recognize that.”

Training Without an Army of Annotators

One host noted that most AI successes depend on properly labeled data. Tesla, for example, learns from drivers who take over the wheel. Betzig says the fluorescent labels already carry much of that information, and he can teach the LLM all it needs to know about the cell. If a T cell and its target glow in different colors, the model can see when they meet. The rest has to come from the data itself. “This is why you have to do something through a transformer, because it has to be self-supervised with minimal annotation thereafter. That’s the only way it’s going to work because you just can’t use human annotation at scale.”

The hosts raised self-driving cars as a comparison, and he quickly agreed that it’s the best one. “You have hit on the closest analogy to what we need exactly.” He added, though, that Tesla’s cameras still only estimate the third dimension. He tried to interest xAI, along with many other AI companies, in his project but without success yet.

Is this Even Possible?

Betzig isn’t saying the problem is easy. “Even I, and I’m not an AI guy, feel like this is at the very bleeding edge of what’s tractable. It is a big scale. It would make AlphaFold look like a picnic.”

He’d like to test his idea with smaller ablation studies first. But compute resources are scarce, and his funding is limited. “I would kill to have 128 B200s.” His lab recently bought 32. It took nine months to get them, and then they had to “fight like hell” for enough power and cooling to run them.

Hiring the right engineering talent is also hard. He works at Berkeley, a few miles from companies across the bay in Silicon Valley that pay AI engineers far more than a university can. “At fucking Cal, how much do you think I can hire an AI engineer for compared to what he can get five miles away from here? You have to find the crazies. The crazies like me and Harald who don’t give a fuck about the money. They give a fuck about the problem. They’re hard to find.”

Even viewing the data is slow. “A 10-second movie can take days,” he says. His team is looking at Gaussian splatting as a lighter way to represent the data, although he’d prefer that someone else build it. “The last thing I want to do is reinvent the wheel if there’s people better at this type of shit than us neophytes.”

He already uses AI in his own daily work. He used to design optical systems in Zemax. Now he describes the lens he needs to Grok, and the design comes back in about five minutes. He estimates the full project would cost about $50 million. So far no funders are interested, including the Howard Hughes Medical Institute, which still supports his lab. Betzig also admits that he’s no businessman. He’s an engineer. A scientist. He could also use some business types who can think about these problems like he does. They, too, are rare.

How it Could be Used

Besides basic biology, Betzig sees a practical use in drug testing, which brings him back to the failure rate he complained about earlier. His lab uses zebrafish because they’re transparent vertebrates that share much of their genetics with humans. With a trained model and a baseline of normal activity across organs and cell types, a company could screen drugs for off-target effects before human trials. “If they got something they’re about to put in phase one, they bring it to us, we put it into the fish, and we see if bad shit happens.” That could revolutionize the pharmaceutical industry by solving problems it’s had for decades. He didn’t mention this in the interview, but could you imagine the images you’d get in future MRI and CT scans if the resolution could be increased to the point that Betzig imagines?

No Virtual Cell Yet

He has little patience for efforts in the field to simulate a whole cell from partial data. Most of those projects rely on spatial transcriptomics or protein structure. “It’s one part in 10 to the tenth of what’s going on in the cell, and they’re going to recapitulate all of cellular behavior on the basis of that. It’s naive to the point of craziness, in my opinion.”

He has a broader complaint too. “Everybody’s creating huge atlases, and so we’re getting all sorts of data and no understanding. Just collect, collect, collect.” But at least having the data can lead to the perfect opportunity for AI to use the data in a way that’s never been possible before.

When the hosts kept coming back to modeling, Betzig stopped them cold. “I haven’t gotten through to you guys how complex it is. We’re not ready to model. We have to observe.” He compared the field to astronomy before Kepler. “We have to be Tycho Brahe and guys like that. We have to be looking at the thing and seeing these epicycles. We don’t fucking know the inverse square law yet, but we have to see the epicycles first.”

And he’s clear that just storing piles of data on disks doesn’t count as observing it. “Putting it as bits on drive is not observation. Having understanding either through a machine or through humans, or best, humans and machines working together, is the only path forward.”

It’s Mystery Everywhere

Near the end of the interview, Betzig described why he spends his last working years on this work.

“It’s the last frontier. It’s mystery everywhere. We understand so little. We’re in the pre-Kepler era. We’re in the phlogiston era of understanding biology. There are so many things we think we know that are just bullshit. It really is the final frontier because it is the most complex matter in the known universe. And it’s going to take generations, if not eons, if ever, to have enough of a mechanistic understanding to really do a virtual cell like they say they want to do now. People just don’t understand exactly what big of a mystery it is.”

And More

The interview goes on and on to explore many other subjects such as peer review, the culture of science, politics, funding, grant writing, engineering, rockets, space exploration, and many bits I can’t even remember. I only really covered the AI section. Betzig thinks as he talks so he can bounce around a lot. But it’s that free-form thinking and his enthusiasm that’s addictive because he makes such interesting connections. In fact, he ends the interview by saying that the best place for young people to go work right now is SpaceX. He says SpaceX is the new Bell Labs since they are doing the best science out there at the moment. If he were back in his 20s, he says he’d “beg for a job at SpaceX.”

Betzig is quite the character. Personally, he reminds me of Kary Mullis, who won the 1993 Nobel Prize in Chemistry for his work on the polymerase chain reaction (PCR). It’s a shame Mullis died so early. We need guys like Betzig and Mullis in science who speak so freely about what we don’t know now and the possibilities we all have to create new realities in the future.

Here’s a Betzig technical presentation where he explains the details of his story:

Eric Betzig, Is anyone there? Information transfer in biology from proteins to organisms, 4/2026

This session flows really nicely and more precisely than the free form podcast interview above, so it’s convenient to understand Betzig’s timeline in its entirety. He’s remarkably articulate. He also shows videos of Lattice Light Sheet Microscopy in multicellular organisms, so you can see cells performing in living tissues. It’s amazing. We also learn about his new initiative that he calls The Cell Observatory, which includes his lab’s massive effort to solve the problem of viewing live cells from molecules to organisms. What’s most interesting about the session is that it provides all the technical detail that helps explain the frustrations he expressed in the 632nm Podcast interview. He’s really on to something here, so it’s easy to share his passion. The talk also lacks his colorful curse words if you care about that as well.

Think First, AI Second

Here’s a cool new study on learning using AI:

Think First, ChatGPT Later: Guiding Human–AI Collaboration for Learning Gains in Independent Human Creativity (April 2026)

The authors, Sarah Shi Hui Wong and Sophia Xuefei Qiu, ran a simple experiment with 196 university students and published the paper earlier this year in Educational Psychology Review. What struck me about the study was that the results seemed so obvious. I must be missing something. But maybe I came to that same conclusion after using various AI tools and just figured out on my own what they do best and what they don’t do very well at all. For me, I found that if I engaged the AI with something already in mind I was far more productive and the experience was more satisfying, whereas if I just randomly poked around with the AI with an empty mind then we both just ended up running around in circles and I learned nothing.

From the abstract:

“Thinking of one’s own ideas first, then collaborating with ChatGPT to improve them, promotes learning gains in independent human creativity.”

And with that, I knew I was on the right track and read the rest of the paper.

Wong and Qiu are researchers working in cognitive and educational psychology, and their study draws heavily on the memory and learning research that those fields have built up over the past several decades. The tradition also traces back to people like Robert Bjork, whose PhD is in mathematical psychology, and other researchers studying how the human brain encodes and retrieves information. These questions predate generative AI, but now they can be studied in the context of AI to explore new possibilities to use the technology. The findings of the ChatGPT study support my own experience since I’ve been using AI specifically for learning from the beginning. It’s the best teaching tutor I’ve ever had by far. The possibilities for education are endless. I wish I had these tools when I was a kid in school.

Here’s what they did with ChatGPT, but it works with any chat-based AI tool. They split students into three groups. The first group solved a creative task alone. The second group used ChatGPT however they liked. The third group followed what the authors called a “think first, ChatGPT later” method. So that third group of students had to think of their own ideas first and then iterate with the AI to improve those ideas. And then they’d submit a final answer on their own like the other groups.

On the first task, the group that used ChatGPT freely did the best. And if you’ve used AI at all, it should be obvious that the first group came out ahead initially. They seemed more productive. But they outsourced their own thinking entirely. Why does that matter? Well, from a longer term learning perspective, this is where it gets most interesting when you look more carefully.

Next, all three groups did a second and harder creative task with no AI access at all. At that point, the free-use group’s advantage vanished. Their scores dropped back to the level of students who never touched AI. But the guided group using AI was different. They hadn’t looked especially strong on the first task, but on the second task working alone they ended up beating the other groups. The researchers then read the chat transcripts and found a possible reason. Students using ChatGPT freely mostly just asked it for external ideas outright, just like I did when I was aimlessly messing around with AI. The guided group, however, mostly brought their own ideas to the chat conversation and asked ChatGPT to push on them. In other words, they iterated. That difference in human behavior, not the AI itself, predicted who actually learned something that resonated over time.

So why is this important? Use AI to help you learn! Simple. And we now have some data in a controlled study to substantiate that AI is a useful tool for learning beyond just anecdotal experience. Sure, AI can make you more productive really quickly, and I use it for that purpose all the time. But to actually leverage the tool it makes more sense to go in with a plan in mind so you can be both more productive and learn something valuable in the process. Then you can take what you’ve learned to focus your prompts to the AI so that you can move even more deeply into the subject. Again, this process is obvious if you’ve spent any time poking around with AI.

Others have also discovered this phenomenon as well. Tamara Tate published a piece called Think First: Why the First Idea Shouldn’t Come from AI in March 2026 before the ChatGPT study, and she made a similar argument about generative AI and writing.

As I read both pieces I found the concepts familiar because they reminded me of a book on learning I read years ago: Make It Stick: The Science of Successful Learning. Make it Stick is probably the best book on leaning I’ve ever read. Every page is pure gold for people who are motivated to learn on their own. The book is based on the same body of cognitive psychology research as the ChatGPT study, particularly the work of Robert and Elizabeth Bjork on what they called “desirable difficulties.” Wong and Qiu cite that same research to help explain why the AI guided group may have performed weaker initially but excelled later. The key is that the students were doing more work up front. Attempting retrieval on your own (thinking first) before instruction beats instruction alone followed by retrieval (testing) afterwards. That’s not a new idea. It’s just being tested on a chatbot now. We now have a new tool at our fingertips to implement an old principle.

Aside from going into an AI conversation with at least a rough plan up front, it also makes sense to spend some significant time iterating with the AI at length. The ChatGPT study only covered a 12 minute timeframe before answers were submitted. But try going longer. I can talk to the damn thing for hours just going back and forth asking and answering questions, brainstorming ideas, taking quizzes, working thorough tutorials, and anything else I can think of. I tend to have hand-written notes to the side, though, to track things that are important and need more probing. I learned this from an engineering manager I worked for at Sun in Solaris engineering where I observed her managing meetings of engineers. She asked one question after another until they solved the problem. I was surprised how deep she could go to get the engineers to think differently about the issue. She always had another question to ask, which led to additional questions the engineers would ask each each other as well. Those meeting sessions would also often lead to further conversations on the team’s mailing list. It works. And AI has infinite patience for this iteration. I guess Socrates was right.

Here’s the final sentence in the ChatGPT study:

“Generative AI is not to be a substitute for human creativity but, when strategically harnessed, a powerful tool for enhancing it.”

Makes sense to me. Working it every day.

The Value of Contributing: Mentorship

I found this interview on the JetBrains YouTube Channel really interesting: Zig 2026: No-AI Policy, $670K Foundation, Left GitHub & Why Zig Isn’t 1.0 – Andrew Kelley Explains

Some developers really love to code. They share their code and they mentor others to share their code as well. The entire process is a challenge for them. It’s their passion. It’s their craft. And for these guys they aren’t casually dumping their entire development experience for the latest automated AI system to come along.

Here’s Zig Software Foundation lead Andrew Kelley during the interview:

“I love computers. And I love learning about what people are doing with them. And there’s a sense of mystery and magic that you can get from reading someone’s explanation of a project that they did that took them a very long time. They had to learn lessons, and they had to increase their skill as a programmer and as a user of computers in order to accomplish this goal. And when you read a blog post like this, it’s brilliant. It captures the imagination. It makes you think about what you could do yourself. It teaches you something. It connects you to them emotionally.”

And that’s why Andrew bans the use of AI on his Zig projects.

Instead, he’s seeking that close connection with his contributors. He’s seeking excellence and discovery from the many procedures involved in software development. I picked up on this attitude right away because I’ve interviewed many engineers about the value of contributing. Many of them hold remarkably similar views to Andrew, and they talk at length about the importance of contributing to the community.

Here Andrew explains it beautifully:

“The main point of doing code reviews and having contributions, instead of us just doing all work ourselves, is mentorship. The whole point is that a contributor can become a core team member eventually or a more valuable contributor. This will help the project because we’ll have more people who can contribute to Zig skillfully. And it will help their resume because they’ll be a better systems programmer, and they can then take those skills elsewhere.”

But the contributions Andrew used to get from developers using AI were “garbage and had no value whatsoever,” he says. That’s pretty strong language. But Andrew has very high standards and would prefer to engage developers directly on their contributions during the code writing, review, and integration process. That’s where he builds valuable, long term relationships where contributors become more skilled, the core team benefits, and the product maintains the highest quality possible. AI, he says, only gets in the way of that process.

Education is also a big reason for why Andrew holds this view.

“This policy just makes sense because the Zig project is also an education project. That’s part of our mission statement. We’re providing guidance and education to students. So we’re all trying to learn. We’re all trying to get better at programming. And so people who are sending AI pull requests, those people are not helping this goal. In fact, I think that they’re detracting from this goal. So, for our project, I think that the strict no AI policy is an appropriate policy.”

Many times websites bury their organizational mission statements down at the bottom of the page or tucked inside a nav element. But on the Zig Software Foundation site, it’s printed right at the top as the first item.

For developers using AI, Andrew says, it’s simply not worth investing in them. “They aren’t going to join the core team later, not a chance,” Andrew says. It takes too much time away from the reviewers, and the contributors don’t learn anything in the process and are less likely to stick around. Andrew isn’t just saying this, though. His team has tried taking AI contributions but it hasn’t worked out, and they still have hundreds of contributions that still need reviewing. So he’s speaking from his own experience. He’s actually quite thoughtful about the whole thing, so it’s hard to argue that he’s just making a mistake or that he’s blind to the emergence of AI.

So, what’s the lesson? Learn your craft. Learn it well. Go deep into the details so you can demonstrate your expertise so that it’s unmistakable to more advanced developers. And look at the entire process as a community building exercise. In other words, it’s personal. It’s more than just code. Don’t outsource that opportunity to a machine. And then you can build lasting human relationships with the core team via iterations about your code that you submitted yourself. That’s a powerful signal to send into the noisy world of software development. It’s also a simple technique that can be applied to any field you choose. All craftsmen in all trades know this.

Andrew’s position on AI may fly in the face of the current trend of developers embracing AI tools, so we’ll have to see how he and his project evolve as AI grows to pervade more levels of software development. So far, though, Andrew remains emphatic. And judging from many of his comments throughout the interview and how deeply he gets into system development, I bet he sticks to that view. At the end of the conversation, Andrew says he’s probably unemployable at this point and is best suited to be on his own. He’s clearly an entrepreneur at heart, and his standard is “uncompromising perfection.” Can’t argue with that.

And finally, here’s a shout-out to the JetBrains team for talking to Andrew, who said flatly in the interview that he doesn’t use JetBrains products because he only uses FOSS software. Even with advanced IDEs on the market, Andrew’s development set up remains simple and open: “It’s just a terminal and Vim.” Like I said earlier, he loves to code, And he seems as hard core about that as you can get. But the important bit for JetBrains is that they had him on at all. The interview has nearly 900K views in just three weeks, which is vastly more traffic than other recent videos on the channel. Most software vendors would never do this, no matter what traffic a guest could generate. Most vendor channels exist to simply sell products. But JetBrains with this decision demonstrates that they are also concerned with promoting excellence in software development and in building the overall software community. Good for them.

A Double Standard Around AI?

I’ve noticed something interesting around all the conversations on AI these days. The issue isn’t necessarily new, but I guess I now finally have an opinion on it since I live in both worlds.

If you follow technology, it’s obvious now that many software engineers speak openly about handing over their coding tasks to AI agents. And most of them are praised for that move. On the other hand, writers who use similar AI tools to draft their prose usually get a vastly different response. They actually face harsh and sometimes nasty criticism. Why the difference? Both groups use AI technologies, yet the social judgment for each runs in radically different directions. Sure, the two crafts share some things in common, but they certainly aren’t exactly the same and comparing the two is difficult. So, does the disparate reaction really represent a double standard? Let’s see.

How Developers Worked Before AI

Go back. Picture how developers built software before AI. Roughly speaking, they started with a problem, broke it into pieces, and thought through the architecture before typing any code. They chose data structures, wrote functions, planned testing, and reviewed edge cases with team members. When the code failed they read the errors, found the bugs, and fixed them. They reviewed their own work and asked colleagues for reviews as well. Every choice in the finished code passed through the developer, the team, and various systems. Human judgment stayed right at the center. The process was slower by today’s standards, but that slower pace often forced developers to think carefully about what they were building before anything integrated into the final product and shipped to customers.

How AI Changed the Developer’s Process

Compare that older process with how many developers work today with AI. They still start with a problem and break the work into pieces. The difference now, though, is that they write detailed prompts into an AI system instead of writing their own code. A prompt can name the files, state the tech stack, describe the desired behavior, and list any constraints. An AI tool then produces a draft, sometimes across many files that touch many parts of the system. Developers review the output, run the tests, fix errors, update the code, and tighten the result. They iterate with the AI as needed by refining prompts or making direct edits to the code themselves. For routine tasks such as boilerplate, tests, or simple refactors, the AI seems to save time.

Moving up to more complex logic or significant architectural decisions, developers still largely do much of the thinking and remain responsible for the final code. But recently the job has shifted significantly toward directing multiple AI systems up front and verifying outputs at the end. Of course, careful review remains essential for serious work. This is generally described as managing agents, and it’s characterized as a sign of technical sophistication. But a smaller group of developers say they now rarely open the generated files at all because they believe the systems are reliable enough. That more extreme group sometimes draws criticism within engineering teams if developers can’t explain or fully own the result. Nevertheless, this trend toward the edge continues and seems to be increasing. Will it mark the new standard in the years to come? There are good arguments on both sides.

How Writers Used to Work Before AI

Now let’s shift to writers. They went through something similar to engineers but on their own, much longer timeline. For them the old way meant research, outlining, interviewing, drafting, and editing using their own editorial or manual tools. Writers wrestled with ideas for days or even longer before they carefully crafted concepts or scenes onto the page. They revised their text, cut what didn’t work, and sharpened what did. This iterative process could go on for quite some time during the drafting process. Aside from formal copy editing that came later, the voice on the page belonged exclusively to the writer because every word was crafted by hand.

How AI Changed the Writer’s Process

Modern AI systems changed all of that. Writers can now hand an AI their rough notes, ask for some background research, and get a full draft back in mere minutes. That may be an oversimplification, but it’s not that far off. From there they can edit to whatever degree they want to make sure that more of their own voice survives from the original AI-generated text. So, the starting point for writers has moved from a blank page to a rough draft or even sometimes a seriously advanced machine-generated draft extremely quickly. That shift is similar to a developer’s starting point moving from a blank file to a block of generated code that just works as specified.

Reactions to the Two Fields Split Radically

The two examples above aren’t exhaustive or even totally parallel, of course. Developers are much more used to automated systems that aid in building software, whereas writers have not traditionally had tools capable of generating substantial amounts of finished prose immediately. Here’s where the paths split significantly. It’s not really in the work itself but instead in how the world responds to the work and the behavior of the people initiating the work.

When developers say their agents wrote most or all of their code, people in the field and observers online often call it efficient and innovative. Companies hold it up as the future because faster delivery requiring lower headcount translates into competitive advantages in the marketplace. That argument is pervasive right now. And engineers who still code everything by hand often hear that they’re falling behind the times. So the bias toward using AI is the clear standard for software development.

However, if writers say the same thing the reaction totally flips. If writers brag that AI drafted a chapter, even one rewritten by hand many times, readers and critics immediately express outrage. Editors ask questions they’d never ask developers. Some publications ban the practice outright. Many schools do as well. The finished piece gets picked apart for tells, real or imagined, in a way finished code rarely does.

To give just a small example of how fast this issue changed on the tech side and how it still lags on the editorial side, check this out. We were told at my last job just last year that we were forbidden to use AI for anything, and then just a year later, virtually overnight, we were required to use AI for pretty much everything. Rapid changes in strategy like that is normal for tech. The editorial community, however, still grapples with even touching AI for any of their tasks.

Why the Judgment Differs

Part of this disparity comes from how each field values quality work. Computer code has a clearer set of qualifications. It runs or it doesn’t. It passes the tests or it doesn’t. Users either get what they need or they don’t. But prose has no such test. Readers respond to the human voice and the sense that a particular writer shaped those particular words. And that feeling seems pretty sensitive. If readers and editors learn that a machine had a hand in the editorial process then the literary spell breaks, even when the finished text shows no obvious signs of having been produced by a computer.

There’s another piece to this puzzle, though. Code has always been written primarily for a machine to execute. A developer might read it later, but the first audience was the compiler. Success meant the machine did what it was told to do. Prose never worked that way. It was generally written by one person for another person. It only started drifting toward a machine audience recently, first for search engine optimization and now for AI systems that write, summarize, and reuse text.

That difference in history helps explain part of the discomfort readers feel when they consume AI content. Writers may not realize this but engineers using AI agents are actually extending a close relationship with machines that already existed. Developers have always used the latest tools to help in the code writing process. But writers leaning on AI are now altering a world that used to belong almost entirely to readers. And readers can feel that shift even when they can’t always name it clearly.

The tools themselves are also at different stages. AI coding assistants have become reliable enough for real production work on many routine tasks and even some advanced ones. But the ways developers can test and evaluate generated code are often more concrete than the ways writers can evaluate the quality of generated prose. That line for writers seems more distinct and rigid. For now, anyway.

Internal Standards in Software

But even inside the software industry a version of this line already exists. Programmers who let agents generate code without closely reading or understanding it are often described as vibe coders, and plenty of other engineers still push back hard against the practice. For them their standard sounds a lot like the one writers get held to.

If you review what the AI produced, test it, and can explain how it works, that counts as real development no matter how much the AI contributed. If you can’t do any of that, though, it doesn’t count. Writers face nearly the identical test. The difference is where the line sits and how easily it’s moved. Inside engineering, plenty of AI use still clears the bar and earns credit. But in writing, even heavily revising an AI draft, it seems that writers can’t clear the line at all, at least not in the eyes of the people judging the finished product.

That explanation only goes so far, though. It doesn’t fully account for why engineers who fail their own field’s test, the ones who openly say they barely glance at the code, sometimes still receive praise for their candor or for embracing the future. Writers would likely face a much different reaction for the same admission. So although the work has moved in the same direction in both fields, the judgment certainly hasn’t.

Where Pride in Craft Can Move

None of this is new to developers. Long before AI, many engineers took real pride in writing their own code, not just in what it did. They argued over whose solution was cleaner. They refused to publish something messy. An ugly or inefficient implementation was an embarrassment even when it worked. They took pride in authorship with the same instinct that drives writers to protect phrases, sentences, and characters. The difference between the fields was never that developers lacked that instinct. It’s that the instinct had somewhere else to go once AI arrived. And this is a subtle point that the editorial community doesn’t get yet.

For many developers, though, the code was always a means to an end. The real deliverable was a working system, and the code was the best available way to produce that system. Pride in the craft and the outcome pointed at the same target for a long time, so nobody had to choose between them. AI broke that link. Once a working system could exist without hand-written code underneath it, the pride had somewhere to move. It shifted up a level to the architecture, the judgment behind what got built, and the skill in directing agents to write the code. The identity didn’t disappear. It just relocated to a layer AI hasn’t fully reached yet. It remains to be seen how developers will react when AI can cover the entire process, perhaps writing machine code directly, and enabling non-coders can produce the exact same output as the engineers. Some say we’re getting close to that reality right now.

Why Writers Have Less Room to Move

For now, though, writers don’t seem to have an equivalent layer to retreat to, and the reason is structural. Developers hide code behind a working system that the user never sees. That’s by design. Two completely different codebases can produce an identical experience. The code, no matter how it was written, sits quietly behind a layer of insulation.

But prose has no such layer. A reader doesn’t experience some abstracted effect that the sentences produced somewhere upstream. A reader experiences the sentences themselves directly as a new world gets created in their minds as their eyes scan the text. Change the words and you’ve changed the actual thing being read and the world that’s being created. The sentence is often the product in a way code usually isn’t. So there’s no equivalent shift available for most writers. Will that change as readers grow used to AI-generated text? And what happens when AI agents can craft long works directly in a specific writer’s voice without any human or machine detection?

Markets and social pressure reinforce these views for now. Software users and companies primarily care whether the system works, how much it costs, and what value it produces. But many readers of prose still treat the human relationship with an author and the sense of a particular mind behind the words as part of what they are buying. Countless readers return to text over and over again simply to experience the pleasure of how specific passages are crafted. Authenticity is not only a timeless cultural preference here. It’s part of the product itself for a large share of the writing market.

Copyright and the Writer’s Dilemma

There’s another structural difference that writers face and developers largely don’t. That’s copyright. For a writer, publishing text that contains significant AI-generated material creates a genuine legal and commercial question. Under current United States copyright law, purely AI-generated expressions aren’t protected simply because a human prompted the system. Although a larger work can still receive protection when a human author contributes sufficient expressive elements, this line will surely be tested in courts at some point as the things blur. For now, humans at publishing institutions are making those judgments through various editorial policies.

Regardless, this distinction is still uncomfortable for many people. A writer who generates a chapter and publishes it essentially unchanged may have little or no copyright claim to the AI-generated expression itself. But a writer who uses AI as a starting point and then substantially rewrites the material may have a stronger claim to the human-authored portions. The machine can make the work easier to produce while making the boundary around authorship harder to define. This is one reason the reaction to AI-assisted writing may not rest entirely on questions of linguistic purity. There’s a real intellectual property question here.

But developers face copyright and licensing questions as well, although they often differ from the problem writers face. The code engineers ship usually lives inside a broader intellectual property framework involving employers, contracts, proprietary software, or Open Source licenses. A developer may need to worry about whether AI-generated code resembles existing copyrighted code or incorporates material subject to a particular license. But the question of whether the developer personally authored the expressive work is usually less immediate. For writers, though, the words themselves are the product, so the copyright question is personal and potentially significant.

Patents and the Developer’s Emerging Concern

Writers aren’t the only ones with legal exposure, though. AI may complicate the situation for developers too. Developers who work on genuinely novel inventions have another issue to consider. And that’s patents. AI-generated code doesn’t automatically create a patent problem, though. The question isn’t simply who typed the code. Patent law is concerned with the invention and with who actually conceived it. The United States Patent and Trademark Office’s current guidance says the same inventorship standard applies whether or not AI was used. AI can assist a human inventor, but only natural persons can be named as inventors.

That creates an interesting problem. Suppose developers know what they want to build and use AI to implement a technical solution they had already conceived. That’s one thing. But then suppose developers give the AI a problem and the AI comes back with the novel mechanism that actually makes the invention work. They may have directed the system, but direction alone doesn’t necessarily establish inventorship. The important question is whether they made the required contribution to the claimed invention.

An AI-heavy development process could eventually create a strange reversal. Developers may end up with more efficient code and a better product but less certainty about whether their contribution was sufficient to establish inventorship of the underlying technology. Here’s where AI could actually earn its characterization as true artificial intelligence.

There’s another issue too. Developers working on potentially patentable inventions need to be careful about what they put into external AI systems. The problem is broader than trade secret law. Developers can have contractual or employment confidentiality obligations even when the information does not qualify as a trade secret. Patent law also depends heavily on novelty and timing, and public disclosure can lead to consequences that vary by jurisdiction.

The safest approach isn’t just to ask whether the AI wrote the code. It’s to keep track of what the human actually conceived, what the AI contributed, and where the confidential invention details were sent. This can get complicated jet fast as AI becomes pervasive and handles more and more complex computer science tasks.

So, although not exactly the same, patents can be a useful counterweight to the copyright comparison. Developers may not have the same immediate authorship problem as writers have, but they aren’t completely outside the intellectual property problem either. The legal question has simply moved to a different layer.

Institutional Rules Differ Too

Institutional rules and agreements reinforce the disparity between writers and developers as well. Publishing contracts, academic policies, and editorial standards still treat AI contributions as something that must be disclosed or even banned outright. But software employment contracts and Open Source development practices treat AI assistance as ordinary tooling with far fewer formal restrictions.

There’s yet another wrinkle in the rules game. International regulations. The European Union is beginning to treat transparency around AI-generated content as a regulatory issue. The EU AI Act’s transparency obligations under Article 50 began applying on August 2, 2026. Among other things, the rules require certain AI-generated or manipulated text published for the purpose of informing the public about matters of public interest to be identified as artificially generated or altered. There are exceptions, though, including situations involving human review or editorial control. So this is not a blanket requirement that every article or book touched by AI carry a label.

That’s interesting within the context of the difference between writers and developers. A person reading an article may have an interest in knowing whether the words were generated by a machine, particularly when the article is intended to inform the public. A person using a piece of software, however, generally has no comparable expectation that they should be told whether an AI agent wrote some or all of the underlying code. The EU rules could therefore suggest a broader distinction between AI used to build something and AI used to produce the thing that another person consumes directly.

The EU is also putting copyright obligations directly on providers of general-purpose AI models. Those providers must have policies for complying with EU copyright law and must publish a sufficiently detailed summary of the content used to train their models. The relevant obligations for general-purpose AI providers began applying in 2025 with the European Commission asserting enforcement powers over those obligations in August 2026. At present, most major model providers are complying or plan to comply, but what happens if they don’t? xAI, which produces Grok, currently isn’t complying, so the issue may result in fines or court cases. There’s a long history of US-based computer companies paying massive fines to the EU for other regulations that have resulted in disputes in international trade with strong statements of protest coming from the Trump Administration.

Even this issue is still developing. Different countries will continue to approach AI, copyright, patents, and disclosure differently. It’s probably too early to tell whether these regulations will create a permanent distinction between writing prose and writing code, or whether that line will blur as AI becomes a normal part of both professions.

What the Split Inside Software May Suggest

Although software has generally embraced AI, not every developer makes the shift easily. The ones who built their identity on demonstrating individual skill, the tightest solution, or the best line of code still feel a loss that looks somewhat like what writers feel. For them the code was never just a means to an end. It was closer to the point itself.

Both fields also worry about the next generation. In software the fear is that junior developers will never build the necessary skills to understand, debug, or design complex systems. In writing the fear is the loss of voice, critical judgment, research skills, and the long apprenticeship of craft. Watching some developers relocate their pride with ease while others struggle to let go of authorship suggests the split may run through personal temperament more than through the profession itself. Writers may be looking at a preview of their own future in that split with some finding a new place to put their pride and others not finding one at all.

Maybe there’s another difference that will become more important over time. Software has always been layered. A developer can move from writing machine instructions to designing systems, then to architecture and product decisions, and eventually to deciding what problems are worth solving in the first place. Writing has layers too, but the final expression remains exposed text. An author can move toward research, structure, editing, or directing AI, but eventually the reader still encounters the words. If those words can be produced almost entirely by a machine, the question of what exactly the writer contributes becomes harder to quantify.

And that may be where the real long-term divide appears. The question may not ultimately be whether AI is allowed in either profession. It may be whether each profession can establish a new definition of meaningful human contribution and quality of work.

There’s no obvious answer yet. Things will remain messy for a while. The technology is moving too quickly and the law is now moving along with it. The European rules are only beginning to take effect, the copyright questions are still being worked out, and patent authorities will have to apply old concepts such as inventorship to a technology that didn’t exist when those concepts were created. It may be that the differences between writers and developers become smaller over time. Or the legal and cultural systems may end up reinforcing the differences.

For now these are just some questions worth watching. Maybe the honest answer to the original question about the double standard is that there really isn’t one. Or maybe it’s just not a clear enough distinction and over time the external reaction to both fields will simply diminish. Code was mostly judged by outsiders on results, while prose was generally judged on the process and on the sense of a specific person laboring over every single word. AI didn’t create that asymmetry. It just made it impossible to ignore. Who knows how this will turn out.

Henri Tremblay at JavaOne 2026

Henri Tremblay at JavaOne 2026 | Duke’s Corner Java Podcast | May 18, 2026

Here’s the second interview I did at JavaOne 2026 in March. Henri Tremblay is a Java Champion, Montreal JUG leader, and EasyMock lead developer from Canada.

Henri’s session at JavaOne covered the Java Memory Model, which is a topic he believes every Java developer should understand well. He’s been to six JavaOne’s and had warm words for the conference, which represents a rare opportunity to meet the people whose code runs on systems and devices all over the world.

He has clear advice for developers: read books, understand how and why your code works, and get out there and join the community.

We also talked about why Java still powers so much of the world’s critical infrastructure, from banks to the Mars rover. Henri pointed out that companies often start in C++ and then move to Java because Java runs nearly as fast once it’s going and is far easier to change later.

On AI, Henri had a balanced view. He uses it for tedious work, like sifting through a gigabyte of logs to find a single error. But he was also clear about the risks. “We should not get lazy at reviewing code because AI will generate tons and tons of code. It’s not bad at reviewing it, but still it makes mistakes.” He warned that AI reflects the average of what’s on GitHub, and most code on GitHub isn’t great. Your role, he said, is to find a better answer.

For students and junior developers, he says they should also leverage AI for learning, but he advises that they internalize the fundamentals of software engineering deeply. “Read books, please, please!” He pointed to Core Java, the book he originally learned from and is now helping revise. Blogs and YouTube videos only tough on surface level issues. Books take you deep and that’s the knowledge you need to grow your career.

Henri Tremblay on LinkedIn: https://www.linkedin.com/in/henritremblay/
Jim Grisanzio on LinkedIn: https://www.linkedin.com/in/jimgris/

Bill Joy’s Future

Was Bill Joy Correct Back in 2000? What His Warning Means for Science, Technology, and AI Today.

Since there’s a lot of distracting and unfortunate AI doom in the media these days, I figured I’d go back and revisit Bill Joy in April 2000 for some history on the pending catastrophes we keep hearing about today. I remember that Joy’s massive article “Why the Future Doesn’t Need Us” hit really hard twenty-six years ago. Joy was a cofounder of the iconic Sun Microsystems, after all. He helped build the Internet. And here he was now warning that three technologies might eventually lead to human extinction. Unlike today, though, that negative perspective was a pretty novel idea back then. The technologies he cited that could end us all included robotics, genetic engineering, and nanotechnology. His core argument was actually pretty simple. These were not like past technologies. These new things could potentially replicate themselves, and that’s the bit that would change everything if they were used as weapons or just got loose by accident.

I first read Joy’s article at a coffee shop in Cupertino, California right across the street from Sun where I worked in software systems marketing. The article was widely read at Sun and also across Silicon Valley and for months also generated wild discussions about Joy and his analysis. Many people just called him crazy. But that only demonstrates to me that those who made such flippant statements never read his article or thought deeply about his arguments. Others knew better, though. They knew full well that Joy was documenting in detail the very real risks of rapidly developing technology and the consequences of ignoring them.

That is, of course, the standard and pervasive culture of Silicon Valley. The valley may be big, but it’s remarkably insular as well. I didn’t know Joy at the time I read his article, but I went on to meet him several times at Sun and worked closely with his teams promoting projects like SPARC, Solaris, Java, Jini, and later on JXTA. I didn’t really know him well, of course, but he was always friendly and professional to me. He was quiet, too, and I always found him a serious thinker who obviously knew far more than he ever expressed. Sun was filled with such characters. They all fascinated me to no end.

The timing of the Joy article is also interesting. He said he had been working on the essay since his 1998 conversation with Ray Kurzweil and continued revising drafts through 1999. When Wired published the piece in April 2000, the tech world was at its peak. The NASDAQ hit a high of 5,048 in March of 2000 just a few weeks before the article dropped. And at that time Sun’s stock reached $250 a share, which gave the company a market cap around $200 billion. That was a significant achievement for 2000. Sun was one of the hottest companies in the valley back then, and it was quite an experience working there. The place was buzzing with activity. I loved it. Many of us did. So, it was into that environment of overt tech optimism running at manic levels that Joy published his thoughts about our potentially perilous future.

Andy Bechtolsheim, Vinod Khosla, Scott McNealy, and Bill Joy at the Sun Reunion in Silicon Valley in October 2019. Photo by Jim Grisanzio.

Boom!

Then everything blew up. The bubble burst. By mid-April 2000 the NASDAQ suffered its worst week in history and dropped more than 25 percent. Companies started dumping workers like I’ve never seen before. Joy’s dark warnings about unchecked technology landed precisely as that optimism crashed. He obviously couldn’t have seen the future, but his timing was remarkable.

Joy’s warnings took on an even darker tone the following year. After publishing the Wired article, he signed a book contract to expand on the concepts. He moved into a hotel room in New York City and surrounded himself with gloomy books covering plagues and nuclear bombs and other such material he was studying about the future risks of technology. Then on September 11, 2001 came the terrorist attacks we’ve come to know as 9/11. I knew a few people who worked at the Sun building in New York City, but I didn’t know Joy was also in the city at the time. He said he stood in the streets with everyone else and watched the impossible happen in real life. The next morning he went back outside and observed a long line of sanitation trucks parked on Houston Street ready to haul away the rubble. Everything below 14th Street was closed, he said. “It was quite a compelling experience, but not really, I suppose, a surprise to someone who had his room full of the books I was reading,” he said in a TED Talk. “I was not surprised that it happened at all.”

Joy eventually abandoned the book project. I point this out just as an aside since the event occurred shortly after he published the article I’m writing about here in this post. Still, it does reflect the feeling of the times. How much had changed in Silicon Valley and the United States in just one year.

Who Is Bill Joy?

Joy was born in 1954 in Michigan, and he was considered a child prodigy. He started school early, was reading by age three, and later excelled in math and science. He even graduated high school at 16. He loved books and thinking, and that became his escape from the world. He also loved science fiction and devoured Heinlein’s “Have Spacesuit, Will Travel” and Asimov’s “I, Robot” with its Three Laws of Robotics. He wanted to be a ham radio operator, which were the Internet hackers of their day, but he couldn’t afford the equipment. On TV, Star Trek inspired his imagination, and Gene Roddenberry’s “The Prime Directive” clearly resonated with him. You can actually see that ethic woven into his writing thereafter.

At Berkeley in the 1970s, he created the vi text editor, which to his surprise, was still widely used more than twenty years later Some hard core developers still use it even now. He also developed the Berkeley version of the Unix operating system and was a key contributor to the TCP/IP network stack. When the other founders of Sun Microsystems (Andy Bechtolsheim, Vinod Khosla, and Scott McNealy) invited him to join them, he participated in the creation of advanced microprocessor technologies and software technologies such as Java and Jini. As co-designer of three microprocessor architectures — SPARC, picoJava, and MAJC — he helped drive innovations that shaped modern computing.

By the time he wrote his famous Wired essay, Joy was only 45 years old and at the peak of his influence among developers in Silicon Valley. But Joy was far more than just a coder. He was well connected to the broader scientific community as well. That’s what made his article so jarring to so many people. He was not an uninformed critic chiming in from outside with yet another opinion in the media, which we’re all familiar with today as we read the news. He was a core architect of the digital age who was expressing deep doubts about where his own work was leading. His self-reflection was pervasive during this time in his writings and during his conference presentations.

The Kurzweil Meeting

Joy’s concern seemed to begin at George Gilder’s Telecosm conference in 1998 when he met Ray Kurzweil, who was an inventor and futurist. Kurzweil talked about how the rate of technological improvement was accelerating and also how humans may merge with robots or download their consciousnesses to achieve near immortality. I remember attending several talks on this topic of immortality when I moved to Silicon Valley in early 2000. It always sounded so silly to me. I wondered how such smart people could take that stuff so seriously. But even now some people in these circles talk about downloading themselves. It still sounds silly. The other bits, though, about intelligent robots and genetic engineering were much more reasonable given my own experience working in the biotech industry before I joined Sun. Joy had heard such talk before and always felt sentient robots were science fiction. But hearing it from someone he respected changed things. Kurzweil gave him a preprint of “The Age of Spiritual Machines,” which outlined a utopian future where humans gained near immortality by becoming one with robotic technology. I don’t know how far Joy goes with respect to robotic sentience, but it’s clearly more than I’m willing to accept.

Nevertheless, Joy’s unease intensified after reading the book. He felt sure Kurzweil was understating the dangers. Then he found a passage in the book describing a dystopian future where machines become so capable that humans depend on them completely. The passage argued that we wouldn’t consciously hand over control to the bots. Instead, “the human race might easily permit itself to drift into a position of such dependence on the machines that it would have no practical choice but to accept all of the machines’ decisions.” In other words, we would gradually grow dependent. I can surely see that as a potential reality, no question about it.

But that passage came from Ted Kaczynski, the Unabomber! Joy admits this realization was uncomfortable to say the very least since he was taking a point from a terrorist seriously. Many people have said the same thing after reading Kaczynski’s words, and he’s actually still cited even today. Kaczynski’s bombs had killed three people and wounded many others. One bomb gravely injured David Gelernter, one of Joy’s colleagues and friends. But Joy felt compelled to confront the argument because, however uncomfortable, he saw merit in that single passage about the unintended consequences of technology.

The Self-Replication Problem

Joy’s main concern centers around one key difference between powerful 21st-century technologies and those of the 20th-century. Nuclear weapons required huge facilities and rare materials. But genetic engineering, nanotechnology, and robotics, what Joy called GNR technologies, require less infrastructure and can potentially make copies of themselves. A bomb explodes once, but a self-replicating machine doesn’t stop. And that could be a serious problem if something goes wrong.

This matters because knowledge spreads freely. You can’t control ideas at all like you may be able to control uranium. Once people know how to genetically engineer bacteria or design tiny self-replicating machines, that knowledge exists in the world and will move rapidly. A small group, even one person, could potentially cause massive harm. Joy calls this “knowledge-enabled mass destruction.” As he wrote, “I think it is no exaggeration to say we are on the cusp of the further perfection of extreme evil, an evil whose possibility spreads well beyond that which weapons of mass destruction bequeathed to the nation-states, on to a surprising and terrible empowerment of extreme individuals.”

He made that point even sharper later. The real danger, he said, is no longer nation-states but individuals or small groups now empowered with “pandemic power.” These new digital, self-replicating technologies give extreme individuals the kind of destructive capability once reserved only for governments. That shift changes everything because the ramifications of mistakes or ill intent can’t be calculated or controlled.

The advancement of technology was clearly moving faster and Joy knew it. He learned about complex systems and non-linear systems from physicists Stephen Wolfram and Brosl Hasslacher in the early 1980s. These are systems where small changes move in unpredictable ways and where feedback loops create unexpected outcomes. Thus, they are extremely difficult to predict. Later, Joy deepened his understanding of these issues after conversations with Danny Hillis, a pioneer of parallel supercomputers and co-founder of the Long Now Foundation, biologist Stuart Kauffman, and Nobel laureate Murray Gell-Mann. Hasslacher and Mark Reed, a leading researcher in molecular electronics at Yale, also gave him insight into molecular electronics, which is the manipulation of matter at the atomic and molecular level where individual atoms replace transistors. When you get to this point in Joy’s article you can’t help but realize that he’s going well beyond just being a smart software developer who happened to strike it rich by helping found a successful tech company in the valley.

Joy knew that the merger of computers and physical sciences was creating enormous power, which up to that point others hadn’t expressed clearly in the popular tech media. By 2030, Joy calculated, we would likely build machines a million times as powerful as personal computers of 2000. That’s enough computing power to make the scenarios that worried him technically possible. As he wrote, “But now, with the prospect of human-level computing power in about 30 years, a new idea suggests itself: that I may be working to create tools which will enable the construction of the technology that may replace our species. How do I feel about this? Very uncomfortable.” He admitted he had once been too optimistic about nanotechnology. Having struggled his entire career to build reliable software systems, it seemed to him more than likely that this future would not work out as well as some people had imagined.

Three Scenarios

Joy walks through what could actually happen with these technologies. What’s interesting is that evolution itself would be one of the driving forces moving these technologies from a positive outcome to something very negative in the wrong hands.

One — Robots might simply out-compete humans for resources the way better-adapted species have always displaced others. We wouldn’t need a robot uprising. Hans Moravec, a roboticist and futurist at Carnegie Mellon, argued that “in a completely free marketplace, superior robots would surely affect humans as North American placentals affected South American marsupials.” So, economic forces alone could push us aside. Also, the dream of robotics includes downloading our consciousnesses into machines. But Joy questions whether a downloaded consciousness would be human in any meaningful sense. The robots would not be our children, and on that path our humanity might be lost entirely. This is one of the reasons why I take Joy seriously. He sees the obvious problem with the downloading issue in a way that others in this field simply do not.

Two — Genetic engineering gives us power to create devastating plagues, either by accident or intention. Joy calls this the “White Plague” scenario, which references a Frank Herbert novel where a molecular biologist weaponizes his knowledge. We now know these profound changes in biological sciences could be imminent and could challenge all of our notions of what life is, and Joy points out that the public remains skeptical even by the standards of 2000.

Three — Nanotechnology could produce “gray goo,” which are self-replicating nanobots that consume the biosphere. Eric Drexler, an engineer and pioneer of molecular nanotechnology, warned in 1986 that “tough omnivorous ‘bacteria’ could out-compete real bacteria. They could spread like blowing pollen, replicate swiftly, and reduce the biosphere to dust in a matter of days.” Joy notes grimly, “Gray goo would surely be a depressing ending to our human adventure on Earth, far worse than mere fire or ice, and one that could stem from a simple laboratory accident. Oops.”

The Manhattan Project Parallel

Joy uses the atomic bomb as his template. After Hiroshima and Nagasaki, physicists were shocked by what they created. Oppenheimer later said the physicists “have known sin.” But there was a real opportunity to prevent a nuclear arms race by internationalizing nuclear power as documented in 1946 via the Acheson-Lilienthal report and the Baruch Plan. It failed, though, because political distrust and competitive pressure got in the way. Within years, the Soviets had the bomb, and the arms race was on. This is a pattern that will repeat in the future.

Freeman Dyson, a theoretical physicist who worked on nuclear weapons and later advocated arms control, captured the moment: “The glitter of nuclear weapons. It is irresistible if you come to them as a scientist. To feel it is there in your hands, to release this energy that fuels the stars, to let it do your bidding. To perform these miracles, to lift a million tons of rock into the sky. It is something that gives people an illusion of illimitable power, and it is, in some ways, responsible for all our troubles, this, what you might call technical arrogance, that overcomes people when they see what they can do with their minds.”

Up to that point, everyone feared nuclear bombs as the ultimate expression of madness. But Joy feared that we were repeating the pattern with even more dangerous technologies, and the commercial incentives for their production would be enormous. Nations compete. Corporations compete. Individuals compete. Researchers want and need breakthrough innovations. The momentum builds, and pretty soon it’s almost impossible to stop. We are being propelled into the new century with no plan, no control, no brakes. However, the driver is not military necessity this time but instead private sector economic gain and competitive pressure. This is human nature, though. Our history is literally filled with these processes. It’s who we are. Joy was simply stating the obvious.

The Relinquishment Argument

This is the part that bothered many people the most. Joy argues for “relinquishment,” or the voluntary decision not to pursue certain lines of knowledge or technology because they are too dangerous. This goes against everything we believe about the value of knowledge and open inquiry, especially in Silicon Valley and across the scientific community generally.

That may be why Joy’s article struck such a nerve. The general population is used to various power centers attempting to curtail their freedoms for whatever reason of the day. But regular people are largely powerless to do much about it beyond voting, which is time consuming, or protesting, which brings its own personal risks. The scientific and technological elite, however, represents something different. They are one of the power centers themselves, and here was one of their own, a high profile one at that, advocating for relinquishment. You have to give it to Joy. He was brave to express such thoughts directly in the face of powerful people.

So Joy asks, what if the unlimited pursuit of knowledge puts us all in mortal danger? He points out that we have done this before. At a 1989 nanotechnology conference, Joy said, “We cannot simply do our science and not worry about these ethical issues.” The United States unilaterally abandoned biological weapons development because the logic was clear. These weapons were easy to replicate and could easily end up in malicious hands. We would be more secure if nobody developed them, and we embodied this concept in the 1972 Biological Weapons Convention and 1993 Chemical Weapons Convention.

Joy further quotes Thoreau: “We do not ride on the railroad. It rides upon us.” Then he asks directly, “The question is, indeed, Which is to be master? Will we survive our technologies?”

Joy isn’t advocating the stopping of all research. He just prefers being thoughtful and strategic about which lines of inquiry we pursue and which ones we intentionally avoid. It means international agreements, verification systems, and scientists adopting ethical codes like the Hippocratic oath. It requires transparency and cooperation. Within that context, he sounds reasonable, right?

Personal Responsibility

Joy writes honestly about his own sense of responsibility: “I feel, too, a deepened sense of personal responsibility, not for the work I have already done, but for the work that I might yet do, at the confluence of the sciences.” That statement carries some weight because Joy was someone who helped create the technologies that might enable the dangers he himself feared.

But he finds hope in the Dalai Lama’s “Ethics for the New Millennium,” which argues that the most important thing is to conduct our lives with love and compassion for others, and that our societies need a stronger notion of universal responsibility. This awareness of greater spiritual principles at play in a world of technology is rare among the valley elite. Joy was saying that neither material progress nor the pursuit of knowledge is the key to happiness. We need to find alternative outlets for our creative forces beyond the culture of perpetual technological and economic growth.

In the TED Talk he gave six years after publishing his article, Joy made the point even clearer. The solution, he said, can’t be technology alone. We need both better public policy and deeper moral progress. He spoke of the need for “the head and the heart,” which echoes Russell and Einstein. He also argued that scientists, technologists, and businessmen must be held personally accountable under the law for the consequences of their inventions. That must have been shocking to Joy’s peers! Yet, today they face no such responsibility. That, he believed, had to change. What’s striking about hearing him articulate these perfectly reasonable points is that they aren’t taken that seriously during the time he was writing or even now. Many thoughtful people say similar things at technical conferences or in political speeches. We all clap and support the concepts. Yet, very little actually changes. Or at least the changes take so long and occur in such small steps that we’re left unsatisfied in the present moment.

Was Joy Right?

Joy’s article was unique for its time in the popular press. When it appeared on Wired’s cover in April 2000, it created quite the rumble in tech circles. Joy even said it generated thousands of letters to the editor, many of them overtly hostile. Wired had been a cheerleader for the digital age for nearly a decade. Its shift from cheering to warning marked an important and surprising moment in the digital community. Also, Bill Joy wasn’t some alarmist outsider like many others. He was one of the core architects. His warning came from inside the cathedral and it certainly resonated. At the time, I was working in marketing at Sun, and one of my jobs was to promote Sun’s technologies to the media. In press interviews during this period reporters would always bring up Joy’s article even if the meeting was booked to cover another issue. We had to draft briefing and messaging documents to prepare executives, managers, and engineers that we brought into all interviews because we knew Joy’s article would always come up in the discussion. I remember many stressful meetings and awkward interviews from this period.

The article also presaged much of what we are experiencing now but not necessarily always in the ways Joy anticipated. His specific predictions about nanotechnology have not materialized yet. And the gray goo scenario is now considered flawed and implausible. Most scientists believe built-in limitations make runaway nanotechnology improbable. But the science is never settled, so we’ll just have to see how things evolve in the future.

But Joy’s underlying concerns were clearly prescient. And we have to deal with that. He worried about knowledge-enabled destruction, powerful technologies becoming widely available, and complex systems we don’t fully understand. All of these have become more relevant as artificial intelligence has advanced today far faster than most people expected in 2000. Interestingly, Joy only explicitly mentioned artificial intelligence once in his article, possibly because he was writing at the tail end of the second “AI winter.” Yet his concerns about self-replicating technologies and systems moving beyond human control resonate well with some of the modern AI risk debates these days.

In 2025, Joy himself reviewed his Wired article in a comprehensive session at Berkeley that covered scientific advancements over the last 50 years. He said he was called a doomer for his article in 2000, but to him it was just a matter of pragmatic realism. “So if we look back 25 years, the risk was real. It’s happened roughly like I said, and we haven’t done anything about it, which is really kind of frustrating and disturbing.” He also connected his original argument from 2000 to the technologies of today. Whereas 20th century technologies require rare raw materials, the thing about “21st century technologies is that they only require information. So, they’re fundamentally different and more dangerous and difficult-to-impossible to control.” One example he cites is that AI today is “verging on being able to do recursive self-improvement,” which is the very process he warned about decades ago.

He also warned about mirror life, which he described in the Berkeley session as a possible threat to the biosphere and cited recent reports from Stanford and Nature. He framed it as just one more example of powerful technologies that could cause catastrophic harm if misused. His concern remains fixed on crazy people using advanced technologies “to do bad things as their first act, and we wouldn’t have the chance to stop them.” His conclusions weren’t encouraging, which certainly fits the tone of his original piece in Wired back in 2000. Technology today has become more powerful, the financial incentives have become more intense, and society still hasn’t done enough to address the dangers he identified years ago. It seems like his predictions were pretty close to an observation of reality today.

In the same 2025 session, Joy said that the early calls for an AI pause and stronger rules for deployment were justified. “For the AI risk I think the people weren’t wrong that we need a pause, and we need rules, and we need to be aware that we’re giving very powerful tools to everybody. I mean, we have the bill of rights and freedom of speech, and we say it’s just information and everybody should have access to everything. That means if we spend a trillion dollars collectively in our society making something which has some good uses and also can be dangerous I have to give it to everybody independent of the fact that I’m then thereby giving it to the people for whom it can be dangerous. Or should I say, well, that’s some specialized knowledge and maybe it should be more closely held.” Fair point. But where to draw the line is the issue and that goes well beyond just talking about technology.

So, Joy recognizes the benefits of AI here, just as he recognizes the value of other technologies. But he also reserves much more of his focus to describe in detail his concerns about the risks. That fits well with the views he expressed in 2000.

He said that the field of advanced technology has quickly become another arms race driven by enormous financial incentives, which repeats the pattern of technical arrogance that he warned about in relation to nuclear weapons in his original Wired article. “People are so excited about what they can do with their minds, they just can’t help themselves. And there’s such a pot of money out there that they just can’t stop.” He said society was now “skating right past” the period when the technology is still flexible enough to control. Once everyone has money and status invested in the race, he warned, changing direction would become much harder. There’s a “social physics” he describes that makes things seriously complicated and extremely difficult for people to consider the consequences. The danger, though, isn’t necessarily the technology itself, which also has many benefits. Instead, the issue is that the technology in the wrong hands can cause problems that can’t be prevented or fixed.

That 2025 session at Berkeley is the first time I’ve heard Joy address his Wired article in detail in the context of today’s debates over technological risk. Hopefully, he’ll keep contributing to the debate because unlike many others in the media at present, Joy can speak articulately about the benefits of the technology without also losing sight of the risks. He has that credibility because he’s lived the history directly. Perhaps he will because now he’s working on an AI startup with his son, daughter-in-law, Claude, ChatGPT, and Gemini as the only employees. So, Bill Joy is very much in the game. It will be fascinating to see what he builds and how that differs from the current implementations of AI.

The Sun Sets

After leaving Sun in 2003, Joy moved into venture capital and focused on green energy and climate change investments at Kleiner Perkins. He also worked as a principal investigator and chief scientist at Water Street Capital. Now at 71, he’s focusing on his AI startup with his kids. And aside from his 2025 session at Berkeley, he’s mostly remained silent in public regarding the debate he actually started 26 years ago. It’s almost as if he said what he wanted to say in 2000, briefly repeated many of the same concepts in 2025, and now we’re left to digest his thoughts and hopefully implement his lessons all these years later. I doubt we will. I saw very little press coverage or social commentary of his 2025 talk at Berkeley.

Over the years, Joy didn’t stay stuck in doom or alarm. After the Wired article he actively tried to move toward better outcomes. He invested in solutions and backed innovations in education, new options to preserve the environment, and a major $200 million biodefense fund aimed at closing the gaps that could lead to a pandemic. He came to believe that we can’t solve the management of dangerous technology with just more technology alone. Instead, we need better policies, markets that price in the true cost of potential catastrophes, and a much deeper moral awareness. That combination of thoughts remains relatively rare today. Perhaps that explains the silence in the market that continues to just focus on doom rather than solutions.

Joy put it simply in that earlier talk at TED. “We can’t pick the future, but we can steer the future.” Over the years technologies have changed, but the fundamental challenge he identified in 2000 remains relevant now. Figuring out how to pursue knowledge and innovation while also maintaining enough wisdom and caution to survive the unintended consequences seems to be a question others should carefully consider today. Yet, few are doing so. In fact, the mania is running faster than ever and being driven by many people who have never heard of Bill Joy.

In his article, Joy compared coding to Michelangelo releasing statues from marble. He described his software engineering in a similar way to those ecstatic moments when the code emerged from his imagination as if it were already waiting in the machine to be freed. He ended his essay with that same image. After eighteen pages of text exploring multiple scientific disciplines and warning about existential dangers from the exploitation of technology, he wrote, “I am up late again, it is almost 6 am. I am trying to imagine some better answers, to break the spell and free them from the stone.”

Well, twenty-six years later, we’re all still up late as well. We’re still searching for those answers. But now we go forward in the wild world of AI, which probably represents the biggest paradigm shift enabling new opportunities in tech since the Internet itself. We should embrace those opportunities, but will we also consider and mitigate the risks?

The AI Doom Vibe Change

For a couple of years now, the single story pounding our heads about AI all day every day has been exclusively about looming disaster. AI takes the jobs, then it takes everything else, then few people get rich. A love story. It was an intentional positioning of the technology, obviously, but the question remains why. Well, Cal Newport has a good hypothesis that I wrote about the other day. But others are noticing as well. So, the phenomenon of the old pitch evolving into something new is probably real.

In a recent episode of The AI Daily Brief, host Nathaniel Whittemore says the previous extremist narrative may finally be cracking, and he cites a fair number sources to substantiate his claim. It’s a different take from Cal Newport’s but there is some overlap. The signals are faint, he says, but they’re showing up in two key places at once so that may mean that the shift will likely have some legs. I think Whittemore may be pulling his punches a bit by saying “the signals are faint” just to cover himself since this shift has been so recent, such as really only the last few weeks. I think the shift is clearly underway. Remember, the IPOs are coming soon, baby! The companies and their simps can’t continue with the doom rhetoric. The American public has rejected that strategy. And it’s interesting that some public opinion polls in China lack this doom positioning. Anyway, back to Whittemore’s daily brief.

The first place Whittemore notices the vibe shift is within the never-ending chattering class in the media. He points to Ezra Klein’s recent New York Times article, “Why the AI Job Apocalypse (Probably) Won’t Happen.” I like the “probably” bit. But coming from a big voice on the political left and also one that’s outside the AI bubble, Klein may carry some weight in Whittemore’s eyes that a similar post from others, say, Marc Andreessen, simply wouldn’t. Klein cites economist Alex Imas from the University of Chicago and also a wider body of economic research to make a case rooted in Jevons Paradox. When something gets cheaper, we tend to use more of it, not less. So although computers may have changed or even eliminated specific tasks, the cost savings created enough new demand that the occupations expanded overall. As Klein puts it, “Every enthusiastic AI adopter I know is working harder than ever because there is more they can do.”

Whittemore points to more data that’s emerging. Software engineering, which is the job category most exposed to AI, is the one where postings have actually increased recently. Citadel Securities cites the increase at 18 percent since May of last year. Federal Reserve numbers also show software engineering jobs at their highest level since November 2023, although the current number is still well under the previous mark three years ago. Also, Stripe Atlas just hit 100,000 incorporations, with Q1 up 130 percent year over year. As Derek Thompson says, “AI agents are better at creating firms than destroying jobs.” A new trend?

The second place the shift is showing up for Whittemore is in markets themselves. Anthropic’s revenue, according to SemiAnalysis, has gone from 9 billion to more than 44 billion this year, which is roughly doubling every six weeks. Atlassian’s stock jumped about 30 percent recently after strong earnings with customers using its new Rovo AI tool growing their own ARR at twice the rate of those who weren’t. The skeptics have been questioning how you justify trillions in infrastructure when seats only sell for 20 dollars a month. Well, that’s being answered by the move from seats to tokens taking place recently in the intelligent agent era. A single engineer with Claude Code might burn through hundreds or thousands of dollars in tokens each month, and the companies selling those tokens cannot keep up with demand.

There’s another piece of the vibe shift worth noting, one which I found most interesting since I’ve worked in both industries. The Associated Press recently reported on construction companies teaming up with big tech to push back on community opposition to data centers. Rob Bear of the Pennsylvania Building and Construction Trades Council told the AP that communities should figure out what they actually want from these projects rather than just saying no. “If you don’t ask, you’re never going to get,” he said, pointing to things like better project plans or money for local schools and infrastructure. Whittemore’s take is sharper. He calls it “an insane indictment of how poorly tech companies have run these projects that the issue has gotten this bad” given how many ways there are to make data centers genuinely valuable to nearby communities at a fraction of the total cost. He’s spot on. The AI companies deserve the public backlash. We’ll see how they adapt to the very real world they are now entering.

Even the AI labs are softening their messages. Sam Altman recently wrote that “jobs doomerism is likely long-term wrong” and that OpenAI wants “to build tools to augment and elevate people, not entities to replace them.” Whittemore says this is a meaningful pivot from a company whose stated goal used to look a lot more like replacement.

But Whittemore is careful not to declare victory too fast. The AI transition will still be painful for specific workers and communities, and history shows we generally don’t help them much at all when economies move through technological advancements. But he ends on a hopeful note.

“I find it extremely encouraging to feel the collective foot being taken off the gas of the AI doomerism for just a moment. If nothing else, it creates an opportunity to have a different type of conversation. One that’s neither doom nor utopia, but about how to adapt to and maximize the opportunity of the change that’s here and coming. I think the more time we spend on that conversation rather than in the extremes, the better off we’ll be.”

The AI Doom Fever Finally Fades

Is the AI Doom Fever Breaking? (It’s About Time!) — Cal Newport, AI Reality Check, Deep Questions Podcast

For many years now executives leading the big AI companies have been telling the public that their own products will destroy the economy and gut the white-collar workforce. As Cal Newport observes this is the rough equivalent of a Pfizer executive announcing a new pill that cures psoriasis but also turns half the population into zombies. But lately the extremist rhetoric on AI has started to soften significantly. Instead of totally replacing entire segments of the workforce, these new AI systems will now simply augment existing workers and also lead to massive new opportunities for employment. That’s quite a radical shift in attitude, especially coming from people whose breathless messaging has been so bold. 

Nevertheless, tech companies are still laying off tens of thousands of employees and citing AI as the reason. So, we’ll see. Newport thinks the previous over-the-top positioning on AI resulted more from culture, whereas the recent shift in tone is likely more tactical. His analysis is comprehensive and seems pretty accurate given that he’s been pushing back on this rhetorical issue for years now. 

A Strange Sales Pitch

In his podcast, Newport runs through many of the recent doom statements. Mustafa Suleyman of Microsoft AI has suggested that AI will be capable of automating most knowledge work within roughly a year. Dario Amodei of Anthropic has warned that the technology will soon replace up to half of all entry-level white-collar jobs in finance, consulting, and tech. Sam Altman of OpenAI speculated last summer about a future in which AI would “do everything,” leaving humans to find new ways to “participate” in the world.

“That’s basically what we’re getting from the AI CEOs,” Newport says. “And I think it’s just lunacy.” Newport has held this position for some time now. And it’s an opinion shared online by many advanced engineers who have been working with AI systems for a long time. However, there are many so-called social media influencers in the AI space who still push the end-of-the world theme. It’s been an odd experience for sure. Granted, software executives have a long history of bragging in their corporate earnings calls that their systems will enable customers to cut expensive employees. It’s hard to think, though, of another industry whose leaders so cheerfully predict that their products will wreck civilization while making their founders and a few insiders rich beyond their wildest dreams. Seems like a difficult sell, eh?

The New Vibe

In late April 2026, Altman posted on X that OpenAI wants to “build tools to augment and elevate people, not entities to replace them,” and added that “jobs doomerism is likely long-term wrong.” A few days later, Nvidia CEO Jensen Huang pushed back even harder. In a Fortune interview posted May 2, he called the half-of-jobs prediction “ridiculous” and warned that becoming a CEO can leave a person with what he described as “a God complex,” speaking as if their position alone gave them the authority to predict civilization-scale outcomes. Huang also estimated that AI has already created more than half a million jobs, because companies that adopt it grow faster and hire more, and noted that demand for software engineers is actually rising.

Newport reads Huang’s comments as a slightly concealed dig at Amodei, but perhaps it was a subtle signal to the industry that things will be shifting. Either way, the tone has clearly migrated from outright apocalypse to smooth augmentation. Who knows. But at least it’s a welcome change for those directly affected by the recent massive layoffs attributed to AI systems that haven’t even been fully built and deployed yet. 

Where the Doom Came From

To understand the old rhetoric, though, Newport argues that you have to look at the tech culture of San Francisco and Silicon Valley, especially among engineers, and especially content articulated in a few internet forums in recent decades. The most influential was LessWrong, founded by Eliezer Yudkowsky and devoted to refining the art of human rationality. Also, the blog Slate Star Codex, written by Scott Alexander, helped push the same themes into a wider readership. Out of this loose network grew the rationalist movement, which is the idea that if you trained yourself to think like a logical engineer, you could overcome cognitive bias and act more effectively in the world. Newport, who was trained in computer science at MIT, recognizes this culture. “I’m around engineers. I am an engineer. I know this way of thinking,” he says, adding that his wife once told him, “Don’t take me to the MIT Christmas parties because you guys are all so weird.”

Two important offshoots followed. One was effective altruism, which applies something called expected-value reasoning to charitable giving and was made famous — and then infamous — by Sam Bankman-Fried when he led FTX. The other was the existential risk community (X-risk), which argued that very rare disasters with very large costs deserve serious consideration right now. Newport summarizes the X-risk crowd as focusing especially on three threats: asteroid strikes, deadly pandemics, and superintelligent AI. Nick Bostrom’s 2002 paper Existential Risks is the foundational text, but the wider X-risk literature also covers nanotechnology, nuclear war, and what Bostrom calls “totalitarian lock-in.” Newport doesn’t mention Bill Joy’s shock article “Why the Future Doesn’t Need Us” in WIRED in 2000 but it certainly fits the paradigm.

The X-risk crew organized a closed-door conference that produced the open letter signed by Stephen Hawking, Elon Musk, Bill Gates, and many of the leading AI researchers of the day. Newport places it in Puerto Rico in 2017, but the actual event was the Future of Life Institute’s “Future of AI: Opportunities and Challenges” conference in San Juan in January 2015. The 2017 follow-up was the Beneficial AI conference at Asilomar in California. Robert McMillan’s WIRED piece from January 2015, “AI Has Arrived, and That Really Worries the World’s Brightest Minds,” captured the elite anxiety that emerged from the Puerto Rico meeting and may be the article Newport has in mind when he describes the era. There were other similar pieces in the elite media during this time period as well. The elites aren’t shy with the media. 

ChatGPT and the Hero Complex

Then ChatGPT arrived. For people who had spent ten years writing footnoted lists about the coming superintelligence, it felt like the moment they had been preparing for. Newport thinks this was both terrifying and intoxicating. “What if we were right about this risk,” Newport imagines them thinking, “and not only were we right, but it’s happening?” The rationalists sensed they were going to be Neo. They were going to be John Connor. They were the ones who saw it coming and would now lead everyone else through it. It may be hard for normal people to think this way, but we are talking about the tech elite, after all. They do actually live in a different world, one that’s in many ways disconnected from the normal reality of people who have to work for a living. Newport stresses that it’s important to realize that the current AI companies we see now all grew from that culture. 

OpenAI, Newport says, originally presented itself as a nonprofit AI safety organization heavily shaped by X-risk concerns. It was started as almost a hobby project for the rationalist crowd before commercial ambitions reshaped what it is now. Anthropic was founded by former OpenAI staff who, according to Newport, felt their old employer was not being rigorous enough about safety. Grok came out of the same orbit. The CEOs, Newport says, were not playing 4D chess with investors. They were just talking the way everyone they knew talked in Silicon Valley. The trouble started when their companies got too big to keep speaking only to their own closed subculture. As Newport puts it, “we’re not, you know, in the Mission District anymore.”

Why It’s Breaking Now

Newport sees three potential forces that may be accelerating the recent change in AI positioning:

First, there is real IPO pressure building. As OpenAI and Anthropic move toward public markets, more sober East Coast investors (who wear suites, Newport says) are quietly asking the founders to stop terrifying the customers they hope will pay for AI products. 

Second, public opinion is turning. A Quinnipiac poll from March 2026 found that 55 percent of Americans now believe AI may do more harm than good in daily life, which is up from 44 percent a year earlier, with about seven in ten people expecting fewer job opportunities in the future.

And third, journalists are running out of patience. Ezra Klein’s May 2026 New York Times column “Why the A.I. Job Apocalypse (Probably) Won’t Happen” reports that the economists he interviewed are skeptical of mass joblessness. Also, a recent Ronan Farrow piece in The New Yorker even raises the question of whether Altman is actually a strong chief executive for a trillion-dollar company.

Newport says there may be additional pressures in the market pushing AI executives to temper their rhetoric in recent months, but his analysis on the three issues above seems pretty comprehensive as a working hypothesis. 

A Welcome Maturation

The Silicon Valley monoculture has finally collided with the rest of the country, Newport argues. Wall Street realism, journalistic scrutiny, and ordinary public sentiment are forcing the language to evolve to the realities of the market. Newport sounds almost relieved. Somebody, he suggests, finally had to tell these founders to “stop talking like you’re Sarah Connor from Terminator 2.” His understated parting advice still applies. Take AI seriously, but not everything you hear about it.

Cluetrain Yesterday and Today

When I lived in California in late 2000 I read probably the best book on modern communications ever written: The Cluetrain Manifesto by Rick Levine, Christopher Locke, Doc Searls, and David Weinberger. The authors actually teased the book on the web in their 95 Theses in 1999, so we all knew what was coming would be equally outrageous. Couldn’t wait.

Back then I worked in Cupertino, and many times I’d take breaks and walk across the street from the office to get some coffee and flip through the pages. I loved it. It represented a radical departure in marketing and communications at the time because seemed obviously influenced by developers in the rapidly growing Free and Open Source Software movement (FOSS). I was already familiar with most of the topics in Cluetrain because I worked at Sun Microsystems and mixed with FOSS engineers every day. Still, it was cool to see the concepts applied to marketing where these ideas were totally unknown.

Cluetrain made bold predictions about markets becoming conversations, about authentic voice displacing canned corporate messaging, and about employees and customers breaking free from rigid command and control structures. We read it widely at Sun, especially those of us managing FOSS projects. Back then Sun was opening millions of lines of code, so the book gave us useful language as we built projects and engaged development communities. It also came in handy for dealing with the media, which was my primary job at the time. Cluetrain also fit Sun’s culture perfectly because the place seemed at times more like a frat house than a corporation.

The Prediction: Markets Are Conversations

The Cluetrain Manifesto’s central argument was clear. Mass media had interrupted human conversation and turned people into passive consumers and markets into targets. Even within Sun’s generally open culture, there were still power centers in marketing and product development talking in terms of “targeting” the media, developers, and customers. That positioning drove me nuts because the engineers never spoke that way. Cluetrain argued that the Internet would reverse that old paradigm because people could now talk directly to each other about products, communities, and companies across traditionally closed corporate firewalls.

The authors based their conclusions on their own observations. They showed how customers were gathering in online forums to compare notes, share experiences, and help each other get things done. Engineers at companies could openly explain their software to peers in the community and to potential customers without a corporate communications department in between. Companies that tried to control these conversations through legal threats or PR spin got mocked online, often without realizing it, so it fell to people like me to explain what was happening to teams internally. That was a painful process, I can assure you. The pushback from the more conservative teams was significant. But I never had to justify the book’s ideas to developers because in the FOSS community open communication was just considered normal.

But Cluetrain wasn’t simply calling for friendlier or more open marketing process. It was challenging the idea that companies should attempt to control the conversation about them. That was a wasted opportunity in the Cluetrain paradigm. The people building the products always knew more than the marketing department, and the customers using those products often knew things the company itself didn’t. The Internet let those conversations thrive at a scale that had never existed before, and it didn’t require everyone to agree to the latest message. In fact, disagreement was generally part of the value. People could challenge companies publicly and develop their own understanding of a product without waiting for the officially approved marketing campaign.

What Actually Happened

The first part of the prediction proved accurate. Conversations exploded online. Developer interactions thrived. Customer and developer reviews became crucial to purchasing decisions. Social media finally gave employees and customers a public voice, and many companies encouraged this change. For example, Sun and Microsoft were among the first large software companies to build blogging platforms for employees to talk with outside communities. Development engineers usually led the way, and within a few years Sun had more than two thousand employees blogging and talking to whoever they needed to talk to in order to get their jobs done. That openness was a real relief from the constraints of corporate life, and for a while the old, one-way broadcast model lost some of its grip. We used to call that old broadcast strategy “megaphone marketing” at Sun, but we were now in a many-to-many world. It was cool.

But the manifesto underestimated how quickly new gatekeepers would emerge, and how differently they would operate. The authors understood the Internet would break down old barriers between companies and their audiences, but they couldn’t have anticipated that the new communication systems would themselves become enormous businesses with their own tools and economic incentives.

Facebook, Twitter, Instagram, YouTube, and other platforms created spaces for open conversation, but they also monetized it. Attention became a product. Personal data became even more valuable than before. Algorithms became necessary to sort through the volume, but those same algorithms also began deciding which conversations got seen. A thoughtful explanation doesn’t necessarily earn more attention than an outrageous claim, and content that produces a strong emotional reaction can be worth much more to a platform than content that’s simply accurate or understated. And we all began to see this in our social feeds. A casual peek at a colleague’s screen would reveal that their online conversations showed an entirely different world based on their preferences that were amplified by the platform’s algorithms.

Bots and fake accounts made it worse and the line between an authentic participant and a manufactured one grew harder to see. And a well funded campaign could manufacture the appearance of agreement without any real support behind it. But developers could always see this. Most early FOSS projects ran on Mailman lists, so those discussions among engineers largely escaped the corporate swamp for years. Over time, though, as more engineers drifted onto social media, they were pulled into the mess too. The Internet didn’t eliminate gatekeepers. It built new ones with far more precision than the old media ever had. Also, whereas the old media tended to be centralized, the new one was atomized and embedded everywhere.

The Transformation of Corporate Communication

Companies adapted to these open changes quickly. They hired social media managers and community managers, built brand accounts, and encouraged employees to act as ambassadors. The language of Cluetrain worked its way into corporate communications and marketing meetings, which made it feel like the old model was changing. But something got lost along the way.

The manifesto called for people to speak from a genuine passion and their firsthand knowledge from working within an open community. However, what companies built instead was a new form of managed authenticity. Posts were getting written by communications teams and checked by legal before going out. Influencers were paid to look organic. Employees were encouraged to speak publicly while at the same time getting detailed guidance on exactly what they could and couldn’t say. Some companies appeared to be part of the conversation while also working to keep as much of the control as possible.

But other companies were more genuine about it. They admitted mistakes and let employees talk freely and participate in the communities. At Sun we aired plenty of our dirty laundry in public blogs, and the community noticed that and appreciated it. Several Microsoft bloggers did the same. But those were exceptions rather than the rule, and over time most companies learned to fold blogs and social media into the same communications machine they’d always run. Once authenticity became something a company could measure and manage, it turned into just another item in the corporate toolkit. It was wild to see people in marketing meetings starting to borrow the language engineers had used for years, at least on a surface level.

The Complexity of Authentic Voice

The manifesto also underestimated how hard an authentic voice is to maintain inside an large organization. It’s easy to say employees should speak honestly. It’s harder when those employees have managers, legal departments, and reasonable worries about what happens if they say the wrong thing publicly. People are complicated. They carry competing loyalties, worry about their careers, get tired, make mistakes, and sometimes just disagree with the company they work for.

Also, authentic doesn’t automatically mean good, either. A customer’s honest opinion can be unreasonable. An employee’s honest take can rest on a bias they don’t even recognize. A developer can be technically brilliant and still be a troll. FOSS communities learned this a long time ago, which is why so many projects eventually had to write and enforce conduct policies. Openness alone was never enough, especially when systems scale.

The manifesto also didn’t foresee how authenticity itself would become a commodity. Brands hire consultants to develop an authentic voice. Corporations spend real money trying to appear genuine, and influencers get paid to look like ordinary people making ordinary recommendations. As a result, people learned to perform realness the way they once performed professionalism, and once everyone is trying to appear authentic, authenticity just becomes another style of engagement. This is just a normal process of how organizations full of complicated people operate. It’s to be expected.

What the Manifesto Got Right, and What It Missed

Despite these blind spots, Cluetrain identified something real about how business would change. Markets did become more transparent, and information asymmetries did shrink. Customers and development communities gained real leverage over corporations. Microsoft’s own path from calling Linux a cancer to actively engaging open source projects is a good example. Companies that ignored what customers were saying suffered brand damage faster than they would have in the old broadcast era. And the hyperlinked organization the authors described turned out to be real to a certain degree. Work became more distributed. Information stopped moving strictly through the hierarchy. A developer could talk directly with a customer, and an employee could find the person who knew how to solve a problem without asking permission to cross multiple layers of management. Some managers didn’t love this, of course, but they couldn’t really stop it. Many even saw Cluetrain as an opportunity for career advancement and skill development.

What the authors missed, though, was that they underestimated how well old power would simply adapt. They believed the Internet would make hierarchy less important, and in some ways it did, but hierarchy didn’t disappear. It mostly moved online. They also underestimated scale. Small communities can sustain honest conversation because people get to know each other quickly. They learn who is helpful, who is selling, and who has a real history in the project. That doesn’t necessarily hold once a platform reaches millions or billions of people, though. Context breaks down, algorithms fill the gap by deciding which participants and messages get seen, and gaming becomes easier.

They didn’t foresee algorithmic curation either. The Internet made more information available, but that didn’t mean people would actually see more of it. Two people can use the same platform and end up with entirely different views of what’s happening because the platform is showing each of them a different slice of reality. And they didn’t anticipate how concentrated the new gatekeepers would become. All these years later handful of companies now provide the infrastructure most online conversation runs through. They set the rules, run the algorithms, and collect the data. The promise of decentralized conversation gave way to a new kind of centralized control.

The AI Complication

Then we got artificial intelligence, which is where Cluetrain gets interesting again. The authors believed people could tell the difference between corporate language and the way real people actually talk. That held up fine in 1999. It’s much harder to claim today, though.

AI can write a customer review, draft a social post, answer a support question, or generate an image that looks real. And it can do all of this while maintaining a consistent personality for years. Now, fake reviews and manufactured enthusiasm existed long before AI. What’s changed most recently is how cheap, easy, and fast it is to produce content at scale, and that undercuts one of the manifesto’s core assumptions. If authentic voice is what matters, but we can’t reliably tell which voices are authentic, what happens to the conversation? We used to be able to easily spot corporate messaging, but it’s getting harder and harder to spot AI messaging as those systems become pervasive.

It gets more interesting once you stop thinking about individual fake messages and start thinking about participants themselves. An AI system can answer questions, defend a company, criticize a competitor, and take part in a developer community to the point where people start developing relationships with it. But human reputation develops through repeated interactions where you have something at stake and have to live with what you said. But now an AI agent can simulate the pattern of that outcome without having any of the human substance behind it. Its apparent expertise actually came from a model trained on other people’s work, and its apparent reputation really belongs to whoever operates it either at the LLM or the harness level. That’s a bigger challenge to Cluetrain than simply noting that AI can write marketing copy.

Automation Is Not AI

There’s a distinction worth making here because it sometimes gets lost in many discussions about AI. Automation has been removing humans from transactions for decades. A vending machine does it. An ATM does it. A self checkout does it. None of those systems need AI at all.

I was reminded of this recently while clothes shopping in Tokyo. I expected a clerk to check each item, scan or key in the price, and fold everything neatly for me That’s usually how it goes here. This is Japan, after all. Direct human service matters. A lot. Instead, I dumped a tangled pile of ten items into a bin, and the system identified everything immediately, listed everything on a screen, and let me pay with a quick tap of my card (I love the security on that last part). No human involved anywhere in the process, and no bar codes for me to scan myself. It just worked. What matters here is that this kind of checkout doesn’t need AI to be impressive. RFID (Radio Frequency Identification) can identify a pile of products without anyone scanning them one by one, which is what makes the example useful beyond AI specifically. Automation has been removing people from physical transactions for a long time. We’re all used to it now even though we may not like it. This may help explain why people are so sensitive about AI because AI now extends that same trend but into our conversations. Now an AI agent can replace the conversation itself.

Where Cluetrain Still Works

I still see the original Cluetrain principles working in the FOSS projects I’m part of, though. A developer asks a question on a mailing list, and other developers respond with real knowledge and no filter. They argue, share code, make mistakes, correct each other, and build things together because there’s context and an expectation of openness and credibility. People develop reputations in these spaces. You learn who has expertise, who tends to cause trouble, and who is genuinely trying to help. That history and experience over time gives the conversation weight that a comment in a text box on a website can’t create on its own.

I’d revise one part of the usual argument here, though. It’s not quite right to say Cluetrain works at human scale and fails at platform scale, since some gigantic communities function well and some tiny ones are dysfunctional. What actually matters is context, reputation, shared purpose, and accountability. A local bookstore can have those things. So can a software consultancy, a FOSS project, or a single team inside a much larger company. But Cluetrain breaks down once a conversation gets so large and impersonal that participants lose the ability to build context and reputation, while at the same time the platform itself gains more say over what everyone sees.

What We Need Now

The Internet did change business, and authentic conversations did become more important. But the story didn’t unfold as cleanly as Cluetrain predicted. Markets became conversations, sure, but many of those conversations are now mediated by platforms with their own economic incentives. Our content is now run through algorithms built to hold our attention, and our conversations are increasingly populated by AI that can imitate humans well enough to fool most people who aren’t looking closely.

That raises real questions. Does it matter whether the voice answering you is human as long as it gives you what you need? Many times it doesn’t matter. And that’s fine. But in most cases, people tend to seek human interactions. The real question is how markets will function at scale if we really can’t tell who, or what, we’re talking to. We don’t know the answers yet since all of this is still emerging. Some companies are being transparent about their use of AI and are disclosing when people are talking to a bot. Some companies are now saying that AI is just a tool to augment humans rather than a replacement for humans. Others, however, will remain less transparent about their use of AI simply because the incentive to cut human involvement is massive. Those companies seem intent on pushing the narrative that AI will just replace people in the future.

I don’t think AI will destroy human conversation or replace people entirely. Markets are changing, that’s true, and many people will be displaced, but that’s been happening forever. Nevertheless, People will keep seeking relationships, expertise, and direct contact with other people. That’s just an innate human characteristic. And the more artificial our digital environments become, the more valuable genuine interactions will get. That may be part of why live conferences still draw big crowds. When you’re standing in a room talking to someone, you don’t have to wonder if an algorithm or someone bot produced the experience. It’s just real.

That’s what Cluetrain ultimately got right. Markets are conversations because markets are made of people. The Internet made those conversations possible at a scale no one could have imagined in 1999, but it also built new gatekeepers, new incentives, and new questions about who or what we’re actually talking to. I still talk about Cluetrain at developer conferences, even though almost nobody in the room has heard of it anymore. The book may be mostly forgotten, but the basic idea holds. We should still want knowledgeable people talking directly with customers and developers talking directly with their peers. And we should still want communities where reputation gets earned through real contribution rather than assigned by algorithms that can be gamed.

AI’s Perpetual Present

I’ve been reading “Why We Need Continual Learning” by Malika Aubakirova and Matt Bornstein recently. I also listened to a podcast interview from Malika on a16z . Now, I’m no AI researcher or developer. But I do like exploring the scientific foundations on which advanced software tools are built, especially since I use these applications every day and hope to leverage them more in the future. So although I don’t fully understand what’s actually happening underneath, poking around a bit is an interesting exercise. What follows below is what I’ve learned from the article. Consider it a work in progress. If you want the expert version from Malika and Matt, go read their original piece for a deep dive. This text here is just me working through things as best as I can at my level. At the end of this post, I include a list of terms and definitions. I’ll make that a standard feature in similar upcoming posts for my own short-term recall practice and also for long term memory consolidation. Memory practice (the human kind) is a hobby of mine.

Anyway, here we go. The authors open their article on continual learning by referring back to Christopher Nolan’s “Memento,” which is a film about a man named Leonard Shelby who suffers from anterograde amnesia that prevents him from forming new memories. Every few minutes his world resets and he wakes up in the same perpetual present with no idea what just happened in the past. He tattoos notes on his body and carries Polaroids as memory aids just to function throughout the day. It turns out that he’s very resourceful because he uses whatever he can in his environment to get by. He even appears pretty capable within any given scene in the movie. But, as the authors put it, his tragedy is that “he can never compound. Every experience remains external.” So, I guess that means he can’t learn based on his present moment to prepare for the future like most of us who have normal memories.

That seems to be a good general description of where AI models are right now. Back before I knew about this issue, I actually inadvertently tripped over it when I first used ChatGPT and Grok a few years ago. It was clear from my chats at the time that the models were not “learning” from our conversations at all. I kept spinning around in circles explaining myself over and over again. And, in fact, some of those earlier models didn’t know even basic facts from current events, which was shocking since AI was sold to us as being so super smart. That’s when I realized that the “learning” for LLMs took place at some point in the past and then they were locked shut while life continued on. That experience of an AI not knowing simple bits in the news rarely happens now so the user experience has improved significantly. However, there’s a lot more to it that I didn’t realize from those first few frustrating conversations.

What’s Actually Happening When You Type Into That Text Box

Here’s what I didn’t fully understand before reading the article. When you type into a chat window and stuff happens before you get an answer, that process is not the model learning anything from your input. It’s reading what you gave it and generating a response. When the conversation ends, the model does not carry that conversation forward in its memory. The next conversation starts from exactly the same place as every other new conversation. Initially, that felt unnerving so I had to figure out ways to leverage the knowledge from the LLM without all that forgetting going on.

The text box we type into is just a door into the system. What matters is the context window behind the door, which is everything the model can see at once. So, your message, the whole conversation history, any documents you shared, and any background instructions — all of these things represent what the model is working with when it responds. And it has a size limit. When it fills up, older content gets dropped to make room for new content. So if you spend an hour explaining your company’s internal processes to an AI assistant and then start a fresh conversation the next day in a new text box, the AI has no memory of the previous conversation. You have to start over. Not because it forgot. Because it never learned in the first place.

There’s a name for this phenomenon. The article calls it in-context learning, which is really just the model making smart use of whatever sits in front of it right now. It’s temporary by design. The model reads, responds, and moves on. It’s similar to glancing at your notes before a meeting rather than actually deeply studying, internalizing, and using the material beforehand. When the meeting ends, those casual notes go back in the drawer and are forgotten.

The Frozen Model Problem

To understand why this matters, you need to know a little about what’s inside these models. During training, a model reads an insane amount of text and gradually adjusts billions of numerical values called parameters or weights. You can think of each weight as a dial on a pipe connecting two nodes in the network controlling how much signal flows through. The model trains by turning billions of those dials very slightly over and over again until it gets good at predicting language. That right there is really impressive to me given the scale of information these models are working with. But when the training process ends, all those dials get locked. That stage represents deployment. The model then goes out into the world with its knowledge frozen in place.

Training works because it’s a compression process. The model can’t store everything it reads verbatim. It has to find the underlying patterns, generalize the data, and build something compact that transfers to new situations it’s never seen before. The authors describe this as lossy compression, and that lossiness is actually what produces what seems like intelligence to us when we talk to an AI. When I first read that I thought of a camera compressing a RAW file to a JPEG file. The RAW image contains all the available data but it’s a massive size and requires editing in post production to produce a beautiful image. The JPEG, however, is much smaller because it’s been compressed by the camera to just what’s needed to display a good quality image at a certain size. I’ve always understood that process in photography, but I didn’t realize that LLMs are going through a similar process.

Here’s another way to think about it. Remember when you first learned how to ride a bike? You didn’t read the entire manual every time. You just got some guidance from a friend or a parent and you practiced. You fell down a few times and adjusted your technique, and then eventually your brain distilled your experience into something automatic and compact. That’s compression. You still remember falling down, but that falling down process is no longer helpful for riding once learning has taken place. What remains is the final skill of balancing to ride. An AI model that memorizes every training sentence perfectly would be less useful, not more, because it could retrieve but never generalize. It would behave more like a simple retrieval system than a sophisticated learner.

The painful irony the authors identify is this. The very mechanism that makes these models powerful during training is exactly what we stop them from doing once they’ve been deployed. We freeze the compression at the moment of release and replace it with what’s called external memory. That clarified the argument for me. The system is layered, and each layer is essentially a workaround for the fact that the compression stopped. Understanding that made the next part of the article click.

The Filing Cabinet

To compensate for frozen models, developers have built elaborate scaffolding systems, such as chat histories, retrieval databases, system prompts, external document stores, and more. All of these things make up what the article calls external memory. They are flexible and they live outside the model’s internal, frozen weights. When you need information, the system retrieves it and feeds it into the context window. Then the model reads it and responds.

This architecture works as is and the authors are honest about that. However, they make a point I hadn’t considered before. “A bigger filing cabinet is still a filing cabinet.” Retrieval is not learning. The model is looking things up, not actually knowing them. It just does it very quickly and uses natural language so you get the impression you are talking to someone who is intelligent.

Here’s another practical example. Say a hospital deploys an AI assistant to help with real world clinical decisions. That model was trained on medical literature through some cutoff date. A major new clinical trial or medical policy comes out afterward that changes how doctors treat a particular condition. The hospital can feed that paper into a retrieval database so the AI can surface it when it’s relevant. But the model doesn’t internalize that new research the way doctors would after reading it, applying it to patients, observing the outcomes, and revising their practice accordingly. The AI can retrieve the abstract. But it can’t reason from the new finding the way someone who has truly learned it can in practice. That’s the limitation these researchers are trying to fix.

The same problem exists in cybersecurity with treats evolving daily. A frozen model can be given descriptions of new attack patterns through retrieval, but it can’t compress and generalize from those patterns the way an analyst does who has spent months chasing a specific class of threat. The knowledge stays external. It never becomes part of what the model actually knows unless the model is updated with a new learning process, which is time consuming and very expensive.

What Real Learning Requires

So what’s the alternative? The article introduces a concept called continual learning, which is the field of research aimed at letting models actually update their weights based on new experience after deployment. Not just read notes. Actually learn live like humans do.

And here’s where the Memento metaphor really makes sense. The authors say that today’s AI is stuck in Leonard Shelby’s perpetual present. The scaffolding, the Polaroids and tattoos, and other memory aids work well enough within any given scene. But the model can never compound in real time. Every new thing it encounters stays external.

Think about the difference between a doctor who simply retrieves a recent study and a doctor who has spent years treating patients with that knowledge fully and personally internalized. Or consider the difference between someone who has your email history in front of them and someone who actually knows how you think over time. The article frames this cleanly. “The difference between ‘Here is what you responded to this email before’ versus ‘I understand how you think well enough to anticipate what you need’ is the difference between retrieval and learning.” Even in normal human memory, immediate retrieval is necessary to manage your present experience. However, it’s also required that your present experience be embedded into long term memory for continual learning.

The authors bring up Fermat’s Last Theorem as one powerful example of the kind of hard discovery problem they have in mind. Mathematicians worked on the issue for 350 years. Eventually the problem was solved by Andrew Wiles. But he didn’t crack it by retrieving the right papers. He solved it by working in near total isolation for seven years, and inventing entirely new mathematical techniques to bridge two previously disconnected fields. That kind of discovery required genuine compression, generalization, and creative combination. Not simply fast retrieval. And the article asks directly whether a model that can’t compound from experience could ever do anything like that. The honest answer is they don’t know yet.

Why Updating Weights Is So Hard

At this point I had to ask myself if real time continual learning is so important, why can’t the LLM models do it now? The short answer is that updating a model’s weights after deployment is genuinely dangerous and technically unsolved at scale.

The most obvious problem is called catastrophic forgetting. When you update a model’s weights to learn something new, it tends to overwrite what it already knew. New learning crowds out old learning. If you fine tune a general model specifically on medical records, it might get better at clinical language while getting noticeably worse at everything else because the new training has nudged weights that were also doing other jobs. The model gets better at one thing and potentially worse at everything it was already good at. When you understand this you can really appreciate how humans have benefited from millions of years of evolution. The AI machines seem rather clunky by comparison. When humans learn, new neural connections are made in the brain that stick for a long time as new learning is layered on top. But even in humans, old learning and memory does actually fade gradually over time if a specific neural pathway isn’t continually or at least occasionally reinforced. It just takes a very long period of time. With AI systems, however, new learning can wipe out old new learning immediately. The authors didn’t address this issue directly in humans, but the example seems similar if you study biology.

There’s also the problem of data poisoning. If a model’s weights can be updated through interactions after deployment, bad actors could gradually manipulate its behavior through carefully crafted inputs over time. Unlike a one-time attack, poisoned weights persist across every future conversation. The damage would live in the model itself so safety alignment would degrade unpredictably immediately or some time in the future. The article notes that “even narrow fine-tuning on benign data can produce broadly misaligned behavior,” which is a sobering thought to sit with. Yet we all know this would happen right away based on our own experience being online every day fighting bots and hackers.

These aren’t hypothetical concerns. They’re real problems without clean solutions yet.

Where Things Are Heading

The article maps out a spectrum of approaches to continual learning that are organized around a question I found clarifying: where does the compaction actually happen? It seems there is a stack of technologies managing the process.

On one end you have pure retrieval. No compaction. The model just reads notes. That’s most of what exists today. In the middle there are modules, which are attachable and specialized components that let a model develop some expertise in a specific domain without retraining the entire thing from scratch. A hospital might attach a medical module to a general model so it performs at a specialist level on clinical questions, while the same base model with a different module handles legal contracts. Each module is swappable independently. That’s a practical and reasonable middle ground for now.

On the far end you have full parametric learning, where the model’s weights actually update from new experience after deployment. This is the goal, but it remains largely unsolved at scale with the current technologies. But there are serious research efforts moving in this direction with things like test-time training where the model runs brief learning cycles before it generates a response. Also there are self-improvement approaches where models like AlphaEvolve have generated their own training data and genuinely improved from it, at least within constrained problem domains like mathematics.

The authors frame the path forward as layered. In-context learning stays as the first line of adaptation because it works now and keeps getting better. Modules offer some personalization and domain specialization. But for genuinely novel problems, adversarial scenarios, and knowledge too tacit to put into words, models may eventually need to compress new experience directly into their parameters after training. Otherwise, as the authors put it, we stay stuck in Memento’s perpetual present.

What I Took Away

I started reading this article as someone who uses AI tools every day without really thinking much about what’s happening underneath. What I came away with is a better sense of the gap between what these systems appear to do, what they’re actually doing, and what they’ll potentially do in the future. Right now they can respond to new information and adapt to what you give them. And most times they feel like they understand you. But the reality is that they don’t compound. They don’t learn. They don’t internalize new experience the way continual learning systems or humans would. Their dials are locked. And until engineers figure out how to update those dials safely and continuously after deployment, the models we’re using now are doing something more like reading notes than actually learning from the experience. That’s a distinction with a very big difference.

Check out the original article and Malika’s podcast for the technical details. Below is a list of related terms and definitions.


Continual Learning: Vocabulary List

This list of terms below is based on the a16z article “Why We Need Continual Learning” by Malika Aubakirova and Matt Bornstein and also the podcast with Malika discussing the article. Some definitions closely reflect the article itself, but others expand into broader concepts from the field for additional context. I error checked the terms and definitions with Grok, ChatGPT, Gemini, Perplexity, and DeepSeek.

Agentic Loops

A mode of operation where the model works autonomously step by step toward a goal without you typing each instruction. Each step produces output that feeds into the next. This process can go on for many cycles. The article identifies two related problems as steps accumulate: (1) the immediate symptom is coherence degradation, where the agent loses the thread and starts making poor decisions, and (2) the underlying cause is that maintaining a growing context becomes increasingly expensive and inefficient. Both concerns together represent why the article frames agentic loops as one of the pressure points on the current in-context learning paradigm. For example, an agent tasked with researching a topic, drafting a report, checking sources, and revising the draft might handle the first twenty steps cleanly. But by step eighty the accumulating context has grown so large and costly that the agent starts losing track of earlier decisions and repeating work it already did.

Attention Heads

A key mechanism inside transformers that allows the model to weigh how relevant each part of the context is to every other part when generating a response. Multiple attention heads run in parallel, each learning to focus on different kinds of relationships in the text. One head might learn to track grammatical agreement between subject and verb across a long sentence, while another tracks thematic connections between paragraphs. Together they allow transformers to handle complex, long range dependencies in language that earlier architectures struggled with. For example, in the sentence “The lawyer who argued the case, despite the objections raised by her colleagues, ultimately won,” an attention head helps the model correctly connect “won” back to “lawyer” across all the intervening words.

Catastrophic Forgetting

When a model updates its weights to learn something new, it tends to overwrite what it already knew. In other words, new learning crowds out old learning and sometimes dramatically. This is one of the central unsolved problems in continual learning, and one of the main reasons models are not updated continuously after deployment. Think of it somewhat like overwriting parts of a hard drive. The new files go in, but the old ones can be partially or fully lost. For example, if you fine-tune a general purpose model specifically on a medical records archive, the model will get better at clinical language but noticeably worse at writing poetry or explaining history because the new training has nudged weights that were doing other jobs.

Compression / Compaction

The process of taking a vast amount of raw information and distilling it into something compact and generalized. During training, a model compresses an enormous amount of human writing into its parameters and finds the underlying patterns rather than storing things verbatim. The article uses “compaction” as a broad organizing term for how deeply new information gets digested, which ranges from not at all (pure retrieval, where facts just sit in a database) to fully (weight-level learning, where the model actually internalizes new knowledge). For example, rather than memorizing every recipe ever written, a well-trained model compresses the underlying logic of cooking: how heat transforms food, how flavors balance, how techniques generalize across cuisines.

Continual Learning

The broader field of research aimed at letting models learn from new experience after deployment, ideally by updating their weights rather than relying on external scaffolding. It’s the opposite of the current norm, where training and deployment are completely separate and weights are frozen the moment a model is released. The goal is something closer to how humans learn continuously from experience without needing to be retrained from scratch every time the world changes. For example, a customer service model using continual learning could gradually internalize patterns from thousands of resolved support tickets over time and get genuinely better at its job rather than just retrieving past examples.

Context Window

The full body of text the model can see at once when generating a response. It includes your message, the full conversation history, any documents you shared, and any background instructions passed to the model. It has a size limit measured in tokens. When it fills up, older content must be dropped to make space for new content. For example, if you have a long conversation with an AI assistant and then ask it to recall something you mentioned earlier, it may not be able to answer because that part of the conversation has already been pushed out of the window.

Data Poisoning

One of several serious governance and security risks the article raises around continuous weight updates. If a model’s weights can be updated after deployment interactions, bad actors could gradually manipulate its behavior through carefully crafted inputs over time, which is a slow and hard-to-detect form of corruption that lives in the weights rather than just in the context. Unlike a one-time prompt injection attack, poisoned weights persist across every future conversation. The article groups this alongside other unsolved challenges: alignment degradation, the impossibility of unlearning toxic knowledge, auditability failures, and privacy risks from user interactions being compressed into parameters. For example, an adversary could repeatedly feed a customer-facing AI subtly misleading information about a competitor’s product until the model begins reproducing those inaccuracies on its own with no obvious sign of tampering.

Distillation

A process involving two models: (1) a large, capable, frozen teacher and (2) a smaller student. The student is trained to match the teacher’s outputs as closely as possible and absorb its knowledge in a more compact form. The result is a smaller, more efficient model that performs nearly as well as the larger model on the tasks it was trained for. It’s like an apprentice learning by closely watching and mimicking a master until the skill becomes their own. For example, a large hospital system might use a massive general-purpose model as the teacher and distill its medical reasoning capabilities into a smaller model that can run efficiently on local hospital hardware without requiring a cloud connection.

External Memory

Anything outside the model’s weights used to store and retrieve information. Chat history, databases, document stores, and agent notes are all examples of external memory. Information gets fed back into the context window when necessary. In current deployment architectures, the model typically does not update its weights from that information during inference. The key limitation is that external memory requires retrieval. The model has to be given the right information at the right moment, and if it isn’t, the knowledge might as well not exist. For example, a legal AI might have a database of ten thousand case summaries it can search, but if the retrieval system surfaces the wrong cases, the model has no way to compensate from its own knowledge.

Few-Shot Learning

The ability of a model to perform well on a new task after seeing only a handful of examples, rather than requiring thousands of training samples. Transformers are surprisingly good at this when examples are provided in the context window. Meta-learning approaches aim to make weight-level, few-shot learning just as effective, so the model can internalize new tasks from just a few examples even without them being available in the context. For example, if you show a model three examples of how you want your emails formatted and then ask it to format a fourth, it adapts immediately without any retraining. That’s few-shot learning in action.

Fine-Tuning

A more targeted form of additional training done after the initial training run. Instead of training from scratch on everything that’s known, you take an already-trained model and update it on a smaller or specific dataset. The new information shapes the model’s behavior for a particular use case without rebuilding it from the ground up, but the process still risks catastrophic forgetting if pushed too hard. For example, a company might take a general-purpose language model and fine-tune it on thousands of their internal support conversations, so the model learns the company’s terminology, tone, and common issue patterns without losing its broader language capabilities.

Gradient Descent

The mathematical process by which a model adjusts its weights during training. It measures how wrong the model’s predictions are on a given example and then calculates which direction to nudge each weight to reduce that error slightly. It’s called “descent” because the process is navigating downhill on a mathematical landscape, always moving toward lower error rates. Repeat this across billions of examples and the model gradually gets much better. For example, if the model predicts “cat” when the correct answer is “dog,” gradient descent works backward through the network to figure out which weights contributed to that wrong answer and adjusts them a tiny amount. Do that enough times and the model learns to tell cats from dogs reliably.

In-Context Learning (ICL)

Everything the model reads and uses during a single conversation without updating its underlying knowledge. You paste in a document, it reads it and responds. You describe a task, it follows your instructions. But when the conversation ends, none of that experience changes the model itself. The next conversation starts with the same frozen weights as always. This is a smart use of temporary information, but it’s not genuine learning. For example, if you spend an hour teaching an AI assistant about your company’s internal processes and then start a new conversation the next day, the model will have no memory of the previous conversation. You would need to paste in that information all over again.

Inference

The act of a model generating a response from input. It’s the opposite of training. Training occurs when the model learns by adjusting its weights. Inference occurs when the frozen model performs and takes what it knows and produces an output. Any time you send a message and get a reply, that’s inference. The term “inference-time compute” (below) builds on this and refers specifically to spending extra computational effort during inference to get a better result. But plain inference just means the model is running, not learning. For example, asking a model what the capital of France is and getting back “Paris” in a fraction of a second is inference in its simplest form. No learning took place. The model generated an output from its existing weights without updating them.

Inference-Time Compute

The current dominant paradigm for improving model performance by spending more computational effort at the moment of response rather than updating weights. This includes chain-of-thought reasoning, tool use, search, and iterative problem-solving, all of which cost more compute at response time but produce better results. The article positions this process as a workaround, a scaling of what already works rather than a true solution to the learning problem. Test-time training is the most aggressive form of this learning because it actually runs gradient updates on new information during inference, which begins to compress it into weights in real time. This process sits at the boundary between the current paradigm and genuine parametric learning. For example, when you ask a model a complex math problem and it works through each step before giving a final answer rather than just guessing immediately, that is inference-time compute. The model is using more processing in the moment to arrive at a better result.

Instruction Tuning

A form of fine-tuning where the model is trained specifically on examples of instructions paired with ideal responses. It’s one of the main reasons modern models are so much better at following directions than earlier versions, which tended to just complete text rather than actually do what you asked. The model learns not just facts but the shape of helpful behavior, including how to interpret requests, how to structure answers, and when to ask for clarification. For example, an early language model asked to “summarize this article” might just continue writing in the same style as the article. An instruction-tuned model understands that the request calls for a concise, distinct summary and produces one.

KV Cache

Short for key-value cache. A technical mechanism that stores intermediate computations during inference so the model does not have to redo them from scratch for every token it generates. The article discusses it specifically in the context of KV cache compaction where the cache functions as a form of non-parametric memory but grows substantially as conversations and agent loops get longer. The authors argue that learning to compress this cache more efficiently is one of the meaningful challenges in moving from pure retrieval toward more durable knowledge storage. For example, in a long agentic task, the KV cache holds the computed representations of everything the model has processed so far. Without it, each new token would require reprocessing the entire history from scratch, which would be prohibitively slow.

Lossy Compression

Compression where some information is permanently lost in the process, as opposed to lossless compression where everything can be recovered exactly. For LLMs, the inability to store everything verbatim during training forces the model to find patterns, generalize, and abstract. That forced abstraction is precisely what makes the model seem intelligent and useful in new situations it has never seen before. A JPEG image is the familiar everyday example. Save a photo as a JPEG and the file shrinks dramatically because fine detail is discarded. But if you zoom in close enough you can see the degradation. For most purposes, though, the image is perfectly usable. The tradeoff is the point. For a language model, the equivalent is that it cannot recite every sentence it ever trained on, but it can write a new sentence in any style on any topic because it extracted the underlying structure rather than memorizing the surface.

Meta-Learning

Teaching a model how to learn rather than what to learn. The model is pre-trained in a way that positions it to update quickly and effectively with just a few new examples, rather than requiring extensive retraining. It’s the difference between educating someone to be a quick study versus simply giving them a lot of facts to memorize. A quick study can walk into an unfamiliar subject and get up to speed fast, whereas someone who only memorized facts cannot. For example, a meta-learned model shown three examples of a new classification task, say sorting customer complaints into categories it has never seen before, should be able to generalize accurately to new complaints after just those three examples rather than needing hundreds.

Modules

The article uses this as a broad middle-ground category on the compaction spectrum that sits between pure retrieval and full weight-level learning. In practice, modules can take several forms: adapter layers, LoRA-style weight updates, memory components, or cached representations. What they share is the ability to specialize a general-purpose model for a specific domain without retraining the entire model from scratch. They offer more than retrieval in that some digestion of information happens, but less than full parametric learning in that the core model is typically left unchanged. For example, a hospital might attach a medical module to a general-purpose model so it performs at a specialist level on clinical questions, while the same base model with a legal module performs at a specialist level on contract review, with each module being swappable independently.

Multi-Agent Architectures

Systems where multiple AI models work in parallel with each one handling a slice of a larger task and communicating results to each other or to an orchestrating layer. If a single model is limited by its context window, a coordinated group of agents can collectively handle far more. But this shifts the problem rather than eliminating it. Each agent still faces its own context limit, and coordinating many smaller contexts introduces its own complexity for the system to manage. It’s a non-parametric workaround for scale, not a solution to the underlying constraint. For example, a research task that would overflow one model’s context window might be split across ten agents, each reading a different section of source material with a coordinating agent assembling their summaries into a final report.

Neural Network

The underlying computational structure of an LLM. It’s a network of interconnected nodes organized in layers, loosely inspired by neurons in the brain. But the analogy should not be pushed too far. Each connection between nodes has a weight that determines how strongly one node influences another. During inference, information flows forward through the layers, gets transformed at each step, and eventually produces an output. The network learns by adjusting those weights during training until it gets good at its task. For example, in an image recognition network, early layers might learn to detect simple edges and colors, middle layers might learn to recognize shapes, and later layers might learn to identify objects. Language models work on the same principle but applied to sequences of text.

Parameters / Weights

The billions of numerical values inside a model that encode everything it learned during training. Each value represents the strength of a connection between two nodes in the neural network. During training, these values get adjusted gradually until the model becomes good at predicting language. After training they are frozen, and the model’s knowledge and capabilities are entirely determined by those fixed numbers. “Parameters” and “weights” refer to the same thing and are used interchangeably throughout the article. For example, frontier models contain billions or even trillions of parameters. Each one is a small dial that was tuned during training and now stays locked in place, collectively encoding an enormous amount of compressed knowledge about language, facts, and reasoning patterns.

Parametric Learning

Learning that actually updates the model’s weights based on new experience, as opposed to in-context learning which uses information temporarily without changing anything permanent. It’s the deeper form of learning the article is ultimately arguing we need more of. When a model learns parametrically, new knowledge gets compressed into its weights the same way training data did and becomes a durable part of what it knows rather than a note it holds briefly and then discards. For example, a parametric update after a model encounters thousands of conversations about a new programming language would leave it genuinely better at that language going forward across all future conversations, not just within the session where it learned.

Regularization

A cautious approach to weight updates that penalizes changes to parameters deemed important to existing knowledge. Before updating a weight, the system estimates how critical that weight is to the model’s current capabilities. If it’s very important, the update is constrained or slowed down. This is one of the older approaches to continual learning and helps manage the stability-plasticity dilemma. But it tends to be brittle at scale. Think of it like a renovation rule that protects load-bearing walls. You can still remodel, but certain structures are off-limits because removing them would collapse the building. For example, EWC (Elastic Weight Consolidation), one of the most cited regularization methods, computes an importance score for each weight after training on a task and uses that score to resist changes when training on subsequent tasks.

Reinforcement Learning (RL)

A training approach where a model learns from feedback signals rather than from labeled examples. It tries things, receives a reward or penalty based on how well it did, and adjusts its behavior accordingly over many iterations. The article mentions RL-based feedback loops as one direction in continual learning research where models could improve from real-world deployment signals like user corrections or task outcomes. However, it’s not the central mechanism the authors emphasize. The core focus of the article is on compaction, weight updates, and memory structures. For example, the systems that learned to play chess and Go at superhuman levels used reinforcement learning by playing millions of games against themselves and adjusting strategies based on wins and losses rather than being taught explicit strategies.

Retrieval-Augmented Generation (RAG)

A common approach to giving models access to current or specialized information without retraining. Instead of baking knowledge into weights, you build a searchable database the model can query at response time. The retrieved content gets injected into the context window and the model uses it to generate its answer. It’s purely non-parametric. The model retrieves information but never internalizes it. The limitation is that retrieval only works if the right information gets surfaced at the right time, and no amount of retrieval can substitute for knowledge the model needs to reason with flexibly. For example, a financial AI might use RAG to pull in the latest earnings reports before answering questions about a company’s performance because that information changes constantly and cannot be baked into training data.

Safety Alignment

The work done during training to make a model helpful, honest, and safe to use. It involves carefully curated training data, human feedback on model outputs, and specific training objectives designed to shape the model’s values and behavior. One of the serious risks of continuous weight updates after deployment is that alignment can degrade unpredictably even from adding seemingly benign new data. It seems that fine-tuning on almost anything can shift the weights that govern behavior, not just the ones governing the specific knowledge update. For example, researchers have shown that even brief fine-tuning on ordinary instructional text can weaken safety guardrails in ways that are not obvious until the model is probed specifically for harmful outputs.

Self-Improvement

An approach where the model generates its own training data, filters out low-quality results, trains on the high-quality results, and repeats the cycle. It learns from its own work rather than from human-provided data and can improve capability over repeated iterations in constrained settings. The article cites AlphaEvolve and AlphaProof as examples of this kind of closed-loop improvement. But these systems operate in constrained domains like mathematics and algorithm optimization, not open-ended real-world learning. The article uses these examples to illustrate iterative self-training loops, and what qualifies as a genuinely new discovery in this context remains debated. For example, AlphaEvolve used self-generated solutions and automated evaluation to discover improvements to algorithms that human programmers could not find because it worked within a well-defined problem space where correctness could be verified automatically.

Stability-Plasticity Dilemma

The fundamental tension in any learning system between staying stable, meaning not forgetting what it already knows, and staying plastic, meaning remaining able to learn new things. Push too hard toward plasticity and you get catastrophic forgetting. Push too hard toward stability and the model cannot adapt to anything new. Solving this dilemma is one of the core engineering challenges in continual learning, and no approach has fully solved the problem at scale. The dilemma exists in biological brains too. From what I understand about biology, human memory consolidation is strongly associated with sleep and offline processing, which suggests the brain has its own version of this stability-plasticity problem built right in. For example, a model trained to be highly stable might refuse to update its belief that a particular drug is safe even after being shown new clinical evidence, while a model trained to be highly plastic might update so aggressively that it forgets basic grammar rules after a week of medical fine-tuning.

State Space Models (SSMs)

An alternative to traditional transformer architecture that the article highlights for offering a fundamentally better scaling profile for long contexts. The article describes them as using fixed memory layers interspersed with normal attention, which unlike transformers does not grow unboundedly with every token added to the context. Traditional transformers scale quadratically with context length, while SSMs aim for near-linear scaling. However, this remains an active area of research rather than a fully settled property. The article treats SSMs as a promising architectural direction for enabling much longer agentic loops rather than a definitive solution to the broader continual learning problem. For example, a transformer handling a 100,000-token conversation requires vastly more compute than handling a 10,000-token request. But an SSM handling the same expansion would ideally require only proportionally more, which could make very long agentic tasks far more practical.

Temporal Disentanglement

A core limitation of parametric memory since a model’s weights do not separate timeless facts from information that changes over time. Both get compressed into the same parameters and are tangled together with no internal label distinguishing what’s permanent from what’s mutable. This makes continual weight updates risky because changing a time-sensitive piece of knowledge can corrupt stable knowledge stored in nearby weights. The article frames this as one of the fundamental unsolved problems standing between today’s frozen models and genuinely adaptive ones. For example, the fact that two plus two equals four and the fact that a particular person holds a particular job title are both encoded somewhere in the weights. Updating the job title risks disturbing the arithmetic, because the model has no mechanism for knowing which facts are stable laws and which are contingent facts about the world.

Test-Time Training

An approach that blurs the line between training and responding by letting the model do a small amount of learning before it generates a final answer. Rather than relying entirely on what it learned during the original training run, the model runs brief gradient updates based on what it’s currently seeing and then responds. The article describes this as running gradient descent on test-time data, compressing new information into parameters at the moment it matters, and treats it as one of the more substantive moves toward genuine continual learning because it actually changes weights at inference time. For example, if a model is asked to analyze a long, unusual technical document, test-time training would let it briefly train on that document before responding, compressing its key patterns into weights rather than just reading it as context. This method potentially produces a much more accurate analysis as a result.

The Bitter Lesson

A well-known observation in AI research. It holds that given more compute and data, general methods that let models figure things out at scale consistently outperform clever human-engineered solutions over time. Every time researchers have tried to hardcode structure and shortcuts into AI systems, the simpler but more scalable approaches have eventually won. The article invokes this phenomenon to question why we still hand-engineer memory and compression pipelines rather than letting models learn to do it themselves. For example, early chess programs used elaborate human-crafted rules about piece values and board positions. They were eventually crushed by systems that simply learned from millions of games with minimal human guidance and relied on scale rather than cleverness. The same pattern has repeated across nearly every domain in AI.

Token

The basic unit of text that a large language model processes. A token is roughly a word, though it can also be a fragment of a word, a punctuation mark, or a short common sequence like “ing” or “un.” Models do not read text the way humans do, character by character or word by word. Instead, they break input into tokens first and then process the sequence. The size of a context window is measured in tokens, not words or characters. For example, the sentence “The cat sat on the mat” would be broken into something like seven tokens, roughly one per word. But a word like “unbelievable” might be broken into two or three tokens: “un,” “believ,” “able,” because it’s less common and gets split into recognizable subunits the model has seen frequently.

Training Run

The large-scale and expensive process of building a model’s knowledge by exposing it to massive amounts of data and adjusting its weights. Training involves feeding these huge datasets through the network repeatedly and using gradient descent to nudge weights toward better predictions. The process runs on clusters of specialized hardware for weeks at a time and consumes substantial amounts of electricity. It’s all carefully controlled, occurs before deployment, and produces a fixed set of weights that define everything the model knows. Once training ends, the weights are frozen and the model goes out into the world as-is. For example, training a frontier model like GPT-4 or Claude is estimated to cost tens or hundreds of millions of dollars and requires specialized data centers. This is precisely why continuous post-deployment learning is so appealing because rerunning a full training run every time the world changes isn’t practical.

Transformer

The dominant architecture underlying most major AI models today including Claude, GPT, and Gemini, and more. At its core, a transformer predicts the next token in a sequence of text based on everything that came before it. It generates outputs token by token at very high speed. That sounds simple but at scale it’s not. The architecture was trained on so much human-generated text that it models statistical relationships in language and attempts to produce behavior consistent with understanding context, logic, and meaning. For example, when you ask a transformer-based model to explain a complex idea, it makes predictions about what a good explanation would look like given your question based on patterns it absorbed from vast amounts of human writing on similar topics. That’s why it seems smart. It’s familiar. Whether the final output constitutes genuine understanding is a separate philosophical debate that the article doesn’t address.