Transcription
It's May 11th, 1997, and Gary Kasparov sits in front of a chessboard on the 35th floor of a Manhattan skyscraper. Across the table sits a chunky monitor and a keyboard. Kasparov is playing chess against a computer in the '90s. IBM's Deep Blue Supercomputer, in the original "Man Versus Machine" test of intelligence. Right now, we're five games into the series. The score: Kasparov won, Deep Blue one, and three draws. This is the final game, six of six. And Kasparov is playing the black pieces. He's the world chess champion in 1997, and no one's been able to touch his title since 1985, a 12-year reign. And 19 moves in, Kasparov resigns. Deep Blue wins the match in series. And for the first time ever, a reigning world champion loses to AI under standard tournament time controls. And the next morning, the world loses their minds. After sudden defeat, it's Kasparov who's blue. LA Times: "Swift and slashing computer topples Kasparov." New York Times. Or my personal favorite, this Matthew Pritchette cartoon.
Remember, this was 30 years ago, before memes and Reddit. People got their news literally delivered to their doorsteps by school children. Computers were around, but not everyone had one, and they weren't the RAM-guzzling machines of today. In 1997, I was three. My dad hadn't even bought his first Gateway desktop yet, but investors were well aware of computers and the internet. I mean, there was nothing hotter on Wall Street leading up to the 2001 tech crash. So, understandably, everyone was freaking out about Deep Blue. So, the most powerful AI demo in the world proves its legitimacy on a global scale. So, how did this not turn into the artificial general intelligence or AGI future that people feared back in 1997? You'd think a breakthrough like Deep Blue would make AI research accelerate faster and faster. Certainly, growth would be exponential now that machines are smarter than the smartest humans. In an AGI world, the limits are only your imagination. We're talking the end of disease, global renaissance, terraforming Mars, automated government and economy, fixing the Earth's ecosystem. But that's not what happened, like, at all. And not because Deep Blue was a hoax. Deep Blue was real, and it was AI. In fact, it's exactly the kind of AI that I studied in my artificial intelligence class at university about 12 years ago. But today, we have a completely different type of AI hyped up in our social media feeds. Now, I'm sorry. I promise we do make plenty of non-AI videos. But the conversation happening today is actually pretty similar to the conversation that stirred everyone up in the 1990s. And today's "growth at all costs" investor mentality is not too different from that. That said, there are plenty of videos talking about the potential of an AI bubble. So, instead of beating that dead horse, in this video, we're going to dive into the technical differences between the type of AI that Deep Blue was in 1997, and what we're working with today with GPT and Claude. If you want to be able to spot the hand-wavy BS that a lot of the biggest AI hype bros are peddling, including some of the CEOs, you need to understand the difference between modern large language models, or LLMs, and what Deep Blue was. So stick around, because we'll be going under the hood, so to speak.
If you look at a zoomed-out view of 1997 to today, information technology does seem to be getting better faster, unless we're talking about a certain operating system that I'm forced to use to play most of my Steam library. But I want to challenge the idea that tech always improves exponentially over time. Sometimes it even slows down, only getting better logarithmically. In other words, there can be diminishing returns. I mean, when was the last time you really noticed a big improvement when you upgraded your iPhone? So, let's open up the Deep Blue machine, talk about the actual computer that IBM built, and go on a tour of the 12-year history behind the old-school AI that it used. We'll talk about why the path seemed to suddenly turn cold before we got to the LLMs that we know and love and hate today. But first, have you ever wished learning to code was a lot more like playing chess? Okay, so boot.dev is really nothing like playing chess, but it is fun like chess. Typically, learning to code has always involved either completing an expensive-as-hell university degree, reading books that are probably out of date by the time you check them out, or watching hours-long tutorials. Gross. Who would ever publish an hours-long tutorial on coding? Okay, nothing's too wrong with any of these methods, but they aren't for everyone. And to us, it really comes down to keeping learning fun. On the boot.dev dev website, we use technologies learned from modern game design to make learning as addicting as possible, in a good way, rewarding you for every challenge that you complete, tracking your progress in each milestone you pass, and pitting you against truly tough, hands-on learning problems. So, if you want to level up your tech skills and learn software engineering and DevOps in a more interactive way, use code BootsTube for 25% off a subscription to all the premium features on the site.
Now, back to the beginning. Before Deep Blue, there was Deep Thought. It's 1985 at Carnegie Mellon University. A grad student named Feng-hsiung Hsu, his nickname in the lab was "Crazy Bird," which is frankly a lot easier for me, starts building a chess-playing chip, like a physical computer chip. It's a simple, tiny chip called ChipTest, based on a design from Ken Thompson, one of the guys who created Unix and one of the designers of the Go programming language, a personal favorite of mine. And it works so well that it actually won the North American Computer Chess Championship in 1987, back when most PC users still didn't own a mouse. Inspired, Hsu keeps going. He builds a better version called Deep Thought. And no, not the one on Magrathea built by hyper-intelligent pandimensional beings. It turns out this Deep Thought can think pretty well. Well, it wasn't actually thinking, but don't worry, we'll get to that in a bit. It became the first chess computer to beat grandmasters in tournament play. So, ChipTest and Deep Thought are often kind of lumped together as the Carnegie Mellon predecessors to Deep Blue. But there are actually some differences that show how Deep Blue essentially grew up over time. It all started with this question: "Can you put a chessboard on a chip?" ChipTest was the answer to that question. A single custom chip plus a Sun3 workstation. Every square of the chessboard was wired up as a circuit so the chip could spit out all legal moves in one electrical operation. The Sun3's job was to handle the search, evaluation, and control in software. In 1986, ChipTest was searching 100,000 positions per second. One year later, it was searching five times as many. Deep Thought used the same basic strategy as ChipTest and multiplied it. The team scaled up to six processors and could search over 2 million positions per second, and they weren't done. Later in 1991, Deep Thought 2 was able to search up to almost 5 million positions per second. But it still wasn't enough. In 1989, eight years still before Deep Blue, Deep Thought played a game against Kasparov, who was at this point the world champion for the last 5 years. But Kasparov crushes it in a two-game exhibition match. Deep Thought loses twice, and it's not even close. Kasparov won the games in 52 and 37 moves respectively, easily seeing through the computer's brute-force algorithm with his decades of human-powered pattern recognition. But by now, a lot of eyes are on Deep Thought. Among those watching closely are reps from IBM, and they like what they see. They hire Hsu right after he finishes his PhD. Then Murray Campbell, Joe Hoane, and Jerry Brody are added to finalize the four-man technical team. Later, they also brought on Grandmaster Joel Benjamin as a chess consultant. All that's left is to give their project a name, and they go with Deep Blue, a play on Deep Thought and IBM's nickname, "Big Blue."
Here's the thing about actually building Deep Blue. This isn't really a software project. Hsu is designing custom silicon, and not just one chip. There are 480 of them. Each one simulating a tiny chessboard, an 8x8 grid of logic gates that generates moves and helps evaluate position in hardware. Yeah, it's running algorithms and data structures, but a lot of its speed comes from silicon rather than software optimization alone. Those 480 chess chips sit inside an IBM RS/6000 SP supercomputer alongside 30 general-purpose processors. So, the chess chips handle the brute-force, chess-specific algorithm, and the main processors handle high-level coordination. It's a hybrid: part general-purpose computer, part custom chess machine. And building the thing takes years. Testing it takes even more years. Joel Benjamin's job is to play the machine over and over and over again. He finds the positions where it does something stupid, and he helps the engineers patch those weaknesses. The team also builds the opening book, a library of known strong moves for the first phase of a chess game. By 1996, Deep Thought is old news, and Deep Blue is ready for a Kasparov rematch. The very first game of the six-game series is insane. Deep Blue wins. It's the first time a computer has ever beaten a reigning world chess champion in a game under standard tournament rules. But Kasparov, smart human as he is, adjusts. He wins three of the next five games and takes the match 4-2. So what now? Well, IBM goes back to the lab. They're confident in the possibility now. After all, they were making progress. Deep Thought lost to Kasparov in 1989, going 0 for two. But now Deep Blue had actually scored two points against it. They just needed more compute. So they double Deep Blue's processing speed for the rematch. The improved version can now evaluate 200 million chess positions per second in 1997. Impressive, considering I still can't get my printer to connect to Wi-Fi.
Anyways, it's time for the famous 1997 rematch. But game one surprises everyone with a very odd moment. Towards the end of the game, Deep Blue makes a completely unexpected, counterintuitive move. Kasparov, playing white, had the advantage and was pressing for a win. But instead of trying to make an immediate gain or launch with a tactical counter like a computer normally would, Deep Blue simply moved its rook from D5 to D1. A sort of passive waiting move that seemed to anticipate Kasparov's deeper strategic traps. My understanding, not being a chess expert, is that it was a "do nothing" move, kind of like checking in poker. Kasparov looks at the board and thinks, "That's too human for a computer." Which is funny, because these days we see the exact opposite accusations flying. Magnus Carlsen, today's reigning chess champion, famously accused Hans Niemann of cheating in their 2022 Sinquefield Cup match. He claimed Niemann was completely relaxed in critical positions while outplaying him with computer-like precision. Anyways, so that odd move psyched Kasparov out. And even though Kasparov won that first game, it did rattle him. So in game two, on move 36, Deep Blue, playing white, has a chance to win two pawns and gain a material advantage, which Kasparov is expecting. He knows computers are typically greedy, so he's prepared for Deep Blue to fall into his trap. Instead, Deep Blue executes a surprisingly sophisticated move to shut down Kasparov's counterplay. White pawn on A4 captures Kasparov's black pawn on B5. Kasparov then captures back with his pawn from A6. And then Deep Blue calmly moves its bishop to E4, which according to Grandmaster Yasser Seirawan, sent Kasparov into a tizzy. This is the position that set Gary Kasparov off because the computer made an unbelievably refined positional move. Bishop on C2 E4. The idea of bishop C2 E4 is just to cut out all opportunities of counterplay. But Gary Kasparov, who was shaking his head on stage, understood that he was now lost. >> "What just happened was the computer ignored immediate material gains, which would have been to play something like Queen to B6, and instead it favored a long-term, frankly big-brain positional advantage, completely flipping the script on Kasparov, who used the same kind of tactic to defeat Deep Thought in 1989." As Seirawan later told Wired, it was an incredibly refined move of defending while ahead to cut out any hint of counter moves.
But as we now know, Kasparov loses that series, and he's not happy about it. He asks IBM for printouts of the machine's calculations. At the post-game press conference, with cameras rolling, Kasparov suggests there may be a grandmaster hidden somewhere inside the system feeding Deep Blue moves. >> "Maybe I'm speaking out of turn. Is is do you think there may be some kind of human intervention on on the part of this game or?" >> "No, we don't." >> "Is there a suggestion of the possibility?" >> "Reminds me of a famous goal that Maradona scored against England. He said it was a 'hand of God'." But then 15 years later, we learn what really happened. Murray Campbell, the one from the Deep Blue team, told journalist Nate Silver that the weird move in the first game, moving its rook to D1, was a bug. Deep Blue couldn't pick a move. So, it triggered a fail-safe and picked one at random. It turned out not to be a terrible move, just an unexpected one. Kasparov had concluded that the counterintuitive play must be a sign of superior intelligence. So, Kasparov demands a rematch, but IBM says no. Apple later said IBM felt it had achieved its goal and moved on to other research. Frankly, lack of ambition, if you ask me. And that's the end of the story of how Deep Blue changed chess. But the really interesting part, if you want to understand how AI impacts everything we're doing today, is how Deep Blue actually worked under the hood. So, let's talk about that.
Even though the project started as Deep Thought, it wasn't really thinking, just like modern LLMs don't really think either, but for a different reason that we will talk about. See, Deep Blue was just searching. Chess is like the perfect game for computer scientists. The board has 64 squares in a nice little grid with a controlled set of simple moves available. Both players see everything. There's no hidden information, no fog of war, and no real-time actions. It's one of the easiest games in the world to simulate. Take this toy game board here. Sure, I could move any one of my three pawns forward. Or I could move my knight to one of eight places. 11 total options. But for a computer, 11 isn't that many. It's not like trying to predict which sentence you're about to say next. So midway through your average chess game, you may have around 30 legal moves. Admittedly, it can vary a lot. But to which your opponent may also have, let's just say around 30 counter moves. Multiply them together and that's 900 move combinations or game states after just one move each. Two moves deep, around 810,000 positions. Three moves deep, close to 729 million. All of those moves, which create a big branching mess, is called a search tree. The current board is the root, and every legal move is a branch. Every reply is another branch. And as the computer, you just keep exploring moves until you hit checkmate or a draw, or your computer runs out of time to keep searching. Now, a human grandmaster doesn't actually think about all 700 million possible moves when they're thinking three moves ahead. They look at a handful of the most seemingly promising moves. They follow the interesting lines and use the pattern recognition they've learned over decades of training to narrow it down to less than a hundred real options. But Deep Thought doesn't have mortal constraints. As a machine, it looks at more futures than a human could ever look at in their lifetime in less than a second. Campbell told Scientific American that the 1997 version searched between 100 and 200 million positions per second, usually seeing about 6 to 8 pairs of moves ahead, sometimes much deeper in tactical lines.
Now, you might be thinking, "How is searching enough? That's not decision-making. That's just simulating possibilities." But if you simulate all the possibilities and you can see all the futures where you win, you just keep picking the moves that are most likely to take you to one of those outcomes. But it's still a lot of moves to search. Even if we assume just 10 possible moves on each turn and just 15 turns, so 30 total, 15 for each player, that hypothetical game represents 10 to the 30th possibilities. That's one nonillion possibilities. So, Deep Blue needed a judge to do something like what the best players in the world do when they can't calculate every line to the end. That judge is called an evaluation function, or a heuristic if you're pretentious. Instead of searching every branch all the way to checkmate, it can stop at the edge of its search and score the positions it sees there. If you've ever heard a chess streamer say, "I'm up four pawns here," that's exactly what we're talking about. A good evaluation function can take a chess position and spit out a number that estimates how good the position is. A very simple one would be: given a board state, how many more pieces do I have than my opponent? Positive number is good, negative number is bad. Then it chooses the branches that lead to better-looking positions. Now, in practice, Deep Blue's evaluation function was built directly into the chess chip. It had about 66,000 gates dedicated to scoring positions: material count, king safety, pawn structure, mobility, control of the center. All of that messy chess judgment was compressed into a single score that could be derived from any board state. The opening system drew on a few thousand hand-built book positions in a database of 700,000 grandmaster games. So, to recap, a Deep Blue move worked roughly like this: List the legal moves. List the opponent's legal replies. List your replies to those. Keep branching until the clock or the search depth says to stop. Then score the leaf positions at the edge of the search. Push those scores back up the tree and assume both players are actually trying to win. No matter whose turn it is, the algorithm picks the branches with the best score. If it's your opponent's turn, assume they pick the worst score for you. This algorithm is called Minimax. You maximize, they minimize. And this is the primary algorithm we talked about in my intro to AI class back in 2015. No LLMs to be found. So because branching is expensive, you cut off branches that can't change the decision anymore. Sure, if you had infinite compute, you wouldn't even need a heuristic, but because we live in reality, we need our algorithms to be faster. This kind of shortcut, it's called alpha-beta pruning. Deep Blue's whole stack was built around doing that kind of search in parallel across multiple general-purpose processors and custom chess chips. So that was Deep Blue's intelligence: a giant tree of possible futures, an accurate scoring system built by domain experts who are chess players, and a machine fast enough to execute the calculations. There was no learning in the human sense. No inner voice saying, "I'll make Gary sweat and squeeze him in the endgame." Campbell put it plainly: "When IBM started the project in 1989, machine learning methods for game-playing programs were too primitive to help. So they focused on efficient search and evaluation." Deep Blue was just looking, not learning, and not adapting.
Search and evaluate works beautifully, but only for a very specific kind of problem. You need to be able to describe the whole state. You need a list of all legal moves, full visibility, and you need a decent way to score a position. Chess has all of that. After centuries of play, humans have a pretty good list of what good positions look like. Now, try that with the task to write a birthday card for grandma. What are the legal moves? How do you score how good the wording is? Or try it with code generation or autonomous driving through rain. Minimax just doesn't fit well on these kinds of problems. That's why Deep Blue could beat Kasparov, but could never automate an email marketing campaign. So, Deep Blue was a chess machine. The greatest chess machine ever built at the time, but still just a chess machine. Not the gateway to AGI.
Okay. So, what's the difference about the AI that we can't shut up about today? Well, eventually, we got the machines to start learning for real. But it took a really, really long time to get here. Neural networks, programs loosely inspired by how neurons in the brain connect, have actually been around since the 1950s. Of course, for most of that time, they were an interesting idea with mostly awful results. There wasn't enough data, and certainly not enough compute. But by the mid-2010s, two things had changed. GPUs, originally built for video games, turned out to be great at the parallel math that neural networks need to efficiently scale. And the internet had produced absurd amounts of text. Billions of pages of the written word, Tumblr fanfic, mommy blogs, bad takes on Twitter, and billions of lines of open-source code, all ready to be scraped on the open web. But there was still a bottleneck. The best language models at the time used recurrent neural networks, RNNs. They read one word at a time, left to right, like a person reading a book. So you couldn't process words in parallel, and training was pretty slow. And the longer the sentence got, the more the model forgot about the beginning. Google's massive translation system, you know, Google Translate, turned out to be a great example of this. Around 2016, they shifted from phrase-based statistical translation to a deep network based on Long Short-Term Memory, or LSTM, with attention. The old phrase-based statistical method essentially looked up chunks of words and swapped them based on what was the most mathematically probable fit. But these chunks, often sloppily glued together, fell apart or felt clunky because the machine lacked a holistic view of the sentence. The new system was called Google Neural Machine Translation, or GNMT, and it worked way better than the previous phrase-based translation system. Between 55% and 85% of translation errors dropped overnight. Unlike the previous method, this RNN kept a memory of what it saw earlier, so it could keep better track of languagey things like gender, tense, and grammar across a sentence. But recurrent models were still slow to train. And Google noted even GNMT still translated sentences in isolation rather than considering the whole paragraph or even page.
Then in 2017, eight Google researchers published a very famous paper called "Attention Is All You Need." They called their new deep learning architecture the Transformer because they thought it sounded cool. An early design document even had an image of six Transformers from the franchise, which is pretty sweet. But the paper was very serious. They proposed ditching the sequential reading entirely. Instead of processing words one at a time, the Transformer looks at every word in a sentence at once and figures out which words relate to which other words. That mechanism they called attention. If I say, "The bank was steep and covered in grass," the word "bank" could technically mean a financial institution or a river bank. An attention mechanism lets the model look at "steep" and "grass" at the same time as "bank" and figure out which meaning makes more sense. It doesn't have to wait until it gets all the way to "grass" at the end of the sentence. And because it can process all the words at once, the Transformer can run on GPUs in parallel, so training gets way faster. The original paper was about translation: English to German, English to French. They trained the model on eight Nvidia P100 GPUs. The base model took about 12 hours. The big model took around 3 and a half days, and both beat the state-of-the-art at a fraction of the cost. But remember, these researchers weren't trying to build artificial general intelligence. They were just trying to make Google Translate better. Well, turns out they were on to something much, much bigger.
In June of 2018, OpenAI published a paper called "Improving Language Understanding by Generative Pre-Training." This was GPT-1. GPT as in Generative Pre-trained Transformer, a 117 million parameter model with 12 layers of a Transformer decoder. First, pre-training: feed the model a huge pile of text and train it to predict the next word. No labels, no human annotation. Just given these words, what word is most likely to come next? And second, fine-tuning: take that pre-trained model and adapt it to specific tasks with a much smaller labeled data set. For example, not just "here's some random Python code that might work," but "here's Python code that works and here's Python code that doesn't." GPT-1 beat previous benchmarks on question answering, textual entailment, and common sense reasoning by 5 to 9% on some tasks, which basically means it had a much better understanding of the relationship between different pieces of text. For example, if the premise is "the dog jumped into the lake," it could predict that "the animal got wet," even though there are no directly overlapping words in those sentences, at least not the important ones. And it did this with very minimal changes to the underlying model between tasks. Same basic architecture, but different fine-tuning data.
Then in February of 2019, OpenAI released GPT-2 with 1.5 billion parameters, about 10 times larger than GPT-1, trained on 10 times more data. They scraped 8 million web pages linked from Reddit posts that had at least three upvotes, which was their way of estimating this was reasonably good text. GPT-2 could generate whole paragraphs of plausible writing. Give it a fake headline and it would write a news article with fake quotes and fake statistics. The Guardian called the output "plausible newspaper prose." OpenAI was so worried about misuse that they didn't release the full model at first. If this sounds familiar, you're right. It wasn't the first or last time that an AI release will be wrapped in "too dangerous for the public" messaging. To us in 2026, it's hilarious because even GPT-3, which paved the way for ChatGPT, feels really stupid compared to 5.5, which still can't write an email that I'd be comfortable putting my name on. But GPT-3 was over 116 times bigger than GPT-2. It just goes to show that the AI industry has a long and storied history of exaggerating their systems' capabilities, probably for that sweet, sweet investor capital. Reactions in the tech community were mixed. Anima Anandkumar at Caltech called the decision not to release it "malicious BS" and said there was no evidence that GPT-2 posed the threats that OpenAI described. The Allen Institute for AI announced a tool to detect neural fake news. Researchers worried that Twitter, websites, and email would be overwhelmed by text that sounded reasonable and context-appropriate but was actually bogus machine writing. Unfortunately, I think today we've proved them right. But by November of 2019, OpenAI changed their mind. They said they'd seen no strong evidence of misuse so far, and they released the full model. The Verge correctly pointed out that we've already had programs that can generate plausible text at high volume for little cost to humans.
Then in May of 2020, the GPT-3 paper dropped. "Language Models are Few-Shot Learners." The new version now has 175 billion parameters, trained on 570 gigabytes of text: Common Crawl, WebText, English Wikipedia, and two book corpora. GPT-3 could translate French, answer trivia, write code, and summarize articles, often just by giving it a few examples directly in the prompt. Minimal fiddling required. So, how does an LLM like GPT, Claude, or Gemini actually work behind the curtain? Unlike Deep Blue, it doesn't start with a board. Legal moves are a game state. It starts with tokens. Tokens are chunks of text. Sometimes a whole word, sometimes a piece of a word, sometimes just some fancy Unicode punctuation. Your prompt gets chopped up into tokens, and the model runs them through layers of attention and learned weights. At the end, it produces a probability distribution over the next token. It's not predicting the right answer or the moral answer. Just given everything I've seen so far, here are the chunks of text that seem likely. Then it picks one, usually with a little randomness that you can actually tune if you want. And then it does it again and again and again until it produces a "done" token. That's a language model, a very expensive autocomplete machine. Now, it sounds like I'm being dismissive, but as it turns out, if the autocomplete is good enough, you can get a lot of value out of it. Language is everywhere. Contracts are language. Code is language. Emails, bug reports, product docs, coordinates on a map. A huge amount of work that gets done is squeezed through words before it does anything in the real world. So, if you get very, very good at modeling language, you can do a lot of useful work. And then you give it access to tools. An LLM can write a Python script, but then you can give it a harness that allows it to actually run that Python script, look at the output, whether success or failure, edit the script, and run it again. That's all an agent is: an LLM with the ability to call tools in a loop. You can write an agent in just a couple hundred lines of code. We actually have a course on boot.dev about building your own AI agent from scratch. Anyways, LLMs are unique to AI systems because they handle the messy human language wrapper, but can still hand off the mechanical parts to plain old software. So, while basic GPT-3 might not be the best chess player in the world, you can actually hand it access to Deep Blue or a Deep Blue emulator, and it could then wield that tool to not only beat Magnus Carlsen, but also throw some sweet, sweet shade at him when it wins. So LLMs feel general in the AGI sense in a way Deep Blue never did at all. Deep Blue did one thing. LLMs appear to do anything you can shove into a language problem.
But appearing general-purpose and actually being general-purpose aren't the same. And the people who wrote the early papers actually already knew this. The GPT-3 paper flagged datasets where few-shot learning still struggled and warned about methodological issues from training on web corpora. Even the GPT-1 write-up noted the limits and bias of learning about the world through text and warned that deeper learning models still behave in surprising ways under adversarial or out-of-distribution tests. For example, in 2019, researchers at the Allen Institute for AI discovered that they could break top-tier NLP models by adding seemingly nonsensical phrases. Appending something like "zoning tapping fiends" to the beginning of a movie review forced a model that was 99% sure a review was positive to just flip its position, suddenly labeling it negative. Now, it's a bit philosophical, but to me, it seems that reading and writing about the world in a statistically most ubiquitous way isn't reasoning, at least not in the same way that humans do it. And to be fair, we still don't really understand how humans reason. And is it even fair to say that AGI has to reason in the same way that humans do? I'm not sure. But at the moment, it seems that LLMs can do a lot of things, but they do still seem far from the type of raw computational intelligence dreamed about in 2001's "2001: A Space Odyssey."
Now, I think it's not quite fair to say people overhyped Deep Blue and the 2001 tech bubble, so people must be overhyping LLMs. Deep Blue and the Minimax algorithm are obviously far less general-purpose than LLMs. Deep Blue could literally only play chess. But it's also probably a mistake to assume that just because LLMs are broader than Deep Blue means we must be on a straight exponential growth trajectory to an AGI singularity. >> "Can a robot write a symphony? Can a robot turn a canvas into a beautiful masterpiece? Can you?" The 70-ish years of AI research hasn't just been one exponential progress graph up and to the right, moving faster and faster. We live in a world where we often have something that looks promising and improves quickly, but then the S-curve flattens, and we see diminishing returns. Some have speculated that we're already there with Transformer technology. As a 2022 paper from Epoch AI warned, "public human text data could become a bottleneck for LLM scaling between 2026 and 2032 if current trends continue." So, where are we on the graph, like, right now? It might keep getting better quickly, but that's not a guarantee. We might just be hitting yet another local maxima on the real zoomed-out progress curve. And in that case, it will take another breakthrough that doesn't involve simply adding more data centers to the compute capacity and Reddit posts to the data set. All this to say, I'm not sure. But I'm optimistic about these tools when used appropriately. But I personally don't see a way to get really good and tasteful work done without at least a bit of the human element.
So, thank you for watching. And if you're passionate about being good at things and believe that skilled humans are needed to most effectively use AI tools, then check out boot.dev. We teach back-end development, DevOps, data analysis, and computer science in human-built courses that are heavy on the fundamentals. Our curriculum is even free to read. But if you do want to have the best experience, try our interactive features that make the learning so much more addicting, in a good way, and fun. Use code BootsTube for 25% off an entire year of the premium features on all of our courses. I really hope you enjoyed this video. Please like, subscribe, and share with your friends to make sure that you keep getting recommended these deep dives into your algorithmic YouTube feed. Until next time.