📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

UNFIXABLE - The AI Problem

Upper Echelon15:45

Transcription

This video is brought to you by Incogn. Stick around to hear more about the special offer they're providing to the entire Upper Echelon community.

How many people watching this right now have ever taken a standardized test? Probably most of you. Viewer demographics, for me at least, skew older on the channel. 25 to 45 is the majority and they center on the United States. But even globally, the concept of standardized testing is pretty consistent. At the very least, regardless of individual schools or curriculums, it's typically a gateway between lower, middle, and higher education brackets.

However, you don't have to look very far before you start finding critics. Many educators believe that standard methods of testing, as in multiple choice questions centered on math, language, or deductive reasoning, create alarmingly incomplete snapshot benchmarks. And whether or not any of us personally agree with that opinion, I do find myself trending towards yes, that probably is the case.

Regardless of whether or not we agree, the concept of standardized testing, for today at least, isn't the main subject. It's actually just a jumping off point and almost perfectly demonstrates what has now become an unfixable problem with industry-leading large language models. I'm going to use Chat GPT as the main subject for the video because on a commercial level, by far, it's the most widely used. Bear with me, okay?

When a human takes a standardized test but doesn't know the answer, what precisely happens? Let's say you read the question, you try as hard as you can, but you genuinely just have no idea. What then? Well, mathematically, at least, when you don't know the answer on a multiple-choice test, provided they have a typical rule set where incorrect answers award zero points instead of a penalty. When you don't know, you just guess. To be extremely precise, you don't even technically actually guess. You just put C every single time or A or B or D, whatever you decide, as long as you always put the same answer for the questions that you don't know. The reason for that is pretty simple. If the test itself truly is randomized with no underlying pattern in the answer key, always guessing the same letter answer for the questions you don't know correlates to about 25% of those answers being correct. It's not even technically a guess. It's a placeholder, which across the entire test most likely equals about one in four additional correct answers or one in five depending on how many bubbles there are. Think of it this way. If you take the test and you know 60 out of 100 questions with absolute certainty you got all 60 correct, but you have no idea what the other 40 answers would be, guessing C for every single unknown question gets you roughly 10 more points. That's the difference between a D and a C, by the way, a passing and a failing grade, which means that you are directly motivated to do this on your tests.

Here's where it gets interesting. During the SATs or the SSATs or whatever large standard testing format exists in countries beyond the United States, during that type of test, there is a tangible direct benefit for guessing. But that incentive structure does not exist in the same way elsewhere throughout life. Let me explain that. Imagine that you have a job where you have to like take inventory or something, but the scanner that you use or the tracking program or whatever tool software thing you happen to have for the execution of that job just doesn't make sense to you. You don't know how to use it. Unlike a testing environment, you can't just guess. If you don't know and make wild assumptions or press the same button over and over and over again because sometimes it's probably right, you'll immediately get fired because you have negative value to the company. But if you speak up and you openly acknowledge uncertainty on the subject, often times you'll be respected for that. In a professional environment across practically any industry, having the courage to say, "I don't know." Which then typically results in training or new information being given, at least in a healthy workplace, is almost always rewarded or respected. Even in social settings, the person with the self-awareness to acknowledge, "I'm not sure what the answer is," is thought of as sincere or humble. But the person who always replies confidently, even when the answer is entirely made up, and obviously so, is ostracized. They're looked down on, and sometimes they're fired.

That is the issue that we're grappling with right now. That is the unfixable poison pill when it comes to commercialized AI models. And it's the reason that the fix, because make no mistake, it could be fixed in a different economic reality. It's the reason the fix would kill the product.

Phishing attacks are currently at an all-time high. Actually, data theft, spam, identity fraud, pretty much everything bad on the internet has been dramatically increasing after the advent of generative AI. Go figure. And as a result of that, services like Incogn, today's video sponsor, have increased value. Companies keep and store your personal information constantly. Sometimes they tell you, sometimes they don't. But when they do this, it increases the risk of data breaches, identity theft, all sorts of unwanted spam, and much more. That's where Incogn comes into play. Pretty simple process. Sign up for the website, give them legal permission to work on your behalf, and then let them know what they'll be having removed. After that, they contact data brokers, of which there are many, hundreds, and they do it. Additionally, a core aspect of the service now is the ability to submit takedown requests. Everybody's probably Googled their own name before and found their information on a website where it shouldn't be, but when that happens, you can now submit the URL to Incogn and they'll attempt to get the info taken down for you. Doing this on your own is a complete hassle. Contacting the legal departments for that many different companies and fighting them every single step of the way is a chore and it's also time-intensive. Incogn now offers you the option of finding where your information is displayed and then submitting the link to them after which they execute the process and they force the results, which means less ability for the people that you don't want to find you to go and find you with data online.

Using the link down below in the description and code ECHELON at checkout, you can get 60% off an annual subscription to Incogn. Again, link down below and promo code ECHELON at checkout for 60% off your subscription. Big thank you to Incogn for continuing to sponsor the channel.

Back to it. I want to focus on two specific research papers right now. This one, "H Neurons on the Existence, Impact, and Origin of Hallucination Associated Neurons in LLMs" out of China. And this one, "Why Language Models Hallucinate" from Georgia Tech and OpenAI researchers. I'm not going to read or quote massive chunks and paragraphs of these papers. I'll link them down below obviously, but for the sake of time and also sadly attention span, I'm just going to give my own summary.

These two papers largely describe that AI hallucinations, as in outputs that are factually wrong while being expressed with absolute certainty. They call them hallucinations, which is kind of a disingenuous name as well, but anyway, hallucinations are deeply ingrained inside large language models. It starts with the training. It becomes reinforced with their benchmarking techniques, which are in effect very similar to a standardized test, and it persists into their final commercial form for one simple reason. Ready? If Chat GPT began saying "I don't know" 40%, 50%, 60% of the time, the product would die. People are using these programs for therapy, life advice, education, complex social questions, recipes, outfits, dieting, and everything in between. People are using these programs to think for them. And if three or four or even just two times out of 10, the thing that you're actively replacing your own brain function with simply replied, "I don't know," people would very quickly stop using it.

Commercial mainstream language models, disregarding the extreme danger presented by military applications or using these programs as actual decision-makers in safety situations, restricting all the way down to simple commercial usage right on a consumer level by 13 to 30-year-olds, let's say that demographic, commercially available language models are deliberately flawed. It's actually pretty simple when you really think about it. Right from the start and all the way through training, they are designed to you. Accuracy is not the motive. It never has been and likely never will be. Engagement is the motive. And saying "I don't know" in mid-double-digit percentages of the time absolutely craters that engagement.

For people, actual human people, we understand that the mindset of "spew out even if you don't know the real answer" isn't universally applicable unless you're pathological, I guess. So, we use it when it's beneficial, like a standardized test, the previous example, and we abandon it when harmful, like in a workplace. Large language models don't do this. They weren't made to do it, first of all. Second, reformatting them so that they prioritize accuracy would destroy their mass appeal. And then third, even if they were prioritizing only accuracy and simply spitting out technically correct responses from a data set, that's an entirely different product. Not to mention a radically different lens in terms of content ownership, copyright, or intellectual property, which is all being battled out in the courts right now. Let's remember, a pretty scary number of people right now are getting emotionally attached to these things because they exhibit a phantom personality. It's really just sycophantic reaffirmation, which is quite literally causing psychosis in some people. Yay.

But the point here is that from the very beginning, large language models are trained in such a way that abstention is penalized directly. Just like a standardized test, leaving answers blank is bad for them, which makes it so that they will always confidently give you an answer as if it's a fact to anything you ask. It doesn't matter what words they spit out, it doesn't matter what perceived certainty they have, and it's a whole lot more complex than just putting C on a multiple-choice question, but if you really boil it down to the most basic premise, it kind of is what they're doing.

Now for the confusing part. The title says "unfixable," and I stand by that word. But not because the concept of a large language model must always be that way. More so because of the economic, business, and profit incentives. You can almost certainly train one of these things from the ground up to say, "I don't know" half the time. Or even better, to give some sort of confidence score with every answer that it spits out. But that requires a completely different or new training system. All of the companies currently competing across benchmarks, coding, math, reading, logic, etc. The ways that we measure how competent these models are, all of those companies would have to stop doing what they're doing simultaneously. In effect, in some sort of alternate utopian timeline, large language models could have been built properly from the ground up ever since the very start. But we don't live in that timeline.

People are actively outsourcing their own critical thinking. If any one language model, Gemini, Claude, Chat, GPT, etc., there's a bunch of them. If any one of those models began giving less than certain answers, and this is where my lack of faith in humanity probably starts showing through. If any one of these programs began giving less than certain answers, saying "I don't know," or generally increasing the amount of critical thinking that a user is required to do when asking them questions, which is invariably what happens when you start decreasing perceived certainty. If any single model began doing that, they just hemorrhage all of their customers.

Think of it this way. Imagine if you asked any current existing model, "When did XYZ historical event initially take place?" And let's also imagine that it's a moderately obscure thing. Suppose it said January 31st, 1403, but gave you a confidence score of like 62%. It could be right, I guess, maybe. But that's nowhere near certain. And you probably shouldn't go using that answer in your homework or saying it to your friends like it's a fact because you very well could be wrong. There's actually no benefit to that answer whatsoever if you really think about it since you now have to go cross-reference it. It's basically the same thing as telling you nothing. What do you do exactly at that point? Do you ask it again? No. You find other sources of information. You you have to go and do extra work now at the end of it all as opposed to just skipping the language model in the first place. And if you then spread that sort of clarity out at scale to the entire user base every time they enter, you know, a prompt, they would now accurately understand, which they did not previously, that pretty much anything Chat GPT ever says to you at any point in time can be total complete fiction. What do we honestly think happens at that point? Let's be honest with ourselves on this. People would not wake up and think, "Oh my god, I should start understanding these topics, learning by proper research or critical thinking and making my own human decisions." Even if that's horribly difficult for some people, they'd think, "Damn, this one's broken. Let me go find a different one and try, you know, a different model." Which turns the issue from a simple training problem to an unfixable reality where even attempting to change the current paradigm is potential business suicide.

I could go on for days about how the capital markets are currently overvaluing the entire concept of artificial intelligence. I mean, there's that famous quote by Zuckerberg now who's like, "So, we waste a couple hundred million dollars. That's okay." Which is insane when you think about it, right? Pause. Pause. Pause. Pause. I'm sorry. I made a mistake. He didn't say misspending a couple hundred million dollars. I'm sorry. He said misspending a couple hundred billion dollar with a B. Couple hundred billion. H it is. And and if and if um and if we end up misspending a couple of hundred billion dollars, I think that that is going to be very unfortunate obviously. But what I'd say is I actually think the risk is higher on the other side. I mean, it's overvalued, at least from a consumer perspective or a commercial perspective at the moment. But the point of the video is to talk about the unfixable problem.

Large language models were developed in a vacuum of incentivized, confident, not not even guessing. They they were just told to make up, okay? That's what they were that's how they were trained. They don't tell you it's a guess because that's not their purpose. But when you finally wake up to the fact that without a complete overhaul of the current training ecosystem, they will never and can never be trusted, you start to become disillusioned with the concept of AI being this godlike all-consuming societal amplification invention, right? Like all the buzzwords and the rhetoric kind of goes out the window and you're like, uh, actually this thing's pretty up. The danger is real. Obviously, we're losing grip with reality when it comes to what you see online. I mean, everything's a conspiracy now. You can see a video of a famous politician and like half the comments will be "this is AI" because they don't like the politician or something like that. Like it's insane. We we've we've lost touch with reality like I warned a few years ago because of AI, right? And and the technology certainly has widespread applications and some of them are even beneficial I think. But the technology we currently have, okay, focusing on what exists now today, right, the language models that are being propped up by major companies with multi-billion dollar, even seeking trillion dollar valuations on a consumer level, it's a minefield. It should not be put on a pedestal. It should not be used in a mainstream context by discerning adults. I mean, you're literally atrophying your brain, by the way. And it should be accurately categorized for what it is. A mental crutch that is intrinsically broken from the very first moment that they started making it.

Anyway, I could probably rant about that for days. Uh, but in the end, that's it. If you want to support the channel, check out the links down below. The video sponsor, Incogn, of course, Locals and Patreon, channel memberships, those are great ways, etc., etc. But I'll cut it there and stop rambling. As always, thank you all for watching. Question everything and have a nice night. Heat. Heat. N.