📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Shy Girl AI Scandal: What No One Is Reporting

The Drey Dossier24:36

Transcription

This controversy has changed my life in many ways, and my mental health is at an all-time low, and my name is ruined for something I didn't even personally do.

Those are the words of Mia Ballard, the author of the horror novel Shy Girl, in an email to the New York Times yesterday. You see, Hashet, which is one of the largest publishers in the world, had acquired her book, edited it, and published it in the UK. And then the internet decided that it was written by AI, and Hashet canceled it, pulled it from every retailer, and walked away.

And this is being called a landmark case because this is the first ever commercially published novel from one of the big five publishers ever withdrawn over evidence of AI involvement. We're used to having the conversations around AI self-publishing, you know, flooding the market for years. But Shy Girl just proved that that is not the case anymore. I mean, this has just crossed into traditionally published, edited, vetted literature.

And this news is so much bigger than even that because on the same day that that news broke, the White House sent Congress an AI proposal that would make sure that no institution, no government, no corporation ever would have to answer for what happened to this author. Hey, I'm Dre and welcome to the Dre Dossier. As always, there's an article on my Substack, The Dre Dossier, paired with this video, with all of my sources cited at the bottom of that article. And it's totally free to subscribe. But if you do feel like supporting my work and the lovely people who make this possible, which is honestly just me and Newton, right? Then maybe consider subscribing to the Ruff Riders tier for $5 a month. I don't know. And if that's not for you right now, then honestly liking this video and subscribing to my channel means more than you know.

Which, by the way, did you see I finally got it? If I could get all of your names engraved on it, I would. In January of 2026, a YouTube creator named Frankie Shelf posted a nearly three-hour video with a simple title. I'm pretty sure this book is AI Sloth. And the book in question was Shy Girl by Mia Ballard. And the video at the time of this recording got 1.2 million views.

At the same time, a Reddit thread with someone claiming to be a veteran book editor went viral, walking through the pros line by line and cataloging what they said were the hallmarks of AI generated text. Meanwhile, over on X, a CEO of an AI detection company, decided to run the script through his AI detector, and came back with a result of 78% of it being AI generated text. And within 24 hours, Hashet, one of the biggest publishers in the world, canceled the book, pulled it from every retailer, and issued a statement about their commitment to human creativity and authenticity. And can I just say how much that grinds my gears reading that? Oh my god, it pisses me off so much. Especially when you look at the statement next to all of the evidence I found that supports the contrary.

Okay, so let's just start at the beginning because Mia Ballard originally wrote Shy Girl and published it herself in February of 2025. She didn't have an agent. She didn't have a deal. Just a book on Amazon and whatever audience she could find. And the audience found her. The horror community picked it up on Tik Tok and the Goodreads ratings climbed. Over 4,900 people rated it, averaging 3.52 stars, which is overall a pretty impressive rating. Look, it's not a book for everyone, but some readers really loved it. One reviewer said it was coercive but addictive. And another said that Ballard was undeniably talented. It was dark and visceral, but it was finding its people, you know.

But to be fair, even within those first few months, there were readers who were flagging repetitive phrasing and formatting issues. And a handful of people were saying that it read like it was AI generated. So those voices were there at the beginning. They just weren't the loudest ones yet.

And not long after publishing, the book ran into its first controversy, which wasn't even AI related. The cover of the self-published edition featured a painting of a dog and immediately readers recognized it. Man, that was a really good snap. The image was cropped from a painting that was originally called Dreamer and it was by an artist named Winn Lewis. She creates these emotionally striking paintings of whippids. They're quite gorgeous and she's the daughter of the folk artist Vashi Bunan. And it was a reader on Goodreads who was the one who caught it. They posted screenshots comparing the cover to Louis's painting and called it out publicly.

And according to an account that appeared to be Ballard's own, posted on a legal advice forum, Lewis had then contacted Ballard directly and demanded disclosure of all of the royalties earned from the book during the period which the image was used, plus the removal of any remaining use of her artwork. And look, as someone who makes things and has had plenty of her scripts lifted and used on other people's channels without proper credit, I can deeply empathize with how shitty that feels. Straight up, what Balor did to win Lewis was theft, and that's not okay, whether she did it intentionally or not.

And you know the logic of taking something that someone else made without asking without paying and using it to then build something you profit from is the same story of how these AI tools which are at the center of this controversy were built. The same models that may or may not have helped write this book were trained on on the exact kind of uncredited extraction that happened to Lewis. And you'll find that every single entity in this story, from the largest of AI companies to the institutions pushing their output to the debut of an author pulling an image off of Pinterest is operating from inside a system where taking without asking has become so normalized that the lines of ownership are genuinely blurred for people even though they shouldn't be. And this normalization didn't just happen out of nowhere. It's because there are no meaningful consequences for anyone at any level. If robbing a bank isn't illegal, then why wouldn't you rob a bank?

Now, it's important to note that when Hashet acquired the book in mid 2025, they already knew about this theft of Louiswis's artwork. Allegedly, Hashet considered reaching out to Lewis directly to commission a new piece of artwork from her, but ultimately decided to commission a cheaper and entirely new cover instead. Well, almost an entirely new cover and still looked pretty similar. They hired a new designer named Stephanie Ahes and she created new covers that kept the same whippit energy and feeling as Louiswis's original painting just without involving Lewis at all allegedly. And the US cover was revealed in September of 2025.

Now whether any financial resolution was ever reached between Lewis and Ballard has never been publicly confirmed. But to this day allegedly Lewis is still tracking down her painting being used when other content creators are talking about the book across social media. But it's interesting, right? >> Just interesting. The publisher that publicly positions itself as a defender of artists, knowing that an artist's work had been stolen and used to build the commercial momentum of a book they were about to profit from, responded by working around her.

And believe you me, we will get more into Hashett's broader record on creative rights in a moment. But first, I want to make sure that we're all on the same page about what it actually means to work with a publisher, especially as a new author. Now, I am fortunate enough to have a little bit of insight on this as I've been working with my agents on a book proposal of my own. And it's through this process I've been learning a lot about what representation actually does for you, like what it protects you from and how much of the institutional relationship building that keeps you safe happens through that layer. And it's a layer that most people never see. More than anything, it made me acutely aware of how exposed you are and vulnerable you are without it. And Mia Ballard did not have that layer.

When a major publisher acquires a book, meaning that they buy it, they take on the editorial responsibility of that manuscript, they read it, assign it an editor, handle the contracts, legal review, cover design, marketing, and publicity. The author, especially a green, unrepresented one, is trusting in that institution to bring all of that to bear. And that is the entire value proposition of traditional publishing. That is what the advance is for. That is what the lower royalty rate is for. the publisher takes the larger share of revenue because they are the ones who are taking on the risk and they're also supposed to be the ones who are doing the work of a professional institution. So, Hashett would have done all of that. They would have read the manuscript, assigned an editor, designed a new cover, scheduled the UK release for November of 2025 and the US for May of 2026. They sent advanced copies to reviewers and even got blurbs from authors including Olivier Blake and built entire marketing campaigns around it.

And I'm just having a hard time believing that throughout all of that, you're telling me that a company that publicly positions itself as a fierce defender of authentic, creative human work never thought to check whether the manuscript had been written using AI. Like this is the first they're hearing about it when the New York Times reached out to them. Because what had happened is the New York Times reached out to Hashett with evidence that the book may have been substantially AI generated. And then according to Hashett, they conducted a lengthy internal review before making the decision to then pull the book from the shelves. If you have the ability to do an internal review process, right, a lengthy one at that, then why on earth did that not happen before publishing the book in the UK where it sold 1,800 print copies? Why would that not have been done as part of a discovery process, before signing a contract, before revealing the US cover or sending over advanced copies to reviewers?

And within 24 hours of that call, Shy Girl was removed from Amazon, pulled from Hashet's own website. The UK edition, already on shelves, was discontinued. The planned US release was obviously cancelled. And Hashet issued a statement saying that the company remains committed to protecting original creative expression in storytelling. And the book's author, Mia Ballard, sent an email to the New York Times late on Thursday night. She said that she did not use AI. She said that an acquaintance she had hired to help edit the self-published version had used it without her knowledge. She said she was pursuing legal action and couldn't say more.

Now, before we go any further, I want to just say something plainly. I am not going to be spending the duration of this video prosecuting Mia Ballard. Partially because I genuinely don't know what happened, and partially because women of color already face enough barriers in this industry just to get their work in front of readers, let alone to survive a public accusation of this magnitude with their career intact. Publishing has a documented history of giving debut authors of color smaller advances, less marketing support, and less of that institutional relationship building that protects you when something goes wrong. Like, let's be honest, authors that survive moments like this are the ones who the industry has already invested in protecting. And Ballard had none of that infrastructure when this hit her. And it's that it's that right there that tells you something about who absorbs the cost when institutions fail and the internet decides it already knows the answer.

But I'm not quite done with hashtit yet. So, let's circle back to them. Hashet pulled this book and issued a statement of their principal stance on authenticity and original creative expression blah blah blah. When Hashet canceled Shy Girl, that means a contract was then terminated, which means the income Mia Ballard might have been counting on just disappeared overnight. And for all we know, this income could have represented the first return on years of work that had finally broken through to something. And when you look at it from within that framing, you realize that this is fundamentally more than anything else a labor conversation. This is a conversation about the standard of evidence that should be required before any institution can take away someone's livelihood. Because that standard is what determines whether any working person in this industry actually has real protection or whether they're just one Reddit thread away from losing everything with nowhere to turn.

So what evidence did Hashet actually rely on? Well, the New York Times presented them with findings from two sources. The first is Max Sparrow, the CEO and founder of Pangram, which is an AI detection tool. After seeing a bit of the AI discourse online, Sparrow decided to run the full text of Shy Girl through his tool. The result said that 78% of the text was AI generated. He posted his research to X and told the Times that he was very confident that the book was largely AI generated or at least heavily AI assisted. And that number, the 78% I find has been repeated in every piece of coverage on this topic since. So I went and found the post so that I could take a look at this research for myself. Luckily, Spiro had included a public link to the panagram report. So, I clicked on it and started scrolling through all of the flagged sections. And, you know, I kept seeing something just repeating over and over and over again in the portions that were flagged as AI generated. It was a little URL that was embedded into the text itself. And it said oceanofpdf.com. Are you kidding me with this right now? Oh my.

So, if you're not familiar with oceanof pdf.com, it's basically a piracy website that illegally distributes copyrighted books as downloadable PDFs, which means that the version of Shy Girl that Spiro ran through his tool is not the original manuscript. Like, it was not the Hashett edition. I mean, this being the one that produced the 78% figure reported by the New York Times and allegedly used as part of the basis for cancelling Ballard's contract. This mother trucker ran a pirated script through his tool from a filesharing site that he can't even confirm is the actual manuscript. And the Times didn't mention this. Nobody has reported it.

Now, look, I want to be a little careful here because I'm not saying that this definitively invalidates all of those results, right? What I am saying though is that nobody has asked this question because pirated PDFs go through something called OCR scanning and format conversion and text processing before they can exist as a readable document. And we do not know what the processing did to that text before it was fed into the algorithm or whatever conversion artifacts potentially affected the output. And yet the person who generated that number just posted it publicly and watched it become the evidentary foundation of one of the most consequential decisions in this story. And it all seems to be reliant on a stolen copy of the book.

And the specific phrases that Pengram flagged as these hallmarks of chatbot writing raise real questions about what this process is actually measuring. The flagged examples being like, "The pause feels like a knife in my chest, sharp and unyielding, and I press the phone to my lips, the screen cool and unyielding." I mean, the tool flagged the repetition of the word unyielding across two similes and called it the hallmark of AI writing. Like, what is this actually built on? This is if you ask me.

The second source The Times cited is someone by the name of Thaad Mckelroy, a publishing industry consultant. Mckelroy learned about the allegations from a Pangram employee, got his own copy of the book, which I hope he paid for it, and he ran it through GPT0ero and originality. And the Times reports that all of the tools found the text likely to be largely or partly AI generated. The 1 in 10,000 false positive rate cited in coverage comes directly from Spiro. You know, the CEO telling a reporter his own tools accuracy rate unverified by any independent third party.

Here's the thing with these tools, right? Even if we take this at face value, that one in 10,000 doesn't sound that bad. It's reassuringly small until you remember that more than 3.5 million books were self-published last year. Which means at that rate, roughly 350 authors would be wrongly flagged every year from self-publishing alone. In my opinion, using AI to diagnose AI is like putting a hat on a hat. it's inherently going to be problematic because AI has significant error margins on both sides of it. And I talk about this a lot in my videos about automating warfare and the error rates of AI targeting systems. And I don't know about you, but I don't feel comfortable accepting even a small false positive rate right now across a large population because that would mean that real people get wrongly accused. And when it comes to AI detection tools, this false positive number might even be less reliable than advertised.

Why do I say that? Well, that's because I tried these tools out. That's right. I tried each one of these tools out. And let me tell you what, I got genuinely pissed off by a lot of them. And I wasn't really that impressed with what I saw from Pengram on X.com. So, I decided to try out originality.ai to see what its parameters looked like. And let me tell you what, originality.ai does not give you a single clean verdict at all. What it does is it breaks your document down into all these little segments and it scores each one individually, flags all these specific sentences here and there at varying confidence, and then it pieces together this overall percentage, and then that gets reported out as the headline. I'll show you what I mean because I uploaded two things. One of which was a poem that I wrote when I was 19. So that would have been 2016. Jesus Christ. And the other was a university paper I wrote back in like 2020 or something. The point that's important here is that that writing is before the large language models on a commercial level existed. And overall the scores came back as mostly human, which is already insulting. One of them came back as 91% human. What do you mean 91%? What's the other 9%.

But the thing that gets me is like, okay, when you look at all of these individual flags, like for instance, from the poem I wrote as a teenager, you know, words that came from my actual soul, my 19-year-old soul, flagged at 70% likelihood of being AI generated. If I were 19 today and my poems were being falsely accused as being AI generated, that would be devastating. And it shows you that this 1 in 10,000 statistic doesn't really capture the experience at all. Because if someone wanted to use any one of those individual lines to flag against me, they could.

There's also this really interesting body of research around AI language patterns that the coverage of this story is almost completely ignored. Researchers at the Maxplanks Institute for Human Development analyzed 280,000 video transcripts from over 20,000 academic YouTube channels and found that human beings are unconsciously absorbing this AI linguistic pattern in their own speech. They found that academic YouTubers began using words that AI chatbots favor up to 51% more after chat GBT launched. And this was showing up in spontaneous speech as well. And as the study puts it, the mechanism is straightforward. Humans write online, AI trains on that writing and develops this recognizable cadence. And that cadence saturates the internet and then humans absorb it and then AI trains on the blended corpus all over again. So these styles are converging in both directions simultaneously, which means detection tools calibrated against the earlier moment in that loop grow less reliable every month because the baseline of what human writing sounds like itself is moving towards what the AI output sounds like. And so when a Reddit thread decides that a book is AI, what they might actually be measuring is this feedback loop that we're all living inside of.

Now, let's go back to the company that made the decision, Hashette. That's right. You thought you were off the hook? Nope, not yet. Not done with you yet. Because you see, I was doing a little digging and I saw that Hashette had a published AI policy since November of 2023, updated as recently as March 2025. And part of that policy says that authors must confirm that their work is original and disclose any AI use. And the company says that it would never license an author's work for AI training without explicit consent. But the policy makes a clear distinction and it draws a line between creative AI use, which they prohibit, and operational AI use, which they explicitly permit. And operational use based off of what their website said is anything from metadata creation, keyword optimization, error checking, marketing copy, and using data analysis and behavioral signals to monitor and inform acquisition decisions. In fact, their director of consumer insights is speaking at NYU publishing conference this year about using cultural and behavioral signals and reader sentiment to shape what books they acquire. And industry reporting confirms that at large these major publications broadly are using AI tools in content discovery and author scouting which okay for me raises a pretty specific question. Not an accusation just a question. Is it possible that the social signals that made Shy Girl visible to Hatchet like book talk traction and Goodreads momentum and the viral engagement. Is it possible that an AI tool flagged this book for acquisition as part of Hashett's permanent operational use?

And if that's the case, then you have an AI tool that potentially finds a book, a pirated PDF that gets run through another AI tool by the CEO of a detection company producing a number that's reported as research, a publishing consultant, which runs the book through more AI tools and get similar results. And all of those go to Hashet and then Hashet allegedly uses them to cancel the contract. And so the only person in this entire chain who faced consequences for any AI use is Mia Ballard, the author, the artist. And look, I don't know that to be true or not. I don't know what tools Hashett does or doesn't use, but I think it bears asking out loud for the sake of artists everywhere.

Okay, now let's get to the governance part of this, okay? Because this is really important and what matters to me and you. So, on the same Friday that this shy girl story broke, the White House released this. It's the National Framework Policy for Artificial Intelligence Legislative Recommendations, March 2026. the Trump administration's formal recommendations to Congress for how the United States should govern artificial intelligence. And this document deserves a video onto itself, which I'll probably make at some point. But right now, I just want to show you the specific parts that answer every question we just spent this entire video asking.

So, right now, maybe you're thinking to yourself, well, someone should be held accountable for all of this, right? Someone should have standards in place and someone should be able to protect the next person that this happens to, the next artist that this happens to, right? Well, this document is the answer. And the answer built deliberately into the architecture is nobody. There's there's nobody that's going to do that. So this document is not an executive order. Okay? It's a set of legislative recommendations. The White House is basically telling Congress what it wants it to pass. It still requires a congressional majority, which means that it requires the public to understand what is being asked before it gets voted on so that we can push back and we can let our congressmen and women know what we want. Like for them to quit taking Apac money, right? >> It's interesting. you're like the first to bring up APEC in yours >> and there's a couple of pillars here with a bunch of things hidden in all of them but but on copyright the document states that training AI on copyrighted material does not violate copyright law and then says that the courts should resolve the question meaning that they are already taking the industry side and calling this a judicial neutrality. Let the courts fight it out. Congress don't make any rules for them. You know data and labor violations are not going to be resolved until the courts resolved it.

And on oversight, and this one is really important, so please listen. The document recommends that Congress should not create any new federal rulemaking body to regulate AI ever. Yeah. So, you know how aviation has the FAA and pharmaceuticals have the FDA and nuclear energy has the NRC? It's because each of those industries when they became powerful enough to affect large numbers of people, they got dedicated a federal body with the authority to oversee it in the public interest. And this document asks Congress to establish a statute that AI never gets that not just temporarily while Trump's in office, meaning that permanently they pass a law that AI never gets a regulatory body.

And then there's the preeemption section, which would close every other remaining door that's even slightly open. It says that they don't want states to be able to regulate AI development at all. And on top of that, developers cannot be held liable for third party misuses of their model. So the federal government cannot act because there's no regulatory body. And the states cannot act because they are preempted. Which means that California, you know, the state where Mia Ballard lives, can't pass a law requiring evidentary standards before a publisher can cancel a contract based on an AI detection tool, no matter if it's defunct or not. It also can't require that tools be independently audited. It can't create the right of appeal for someone whose career was ended by a freaking algorithm. This framework does not protect Mia Ballard. It does not protect you. It does not protect me. It protects the institutions and the developers and it preempts the states that might try to fill in those gaps.

And this important document dropped yesterday, the same day as her story, while everybody was talking about whether or not an artist should be punished for alleged AI use. Whatever the truth is about how Shy Girl was written, and I hope that I've shown you today why that is genuinely a harder question to answer than the internet makes it seem. Mia Ballard's name is permanently attached to this story. And look, I'm not telling you that you have to feel sorry for someone. Like, you're allowed to feel whatever you feel. And it's okay, I think, to have conversations about AI's use in creative work and have these healthy debates and discussions about what's okay in a book and what's not. And it's okay to speculate, and it's okay to discuss and have opinions. I mean, this is all a part of being a reader and being part of a community that cares about art. The problem is that we keep ending those conversations at the artist. And the artist almost every single time is the person with the least amount of power and least protection in the entire chain.

Mia Ballard said she's going to pursue legal action. And maybe she can. Maybe she has the resources and the energy to fight on top of everything else that she's already carrying. And I genuinely hope so. Truly, I do. But I think about what that means for a debut artist whose mental health is at its lowest and whose name is attached to a story she says she did not cause. And I wonder how many people in her position would have any real path forward at all.

Hashet can sue Google and OpenAI and mean every word of it and still owe Mia Ballard and Win Lewis a considerably more honest accounting of what happened here rather than a press release about their commitment to human creativity. And it's important that we save as much of this haterade and direct it at our government right now, our current administration as possible. We've got to reframe this fight in our head from I hate AI to I hate that our government released framework ensuring that no government's architecture will ever be built to protect artists from exactly this kind of harm. Because right now we are doing the lobbying work from people who benefit from this chaos for free. Every conversation that ends at the artist and never reaches the institution. Every pylon that burns out before it gets to the policy question. Every outrage cycle that moves on before anyone asks who's held accountable. That is the system working exactly as designed. The regulatory window on AI governance is closing. And once it shuts, the only question that will be left is whether you will be inside the house or outside looking.