📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Gene Hunting with o1-pro: Reasoning about Rare Diseases with ChatGPT Pro Grantee Dr. Brownstein

Cognitive Revolution "How AI Changes Everything"1:33:01

Transcription

Rare diseases are quite common, actually. There's more people with rare diseases in the United States than there are natural blondes. We need to sequence the whole world in order to understand what is actually disease-causing and what is just background variation.

Logging into the Harvard Library, getting that paper, skimming the abstracts—it's not at all what I want—then going back and being able to ask for a summary, be like, "Oh yeah, this sounds good," has changed my life. It's all cutting down on this mundane, time-consuming, really tedious part of the job and getting back to the fun part, which is gene discovery. There's going to be this whole generation of geneticists that aren't going to know how things were done before all this was available, because it's going to be a huge game changer and timesaver.

Hello and welcome back to the Cognitive Revolution. Today I'm speaking with Dr. Katherine Brownstein, MPH, PhD, and assistant professor at Boston Children's Hospital and Harvard Medical School, whose research focuses on identifying the genetic causes of previously unexplained rare and orphan diseases, and who was recently awarded a ChatGPT Pro Grant from OpenAI.

You might be surprised to learn, as I was, that so-called rare diseases are not necessarily all that rare. Any disease affecting fewer than one in 2,000 people, or fewer than 200,000 people in the United States, is classified as a rare disease. And often families spend painfully frustrating years bouncing around the medical system in search of an accurate diagnosis before ultimately reaching Dr. Brownstein's elite team at Boston Children's.

Of course, considering the radical cost reduction that we've seen in genetic sequencing in recent years—which, with nearly a 10,000x improvement in affordability, is one of the very few cost curves ever to rival that of large language models—there's been an ongoing revolution in this space, even before the current AI moment. In 2007, a genome sequence cost upwards of $1 million. In that era, it was used only in the most challenging cases and was often a difference maker. Today, it's just a couple hundred dollars and has become commonplace for individual patients. But that creates new challenges for specialists like Katherine, who now have to comb through a vast and still exponentially growing literature to find candidate diagnoses for their most challenging cases. This new wealth of information, which, as you'll hear, could be growing even faster still with improved regulations and incentives, makes information processing capacity relatively scarce and valuable. And you can probably guess where this is going—a great target for the latest generation of reasoning models.

This conversation, then, is above all a window into how frontier large language models are starting to become useful in highly specialized fields. Dr. Brownstein is pioneering the application of AI to rare disease research in real time. She's using AI to triage potentially relevant research and, in some cases, to connect the dots between subtle clues. She's working directly with OpenAI to develop use cases and provide feedback. And considering that every case represents a real person with a life-altering or even life-threatening condition, she's constantly working to find the right balance between enthusiasm for AI's capabilities and a healthy skepticism for any specific AI output. As you'll hear, she's still figuring out where AIs can be the most valuable, how best to use them, and how much to trust them. That such an established expert is bringing what amounts to a beginner's mindset to such high-stakes cases may be surprising to some, but really, I don't think it should be. Even the most AI-obsessed folks like me have only managed to log a few thousand hours with large language models, and nearly all of that was with earlier and less powerful models. So for the current frontier, we're all still figuring this out together. And there's currently an unprecedented opportunity for people with deep experience in specific niche domains to become the leaders in applying AI to their particular fields.

Importantly, if you are listening to this podcast, this is probably something that you can personally do, even if you're still relatively new to the AI tools themselves. If you're inspired to take on that challenge and think I might be able to help, or if you have any feedback or guest or topic suggestions, please don't hesitate to send me a note. One of the very best parts of doing this podcast is hearing from you, the listeners, and I am pretty consistent about reading and responding to every message I get. Of course, we always love it when listeners share the show with friends, online or offline, and we very much appreciate the many reviews we've received on Apple Podcasts and Spotify, as well as the increasing active comment section on YouTube.

For now, I hope you enjoy this conversation, which I hope will become the first in a series of episodes with ChatGPT Pro grant winners on the application of frontier AIs to high-stakes medical research with Dr. Katherine Brownstein.

Catherine Brownstein, MPH, PhD, and assistant professor at Boston Children's Hospital and Harvard Medical School, you are specializing in the discovery of new genes for rare and orphan diseases, and you've recently been awarded a ChatGPT Pro Grant. Welcome to the Cognitive Revolution.

Thank you so much for having me.

Yeah, I think this is going to be really exciting. As regular listeners know, I have a growing obsession with the intersection of AI and biology, and so when I saw your name on the ChatGPT Pro blog post announcement, I was excited to reach out and learn more about how you are applying the latest AI tools to some of these very challenging and pressing problems. Maybe we could start with just like a zoomed-out, kind of step-back overview of your work. Our listeners are definitely following AI developments; they know about OpenAI, they know about ChatGPT Pro—probably quite a few have subscribed, even at the $200 a month level—but they probably don't know a lot about rare diseases and, you know, what our state of knowledge is, what sort of techniques people use to try to figure these things out. So I'd love to just get—you know, this may be a very tough question, maybe the toughest question—but what's kind of the layman's introduction to your advanced work?

So when I'm asked that question, I usually answer it by saying that rare diseases are quite common, actually. There's more people with rare diseases in the United States than there are natural blondes. So it's actually quite common to have a rare disease, and a lot of it's like how you define disease. Like autism is really common, but autism due to a de novo variant in KCNJ8 is quite rare. So, you know, it's kind of a tricky definition, and it's always kind of evolving as we learn more. But basically, I consider myself a gene hunter, and trying to diagnose the undiagnosed.

I read in preparing for this—I think it was Perplexity—gave me this answer that the definition of a rare disease is one that affects fewer than 200,000 people in the United States. That was a surprisingly large number, isn't it?

It's always wild to me because when you think of like a city that's 200,000, that doesn't seem like a small town, or at least it doesn't to me. But, you know, that's the definition of rare in comparison to common disease, which can affect like millions or, you know, like epilepsy is 1% of the population. So yeah, that's really interesting.

So maybe a little bit more color, kind of background, um, on maybe the patients' journey through the medical system to get to you, and then sort of your experience of encountering new patients. Like I know, you know, it's of course going to be impossible to give like one story because I'm sure they're extremely varied, but how do you end up coming into contact with patients? What have they gone through to get to you, and and what do you do, you know, once you get a case?

I'm really lucky to be at Boston Children's, which is internationally known tertiary hospital, so we get really interesting cases from all over the globe. Usually a patient family starts out where they go to their local medical provider, and they can't figure out what's wrong with the child or person, and then they get referred from specialist to specialist, and they still can't figure out what's wrong. And then eventually they get to us, where a lot of times patients and families have been bounced around for years trying to figure out what's going on, and what's next, and what can they expect, and just looking for answers. So sometimes it's not that case, like we have a lot of really savvy, medically savvy families where they know their child and they know something's wrong and they need the best right away, and they're going to search on the web and find the person who works on that phenotype and call every day until they get an appointment. But a lot of times it's a more circuitous route, and going from doctor to doctor to doctor, and then finally somehow ending up at Boston Children's. And then if they see a clinician who doesn't know, they often refer the case to um the organization that I work with, the Manton Center for Orphan Disease Research, and we get a lot of the negative cases throughout the hospital where they think it's genetic in origin, and then we're able to get the medical records. We're a philanthropically virtual center, and patients can self-refer. So then we get all the medical records, all the genetics that have been done before, and then we have like a huge multi-disciplinary team, and we review the case, go through it, do reanalysis, sometimes we resequence or do a new technology if one is available, like RNA-seq or long-read sequencing, and then we work together to try and figure out what's going on.

When I first started in like 2011, genome sequencing, exome sequencing was quite rare. So if patients were able to get it, a lot of times it was like shooting fish in a barrel, like we would have something like an 80% diagnosis rate. But now genome sequencing, next-generation sequencing is so common that we only see the families if they've already had a negative sequencing test. So we go from diagnosing like 80% now to like 10%, just because we're getting the most difficult of the difficult cases that have already been reviewed by really good geneticists um and getting to us where they just can't figure it out. But you know that's one thing that I think AI can really address is shortening this diagnostic odyssey for patients that really just have been jerked around, not through anyone's fault, but just by the nature of how these things go. And maybe AI can help in analyzing symptoms, or you know maybe you should see this doctor right away, or maybe you need this test, or go to this specialist, and just make things happen a lot faster.

That callback to 10 years ago I think is is quite interesting. Maybe you could kind of give us a little bit of a sense of like the relative pass-through rates at these different levels of filter people go. Initially to their local doctor, the local doctor doesn't know what's going on, they get referred, eventually they get to your hospital, you've got the best of the best there, and there's another related but distinct line of research recently that has been comparing AI's ability to diagnose through natural language with patients against doctors, and it seems like against at least the average doctor the latest models are now like very much holding their own. I don't know if that would be true if if we were looking at the Boston Children's elite uh clinicians and and their ability to diagnose, but they still don't know what's wrong. In the past, if I understand things correctly, because the sequencing was rare, you could often just do a full genome sequence and then be like, "Oh okay, well there's your problem," kind of like the literature has sort of characterized this. Now that we have this additional information, there's a pretty clear match. And now today that low-hanging fruit is getting absorbed somewhere else in the system before it gets to you, and you're now seeing things that are basically not characterized in the literature at all, or maybe just a little bit. And I'm not sure why the connection wouldn't be one that others could make, but yeah, use that prompt and fill in a little more detail if you would.

When I started at my—I was actually hired as a project manager at Boston Children's to help uh clinicians get their patient sequenced. So clinicians, even though they weren't geneticists, were really good at identifying cases that were probably genetic in origin. So they had freezers full of this DNA just waiting for the technology to come online where they could analyze it and figure out if there was something like genetic that could could be discovered. Um, when I—and when I started an exome, which is just 1% of the genome, is just the coding region, just the genes—I mean, it's a good place if you're being economical because like a lot of the variants are within the coding region. So an exome, again 1% of the genome, was $3,800, or it was close to $4,000. Now I just priced out an exome; it's $16 for that exact same test. So it was so expensive that they went through a rigorous selection process. If you were going to get an exome done, you were—it was pretty much we were—if you were going to bet money, you were going to bet money that it was genetic and you were going to be able to figure it out by doing a trio, that is like the patient and the parents, and you're going to see something that's de novo, which is basically not in the parents but it's in the child. So it's like a lightning strike, an error happening during development which causes disease. So the first ones of that were that 80% that I was talking about, and it was because these patients were collected, some in some cases like 20 years ago, and they had the DNA there, and sure enough there was like a premature stop codon or a huge deletion um of one gene that was already hypothesized to be related to this condition or a similar condition, and you could point at it and be like, "Yep, that's it." And then also you would have multiples of the same type of case where you'd see like four families with the same gene missing with the same phenotype, and then you're really confident that that gene is causative to the condition. As the price dropped, it became less of a thing that happened, and it's not because anywhere you could get an exome done, and there's a lot of geneticists, there's a lot of really savvy clinicians, there's a lot of for-profit companies that you could send off and get a report back, get diagnosed and have more precision medicine treatment and go on your way and do very well. So it was the negative cases that were getting referred up the chain to Boston Children's where they've already had a genome and it came back as negative—that is, there were no obvious, like no variant in the known gene that could explain what was going on. So then it becomes a little trickier. We start forming cohorts; Manton Center, we work with clinicians, we have um clinicians in every department of the hospital where they are able to refer patients to us; we consent them to our protocol, and then we collect samples, medical records, and we sometimes wait, and we reanalyze. And when we have four patients with like 10 patients with the same thing, we're able to look at them together as a group and be like, "All right, are there things in the same gene, the same family of genes? What can we come up with a hypothesis here?" And in 2014, I think Zach Kohane had been a previous guest on your podcast; he had the idea of having an international competition to solve undiagnosed families. So we got three families with seemingly Mendelian disorders—that is, that we think they're genetic, we think there's something going on with a clear relationship between gene and condition—and we released their data all over the globe to 23 different teams, and we had them compete and each submit a report on what they think the cause of the family's conditions were, and it was really interesting. Um, a lot actually were diagnosed from this; I think two out of three actually walked away with diagnoses from this process. And also we were able to show that diverse teams did much better—that you can't just have a bunch of like bioinformaticians in a room together looking at cases and expect them to come up with the right answer. It was teams that had a mix of like research assistants, genetic counselors, researchers, clinicians, research clinicians, clinical clinician, clinical geneticists working together, and all those diverse perspectives, they on a whole did—were able to solve more cases. Um, so that was really interesting, and I think that's a recurring theme here when we're talking about LLMs and large language models and AI, you know, none of this exists in a vacuum, like it's helping us along, and maybe it will be enough, but right now having multi-disciplinary, multi-strengths um all working together, we do much better as a whole.

The other thing I wanted to add is um we still see those slam dunks, though. So we just had cases a little while ago where it was one family and three generations all with a rare bone disorder, and the matriarch or patriarch was 90, in his 90s, and we were able to give a diagnosis to this person in their 90s, which I thought was really, really cool, and kind of shows the power of like just having an answer. And you know, he's already gone through like surgeries he didn't need to go through and had his whole life for this condition, but you know something as simple as being able to explain what's going on in like 10 seconds as opposed to three minutes of describing symptoms means a lot to the family.

Yeah, I imagine, especially if you've been dealing with something like that for 90-plus years, that's crazy to think about.

[Sponsor message]

So a lot of different questions I have about all this, but in these cases where you're getting all the way through the entire medical system basically and finally getting to one of these cross-functional teams, mhm can you tell us a little bit more about what the process looks like when that team gets to work? In AI prompting, we talk about thinking step by step and breaking problems down. Maybe one way to frame it would be like, what is the sort of collective chain of thought that the group goes through to take—start with inputs. Inputs would at least be symptom descriptions and results of genetic testing sequences. I don't know if there's any other inputs that you guys get at that level. I guess you have the whole scientific literature also is sort of an input, and then you do some thinking, reasoning process, maybe some additional testing, and finally you get to a result. What are you guys doing when you're doing that?

Okay, so when a case comes across my desk, usually there's a medical record that comes along with it, because again they've been bounced around for a long while, and usually at this point now they've had some genetic testing that we get transferred to us. Um, more and more patients are coming with it on the thumb drive, like, "And here's my genome," which I think is really cool and didn't even happen a few years ago. And we run it through our genomic pipelines, and usually we run it through more than one because they all have their strengths and weaknesses, like some are more comprehensive but they're harder to use, and you'll get more false positives because they rule less things out, and then you have others that are really easy to use, like my kids can use them and understand intuitively what it means, but sometimes they're black boxes and you don't know the reasoning behind why a variant was eliminated or not. I'm a PhD, not an MD, so I usually like things to stay anonymous; I don't want to be a walking HIPAA violation, so I kind of don't want to know the names or meet the families, but sometimes I do, I know who they are. And we go through everything case by case and line by line. There's certain phenotypes where I think more information is better, like you'll get the occasional phenotype that's only linked to one condition, like lack of tears in like one condition, and then that's a really important clue, and then we'll look at that gene. And then also for the what the patient is experiencing overall, generally there's gene lists of what has already been discovered, and you can look at the genomic information for any variation that could be causing disease—like we'll call it for simplicity, take pathogenic variation, though suspected pathogenic variation is probably more accurate to say—in those genes. So you get the new analysis done and you're looking at what's known, and then if you don't see anything, then you start looking like kind of your special sauce, like, "All right, how am I going to approach this?" You look at where—what else in the genome is notable? Is there a huge structural change that hasn't been linked to disease, or like a translocation where it's like chromosomes break and reattach on the wrong spots, or is there some other deletion or duplication? What's rare? What's unique to this patient? Then you also—now there's all these new technologies like looking at epigenetics, which is like you can kind of predict which genes are turned on and off, and even if you can't see a mutation, is the gene of interest like the expression perturbs somehow, where it's constitutively on even though it's not supposed to be? Can you take a look at that? Sometimes in the back of your mind you're like, "Okay, is it multifactorial? It's not just one gene impacting it; it's not some big error in one gene; it's a bunch of tiny little things scattered throughout the genome." And then a different type of a test, like looking at a GWAS, like genome-wide analysis, or S-CAT, where you can look at rare variation weighted by how rare the variation is and how predicted damaging it's going to be to a protein, and look at that and see, "Okay, is there some reason why you think that this is going on?" And a lot of times still, like let's say it's 25 to 33% of genetic testing comes back positive, so that means like what 75 to 66 to 75% are negative. Then you go through this whole process and still most are negative, and then you put it on the shelf and you wait a little bit and you analyze it again a year later. Reanalysis is actually really, really important because things get discovered all the time. You can't be an expert in every gene, every condition, every structural variation, and other people are actively working on it, and a lot of times you'll take something off the shelf and look at it again and it rises right to the top; the number one thing in the genomic browser is the answer, and you stared at it before a year ago and you didn't make that connection, now all of a sudden there is an answer.

Actually, I was asking Alan Beggs and Monica Weck, who are the director of the Manton Center and the medical director of the Manton Center, for success stories, because if they had any that stuck out, and one was 20 years ago—it was three siblings all passed away from a type of myopathy, and they couldn't figure it out, and they kept testing and testing and testing, and eventually ran out of DNA. And then we had a pilot grant at the hospital to do RNA-seq, and Alan and Monica submitted this family because we had some RNA left, and found a variant in CFL2, I think that's the gene name, and even though it was 20 years ago, the surviving siblings were now planning families, and they had an answer, and they could do genetic testing to make sure that there weren't two variants that they were each carriers of one variant—they didn't pass away—so they only had one, not two, but their partners didn't have a variant in the same gene and could ensure that the next generation wasn't going to have this horrible fatal myopathy. So in some ways—and then we had like an interesting discussion, like, "Okay, is that really a success story because like whenever there's multiple deceased, is that really a success? Like, yes, you diagnosed it, but it's not changing anything," but it is changing the future; it is—they're going forward with their eyes wide open and being able to to plan as a result.

Yeah, that sounds like certainly some form of success to me. I have three young kids, and fortunately no crazy medical conditions in my family, but we still did a little bit of genetic testing, and I would say I was probably never more nervous than when opening that report, just to make sure that I wasn't—have to see something really weird or, you know, strange, or that we changed the course of my life. So to be able to be on a, you know, potentially negative course and get the assurance that you could get on a, you know, confidently a path where you'd be able to have healthy children, I think sounds like a—to put it mildly, I would say a life-changing development for those folks. Definitely resonates with me.

Okay, let me dig back in a few points along the way in terms of how you've got this out. Maybe try to summarize a little bit and interject a couple questions. Okay, the pipelines that you're describing, those are—I guess maybe a mix of maybe commercial options or other groups have put things out. The inputs to those are they like highly structured data? I mean, I'm kind of thinking here, my sequence is of course structured, my symptoms are not, right? I describe myself in words, the doctor I'm talking to kind of notes that in words. Is there a way where that gets translated to like specific coded sets of symptoms, or what is the intake of these pipelines? And then are they basically doing like deterministic work where they're sort of running—essentially running down a long checklist and saying like, "If you have this, we check this; you don't have that, so that's out," and sort of just working down like a long set of known possible conditions, or how would you characterize what those pipelines are doing internally?

So I think it's—you're exactly right, like a lot of them you input the phenotype; it's coded to ontology, sometimes HPO codes um or ICD-9, 10, sometimes SNOMED, like there's a bunch of different ontologies. I like HPO the best. Um, be real clear in those sorts of ontologies, something like "no tears" would be like a single HPO item, and then so my condition might be summarized by like basically a set of those. If I had no tears and hair falling...

Out and you know, loose teeth or whatever, that would be like three things. Yeah, that would be like, okay, patient presents with this bundle of things, okay, and it's a huge uh field of research too. Like my friend Melissa Handel, like working with HBO and her site Monarch—Monarch initiative—and being able to map that onto um animal phenotypes and like making sure it's like one to one. Like humans don't have paws, but like, you know, like the phenotype that is most close to that is like being able to be translated.

And then also lay person HPO, where like we're not saying a laca, but we say like no tears or like, like lazy eye and strabismus, you know, and there's a whole mess of work that goes into that and making sure that it's accurate and also culturally sensitive. Like fits for epilepsy, you know, it's all this stuff that you never think of. And if you don't make those translations, then all of a sudden your phenotype is way less accurate than it could be. That yes, it gets incorporated into the model.

And then when you input that with the genetics and you can have raw data, which is like FastQ's and like the zeros and ones that come off the machine, and then a BAM, which is just—it's you're looking at the reads of the sequencing itself. Like um, gosh, I'm not going to explain this very well, but then you have the VCF, which is really processed data, and it's basically every single variant, but it's huge. Like a VCF is a relatively a huge file. I mean, it's orders a magnitude smaller than like a FastQ or BAM, but um it's still quite big.

And then you're putting that into—or BAM or FastQ into these pipelines which process the data along with a phenotype, and then it—it's ordering the variant based on the HBO code related to the gene related to the variant within that gene and how likely it is to be causative of disease. And then the more sophisticated ones can take in like these relational kind of things where it's known that this gene binds to this other gene, and gene A is related to the phenotype, gene B isn't yet, but there's a huge variant in gene B and the patient has the phenotype associated with gene A.

So my actual first ever success story was one of those cases where it's called episodic ataxia, and a patient will just get like—all our patient was getting really stiff and like couldn't move, like would get like locked in position. We did uh sequencing and saw that it was a variant in KCNA1, which wasn't the gene that we were thinking of, but it was related to the gene that we thought it was going to be. And so their KCNA1 just like rose to the absolute top of the list, which was really, really cool.

Hey, we'll continue our interview in a moment after a word from our sponsors. 2025 is shaping up to be a crazy year, and I'm getting a lot of questions about how people should manage their careers. Increasingly, my best advice is to go ahead and do what you've always dreamed of doing. If that involves starting a business, you should know that there's never been a better time than now, and there's never been a better platform than Shopify.

In the past, being a small business owner meant wearing a lot of hats and a lot of times spent doing things you didn't necessarily want to be doing—being your own marketer, accountant, customer service rep, and more. Today, it's increasingly about focusing your time, energy, and passion on making a great product and then delegating all that other stuff to AI. Of course, to get quality work from AI, you have to provide the right context, structure, and examples, and that's actually a big part of what makes the Shopify platform so powerful.

Shopify has long had thousands of customizable templates, and their social media tools let you create shoppable posts so that you can sell everywhere people scroll. Now they're building their own AI sidekick, Shopify Magic, designed specifically for e-commerce. All this makes it incredibly simple to create your brand, get that first sale, and manage the challenges of growth, including shipping, taxes, and payments, all from a single account. And if you need something special, Shopify also has the most robust developer platform and App Store with over 13,000 live apps.

Case in point: I'm currently working with my friends at Quickly to build an AI-powered urgency marketing campaign platform for e-commerce brands, and it will be launching—you guess it—exclusively on Shopify. Establishing 2025 has a nice ring to it, doesn't it? Sign up for a $1 per month trial period at shopify.com/cognitive—cognitive is all lowercase—go to shopify.com/cognitive to start selling with Shopify today. That's shopify.com/cognitive.

Trust isn't just earned, it's demanded. Whether you're a startup founder navigating your first audit or a seasoned security professional scaling your GRC program, proving your commitment to security has never been more critical or more complex. That's where Vanta comes in. Businesses use Vanta to establish trust by automating compliance needs across over 35 frameworks like SOC 2 and ISO 27001. Centralized security workflows, complete questionnaires up to five times faster, and proactively manage vendor risk. Vanta can help you start or scale your security programming by connecting you with auditors and experts to conduct your audit and set up your security program quickly. Plus, with automation and AI throughout the platform, Vanta gives you time back so you can focus on building your company. Join over 9,000 global companies like Atlassian, Quora, and Factory who use Vanta to manage risk and prove security in real time. For a limited time, listeners get $1,000 off Vanta at vanta.com/Revolution—that's vanta.com/Revolution for $1,000 [Music] off.

That challenge of basically what is the sort of graph of interactions and what affects what in the cell, or at the tissue level, or system level, whatever. I mean, that is a—been a fascinating area for me. Recently, I've been really interested to see some new projects. I don't know if you've come across these yet, but there are some that are now trying to predict the evolution of essentially the transcriptome or or cell state from of one timestamp to the next. And I think that that really suggests a major revolution coming soon.

How much would you say—I don't think there's like any answer to this because I don't think we know how much we don't know. But when it comes to those sort of interaction type things, my sense has been that we have a relatively small amount of that space illuminated today. Like of all the interactions, you know, of all the things where something in this gene because that interacts with the other thing could cause a third thing downstream, my sense is that we have a pretty small percentage of those pathways mapped out and well enough understood that we could do this kind of analysis. Is that a good uh summary, or how would you improve on my summary?

No, I think that's totally right. I think like every time I try to look at the impact of a variant on the protein, I'm surprised at how well—first of all, user-unfriendly a lot of these tools still are. And it's because like they're really tough, like they're cutting edge, and you know, protein folding is—it's come a long way, it's definitely super cool, and the people who work on that are totally hardcore, but it's—there's still a lot to be learned, and we're still folding certain proteins like we don't have everything worked out yet.

I just keep thinking about like when the first time we got genome sequencing and how difficult it was to use some of these browsers, and they would crash the computer, and you know, I think like protein folding in some of these tools like STRING—STRING DB—like protein-protein interaction, they're amazing, and they're going to continue to get more and more amazing, more useful as time goes on, especially when they get more user-friendly for people like me.

Yeah, sounds like that might be a a real low-hanging fruit. This has come up on a couple different episodes where the the general observation has been biologists are not programmers and doctors are not programmers. There's a missing layer that would unlock a lot of value if we could just make it a lot easier for doctors and biologists to use the models and other information tools that have recently been created. But a lot of times those are still kind of put out there in open-source project form, and they need like a UI layer on top or an orchestration layer on top to to really make that accessible and useful for a lot more people. So that could be an interesting area for somebody to dig into more.

Yeah, and just little things like I got some sequence back from a new company, and they're like, okay, here's the commands to download your data, and I'm like, whoa, whoa, whoa, what? And it was like—and they had no intentions of helping me either. I learned command line and how to get my data from their server down to mine, or I didn't get my data. And so like I had to have a crash course on getting onto the Harvard—the Boston Children's—supercomputer in order to get my data, and it was like a huge waste of time. And I think they're assuming a level of literacy for some of these programs that just—you know, it's people don't have—you can argue that I should, being in the job that I'm in, but it's hard—it's a learning curve. And yeah, I think there's a lot of opportunity there for making things a little more friendly. And it goes back again to like you don't know what you don't know—like making your tool accessible to a wider audience, they're going to apply it in ways you never dreamt of. So gatekeeping it to only people who know Unix is kind of tough on everybody.

Let's circle back to that in a second because this sounds like one of the candidate areas where you might be getting some good value from your 01 Prog grant. The—these pipelines—are they using any sort of predictive AI technology like classifiers, things like that, or are they kind of working off a sort of accepted known literature of findings? Because I could imagine—and maybe it varies across provider—but I could imagine one form of pipeline is like we want to be just really grounded in things that are very well established, and we're going to, you know, run down this super long checklist programmatically for you and try to find things that fit. And I could imagine another pipeline that would be like if these models exist—I'm not sure to what degree they do—you could say sort of hey, here's my genome, like predict—you know, predict—and give me guesses. Are there models like that? And I guess, you know, to what degree is this all deterministic versus sort of already at those existing pipelines starting to lean into certain kinds of AI?

I think you need both. Like you need to be confident that you've looked at a genome with all the known things that nothing funny—you know, just very validated best practices—and then you need the exploratory pipelines. And that's what we're developing, and—developing as part of my grant with um OpenAI is really—what's the limit? Where can we take this? Where can we like make shortcuts that were before taking a ton of compute and a ton of time? How do we solve cases faster? How do we—what's the minimum required data set in order to make a diagnosis? What's the minimum like compute necessary in order to get a diagnosis? How do we diagnose new things? How do we come up with new hypotheses faster? All using AI.

Well, that's probably perfect teup for your application of the latest models. Yeah, maybe for calibration there before we get into like workflow specifics. When did large language models start to be useful for you? Is it just with 01, or were you already starting to see some value with earlier versions?

Well, we had been using it along with the phenotyping areas more than anything else. I had a prior grant working with Ingrid Holmén, Melissa Handel, where we were trying to take a patient phenotype and map it to HPO codes, and again, the lay person HPO faster and more accurately. One thing that we did use at one point was working with seven questions and like asking—system that's affected and drilling down that way and seeing the ability to get an accurate phenotype through an interactive model and using your own words compared to traditional self-phenotyping like surveys and things that are on the web now. And we're analyzing that still. That we were using it—there are a lot of publicly available tools that I was using like that I've mentioned before, like AlphaFold and STRING DB, and a lot of these protein impact prediction models that are required to do our jobs. Like we need to be able to predict the impact of a variant on a protein. We can't treat it as gospel. Like people who rely too heavily on these algorithms sometimes get tripped up because, you know, some of the known gene-disease relationships went past those filters now because there's just something about that gene that you know—you perturb it a tiny little bit and it causes a phenotype that you wouldn't even think it is, but we know that's true. Um, so if you looked at it at face value, you would have skipped over it. You know, I think a lot of people are using them and they don't even know they're using them, like—and they don't really know what's behind it. They just know that oh, yeah, you—you look at CADD score, you look at SIFT, you look at PolyPhen, you look at protein impact, and then that's a cutoff, and then along with allele frequency, but not really realizing that like the aggregation of allele frequency is really powered by a lot of these models and just a ton of stuff behind the scenes. If you took it away, we would be struggling.

So do I have it right then that with like an AlphaFold type model—this is sort of after a standard pipeline basically comes back negative again—then you would say, okay, let's go into essentially anomaly detection mode for this person's—exactly—sequence?

Yeah, and you have tools for that as well that—conser of say, hey, look, here's a giant deletion or, you know, this gene is stopped prematurely, or this one has been like copied over a bunch of times, whatever. There's—of course, you know, plenty more—I'm sure different ways things can be weird than those—but you identify those, and then you sort of say, hm, I wonder if that maybe is the thing. I will use AlphaFold to take that genetic sequence, see what that protein actually looks like, and then do a structure comparison and see like does that look like that protein is really mangled? And then if so, that becomes like a place to go deeper.

Yep, exactly. And a lot of that comes with experience too. Like there's some genes that are really mutated in pretty much everybody, and so if you don't know, you're like, oh, look at that—that—that's so cool. And then some veteran is going to be like, no, it's not that—it's never that, or it's never lupus. And then you see the gene that you've never seen before, and it has a variant in it that's conserved down to zebrafish or C. elegans worms, and then you look at it and AlphaFold and you see that it's royally messing up the protein, and you get excited. I mean, it's a roller coaster a lot of times. Like even that will fall apart somewhere, and then you'll find out that it's only really common in one specific ethnicity that's hardly ever sequenced, but like the patient is that rare ethnicity. And it goes to show that we need to sequence the whole world in order to understand what is actually disease-causing and what is just background variation and isolated in other populations.

Yeah, I had—there's another fork in the road which question to ask here on the—we'll come back to the data, cuz that—that is a can't-miss area—but just take us a little bit further down this path of like, okay, we have identified some anomalies. Now we run the folding model, we see that the structure looks off, where do we go from there? How—what's like the next investigation after you've identified that?

So back in 2011, you would get really excited about it and like want to publish. But in 2024—the waterline is rising—always for sure. Yeah, the waterline is rising, and now people are like, wait a second, that might just be random. So then you like want other families or other cases with the same type of thing—variants in the same gene—and there's all these sharing tools to be able to do that. One is called Matchmaker Exchange or Beacon, where you put in the variant in the patient phenotype and you see if anyone else has put in that same gene attached to the same phenotype and you match, and then you collaborate. Or somebody's already started a paper with 19 cases of variation in this gene causing intellectual disability, and if you have one, you can add it into that case series and get a better publication out of it that is much more convincing than if you just—com—publish your one case, which looks pretty cool and you're convinced, but other people might not be by reading it. Just the bars continually being raised on this stuff.

Yeah, so that brings us back to data naturally. How would you characterize the data environment? I was struck in reading through a couple of the papers—I don't have the vocabulary to go as deep as I might wish to on all of your papers—but I was able to see quite clearly that the N is small in a lot of these papers—like single-digit numbers of cases. And then I've also kind of noticed a few times you've spoken about like the hospital as sort of the data unit it seems like, and I've heard from a bunch of people actually over time that like, yeah, we have this sort of scarcity of data. And I've always kind of wondered like, is it a true data scarcity problem or is it a sort of man-made—for lack of a better term—data scarcity problem that really is more about like barriers to access and and sharing?

It's a tough situation. I don't want to fault the young researcher who doesn't want to share their super cool case because, you know, they're hoping they'll find another one and be able to publish it as their finding, not as somebody else's finding in a giant research group facility across the world where, you know, they're just going to be a middle author and it's not going to make their career the way that if they held on to it tightly and did everything themselves and got it out there. But the problem is with that is a lot of times it doesn't work out that way, and then that's not benefiting patients. You're not thinking of the patient; you're thinking of yourself. And it's much better for science, it's much better for the patients in general if everyone shares their data and has it open, and if you see something in someone else's case that you're allowed to match it with that group that's already working on that gene and put it out together. I mean, it's tough. Like Children's is really great in that we have this CRDC—this cohorts committee—where you can see other investigators' data, patient data, genetic data—not the phenotype, not their name or nothing identifiable—sorry, I need to make that extremely clear—but if you have a gene that you're working on, you can put it into the CRDC and come up with all patients that were seen in the hospital and their genetic variation in that gene, and then the physician—they have like a de-identified ID number—and you can email the physician to find out more information about that patient. And I've joined like national, international studies that way by like having a candidate gene. I go on to—it's called GeneDX browser now—and query the entire hospital—everyone who's been sequenced and has their data up there—found like four other patients, emailed the investigator, and they're like, oh, yeah, this person in the Netherlands is putting together a case series. Emailed them, got my patients' information into that case series, and now like it's awesome—like they're linked to experts, and we're publishing an accurate like comprehensive view of what that condition looks like. But it's hard—it's—I understand the dilemma. And for the young investigator who really just wants to, you know, they've been working so hard and they want the credit for what they've been working on, they don't want to hand everything over, but it's important that they do—everyone does.

You're identifying a barrier to progress here that I had not even considered, which is the investigator holding information more closely than it sounds like they should in some cases.

Yeah, I guess if we were to imagine an ideal data-sharing scenario and you know—exactly—how do we square the circle on sort of sharing versus privacy? Is obviously a tough question. Maybe there's like a cryptography-based solution that we could imagine, or maybe we just need to sort of change our norms a little bit around like how willing we are to share genetic information. I guess maybe you have thoughts on this, but I've always kind of felt like that doesn't seem to me like a huge risk that I'd be taking to share my genetic information with some international pool of information. But there's multiple different angles here, but I guess I'm—I'm kind of wondering like if we were to move from today's data-sharing reality to an ideal data-sharing reality, how much of a difference would that make for people who have these rare diseases?

I'm just spitballing here, but I think it would be huge. I think there's a lot of cohorts in the back of the freezer that just haven't been sequenced and haven't been shared more because of, you know, not apathy, but it's harder to do so. And also sometimes at a very superficial level, it's—it's hard for the investigator to get there mentally, doing that. But I think if they did, there would be a lot more discoveries, um, and there would be a lot more diagnoses for patients—that's for sure. That's why I always tell patients like—or—or people—if they email me they're like, okay, my child has this, I'm like, well, here, enroll in this program and this program and this registry. And they're like, why not just one? I'm like, you want to like do as much as possible. And registries are really important because when there's a new discovery, they go straight to the registry to find patients, and that way you're—ensuring—making sure that your sample isn't being left in the back of the freezer and they'll get to it when they get to it because like there's just—hitting it from multiple sides, multiple angles.

Yeah, is this sort of akin to—me—there's a few of these like pivot points maybe in the medical system where, you know, a lot of data is of course like locked up in electronic health records, and you know, we sort of have this like nominal interoperability requirement that somehow gets cashed out to like everything gets faxed around, and it's like, what the hell is that? That seems like not what we intended, and yet it hasn't been fixed. And then there's like price transparency is another thing that, you know, that's outside of the scope of this conversation but is definitely, you know, the kind of thing people have high hopes for, you know, if they—if you could get a price menu on the wall, maybe that would help in certain ways. There's like right-to-try is also a big movement, you know, where people are like, you're not gonna let me try this experimental drug even though I'm dying, like I should have that right. This feels like it could be like another candidate for similar reform where if I was going to try to whisper into somebody in the new administration's ear, I might say, hey, look at the requirements around sharing of this information—like could we change the defaults here in a way that would move the needle in a big way?

It's interesting that you say that. Like again, going back to 2011, one thing I was hired for was shifting it—being in the biobank—just your samples, your discards, tissue, urine, anything that wasn't used that they took from you as an opt-out, not an opt-in. And I still think it's an opt-in—how many years later? There's a lot of inertia around this—like to be able to facilitate broad sharing, especially for these tertiary cases where privacy isn't really the number one thing on anyone's mind—it's like moving as rapidly as possible and making as many discoveries as possible in a short amount of time. I really think like decreasing the barriers to sharing and—WR to try—and uh Tim Yu, who made Muon, is two floors down from where I'm sitting right now, and there—it's just like this incredible story of like him seeing an opportunity to make an N-of-one drug and then just like an extremely motivated family breaking down barriers to make it happen. They were so brilliant and motivated and smart and like were able to do it, and you just think like, okay, if you made the hurdles less extreme, like how much more would be possible? That's an incredible story if you don't know it yet.

Yeah, I don't know it, but here's hoping that we might have fewer of those stories and more healthy defaults going forward. That I mean, those stories are inspirational, but they sort of represent like a dark matter of probably a hundred other families that just couldn't—yeah—for some reason get over those barriers. And like some things are just so simple and maddening. Like we have a bunch of cases at the M center where we find the diagnosis, and then we need it to be—CleA confirmed—that is like we do stuff in the research realm, and then you have to get a new sample and verify it in a—a specialty lab called a CleA-accredited lab—to—and—and then have the finding return to the family through a genetic counselor or physician. And sometimes we'll call the physician, and they won't play ball with us; they don't care; they don't want to deal with it; they don't see the value or what it's going to change. And in my own family, I haven't been able to CleA confirm a finding in one of my relatives because the doctor is like, well, I don't have email—well, what's the value of this—like whatever—and it's just like, oh my God, like this is what we're up against. And then you like times that by like people not counseling correctly and getting the families into research programs, and as much as we're trying—as hard as we can—there's still so many barriers. And to bring it back, I'm really hoping that AI can break some of this down and put some of the autonomy in—like our ability to act—into the hands of families and patients so that they're less reliant on some of this infrastructure that doesn't work as well as it should.

Yeah, I mean, this is um an eye-opener for me. I think often about sort of will we end up in a similar spot with respect to AI as we seemingly have with respect to nuclear power, where it's like somehow we have thousands of nuclear weapons deployed, but we're still burning a lot of fossil fuels because we haven't been able to get the nuclear reactors—you know—nearly as many as we have the nuclear weapons. Like something seems very off about that outcome, and I can imagine an analogous version for AI where we sort of have what we need, but through a sort of combination of errors and barriers and, you know, sort of obstinance, we like never quite get to the actual benefits that—that we could get. And it sounds like there is definitely some work to do here to make that change in this area. A lot—lear—L—it.

Yeah, so how do you think this changes going forward? I mean, we could talk about this from the patient level and what they can do. I always say that if it's me, I at this point would not be—I would go with both—both the human doctor and the AI doctor. Like I'm—I always, you know, would have the conversation with Claude or ChatGPT in advance. If they don't want to talk to me, I say I'm preparing for a conversation with my doctor, and that gets them to open up and not—not worry about providing unlicensed medical advice. So the patient experience could be quite different. You could talk about that. Also really interested in just kind of how you are applying these latest models in your own work and where it's saving you time, what it's allowing you to do that you couldn't do before. So pick your favorite approach for that, but

I definitely am interested in the AI-enabled future of all this, so this isn't really that crazy or anything. But I would say the biggest impact AI has made on my research is summarizing articles and genes. Like being able to eliminate the time going down rabbit holes—like looking up a paper, oh it's paywalled, okay, logging into the Harvard Library, getting that paper, skimming the abstracts—it's not at all what I want, then going back and like, you know—and being able to ask for a summary and get it and either be like, oh yeah, this sounds good, or move on with my life—has changed my life. Like it's, and it's kind of wild that it's given me hours back in a day and how much time I used to spend doing that.

I think there's going to be a whole host of new tools where, you know, or new reasoning. I find it funny that like, yeah, sometimes it will like clam up and doesn't want to do it because you're getting too close to medical advice and like maybe just like specialty things that help, and you don't have to ask the same question four times to get it to answer would be really cool. Boston Children's also launched Chat GPT behind the BCH firewall, which is great because then you're not worried about things going out, and they're able to maintain much more control, and it stays much more accurate. I still can't get citations to work properly, which is kind of hilarious, but it's getting way better—like the hallucinations are getting way better. I, I just think it's going to be moving at light-year speed.

Going back to what we were talking about before, I think there's a lot of fear around it that's going to have to be addressed. But hopefully the one bad situation isn't going to be the only thing people read about it, and like some of the really great things that come out of this will also be properly publicized to kind of give a more balanced viewpoint. And again, like keeping in mind that a lot of times these are really severe cases, really severe patients; you know, they're making huge strides and impact, and keeping that in perspective is important, too.

So tell me more about like just some of the things that you actually throw into Chat GPT. I mentioned the one is just kind of, here's my situation and here's this paper—almost like relevance filtering, like is this relevant—what other sort of tasks do you find yourself bringing to, especially the latest models?

Well, I also run the core facility here, so I'm tasked with learning a lot of new genetic techniques really quickly. And if something comes up and I don't know what they mean, I could Google it and find the one obscure paper; I could put it into Chat GPT and like learn about this new type of sequencing that, you know, is only launched at Children's and has one paper attached to it, and here's a nice summary that I can understand as opposed to like weeding through anything. Or I meet with investigators all the time, and being able to summarize their work really quickly allows me to do a much better job in my one-on-one consultations than I would have. Also coming, like there's close to 20,000 genes; like anytime I get a case with, okay, they think it's this, like sometimes I know what that is, other times I do not, and and I am able to print out a summary of the condition really, really quickly and nicely and also get the latest on it, also who's working on it, and just go into a meeting much more prepared and much less time.

Also, when you get a paper, a lot of times for some reason always seems to be reviewer number two is like, there's a whole body of literature on this and you don't really know what they're talking about, and being able to address some of the critiques and kind of put them into context and what they're doing—I mean, it's all cutting down on this mundane, time-consuming, really tedious part of the job that and getting back to the fun part, which is gene discovery. And going through a list of 20 possible candidates and narrowing it down to three that you're going to present in like an hour and a half—true story.

So let's do that true story maybe in more depth. I mean, is that, that's another thing where you're using Chat GPT to help?

Yeah, I mean, why not? Like if you have 20 genes and you have the phenotype and you they all seem pretty interesting, like you can go through and look at the protein impact, so like order a CAD scores or conservation and be able to order it that way, but then doing a really quick relevancy assessment using Chat GPT, it saves a lot of time.

So how do you set that up? Like do you have a prompt template that you go back to over and over again? How much have you had to develop that? How much do you have to give in terms of like detailed instructions or examples? You know, we're getting into the nitty-gritty here, but this is the part where I think both people hopefully can learn from your experience, and if nothing else, you know, just demonstrating that this is possible I think is is quite useful because there's just so many people, including like software developers—it's you'd be amazed—maybe you have seen this, but you'd be amazed by how many software developers tried GitHub Copilot 18 months ago, you know, when it first came out with like the 3.5 model behind it and were like, ah, it wasn't that good, you know, it can't help me. And so I think there's just a lot of value in sort of object lessons of like, here's hard work, you know, that highly skilled, highly educated professionals are doing that Chat GPT, or obviously other models perhaps similarly, but we're focused on Chat GPT in this case, can really help with. So yeah, I love just kind of as much detail as you can get into in terms of how you actually go about setting these things up, how you've iterated on them, etc.

Okay, so for an example, I work on um bladder pain, undiagnosed bladder pain and individuals. Um, it's really severe; sometimes they can't leave their house; they're in—it's called Interstitial cystitis bladder pain syndrome—and um there's no real gene attached to it. We found a couple where it seems like there's way more variation in that gene than you would expect given a general population, and so it's a candidate gene; it's in no way like a slam dunk. But I have around 500 patients in a cohort with this, and then we have done—I've done in conjunction with Josh Modo at Columbia and Ali Garavi—kind of assessments of my cohort and other cohorts, what genes have way more variation in them than you would expect, and you can come up with lists, and then you can also come up with gene pathways—like multiple genes—and then these pathways are interesting because a lot of times they have like a label—like the small molecule transport pathway; there's like 12 genes in it—and then you want to know, are any of these genes like tied to bladder pain? Are they tied to the bladder? Are they tied to bladder cancer? Are they tied to anything? And like being able to ask those questions really quickly, and then sometimes a simple yes or no, and just putting them in in a string and then coming out with yes or no, it saves a huge amount of time. And then the ones that are yeses, then you can drill in, and then I always check the nos too, just in case, like, you know, it's still early yet. But like I, I was doing that last night and was able to just get through these pathway lists and be like, all right, this one has 60% of the genes that have a tied to bladder cancer, specifically bladder cancer, which means then there's they're expressed in the bladder, and there's known perturbations that cause bladder dysmorphology or bladder conditions, so this is more interesting than anything else. One I was actually—I almost screamed—because the gene was linked to urothelial issues, which is a great mechanism of disease, and I'm going to definitely follow up on that; I have a meeting tomorrow morning to discuss it. So it really just helps, you know—I, I imagine I only started working in genetics after the genome was published, so I don't know how people did it beforehand, and I think there's going to be this whole generation of geneticists that aren't going to know how things were done before all this was available because it's going to be a huge game changer and timesaver.

So is that—how much difference would you say you see between, for example, GPT-4, 001 Pro, when you bring those kinds of questions? And because I can see sort of interesting different trade-offs, right? Like in Chat GPT today, if, if I call, call correctly, maybe they've just updated this, but with certainly with GPT-4, you can enable web search; with 001 Pro, search is unavailable. That's what I thought, and that is still the case. So if you're, you have the sort of GPT-4 could like go out online, find information that's maybe more recent than the knowledge cutoff, which could be really useful, but isn't going to reason about it in the same way. 001, 001 Pro, you have this more reasoning, but you have knowledge cutoff issues, inability to go out and supplement at runtime. Do you have sort of a taxonomy of like what models you use for what things and how you know when to trust what it's saying versus when you need to fact-check, and how much is like the reasoning adding over the the sort of 40 for your purposes?

I think I'm becoming more and more convinced over time that this is going to revolutionize things. Like I was skeptical at first; I was like, oh, we're going to have to check every single thing—is this actually saving any time?—and it's just getting more and more accurate; the reasoning is getting better, and sometimes you'll be just so pleasantly surprised—you'll ask a question, I'll say, okay, answering it in the form of a genetic counselor is this, and then like completely surprise you and be like, another way to look at it is this, and you know, and it's, it's just, it's doing an amazing job. I know I'm a convert, but and an early adopter relatively, but you know, I think the sky is the limit really, and we're going to get to a place where, you know, it's going to be solving cases and shortening the diagnostic delay and democratizing access to genetics interpretations and, you know, sidestepping a lot of the barriers that we have right now. It just needs to convince everyone that it's accurate and the reasoning is good, and over a percentage of time it's kind of like hypocritical in a way—I think we're going to have like a higher bar for it than we do ourselves—like we can say like, oh, sorry, I, I missed it; I shouldn't have—and we're not going to forgive it if it misses something. Thing, but I guess that's the way it should be. Yeah, I'm not sure if that's the way it should be; it does seem like it's the way it is. The, um, in self-driving, my general working assumption is that it's going to have to be 10 times safer or have like a one-tenth, you know, the danger rate to be acceptable to people, and I would guess, you know, probably something similar like this will happen here, at least when it comes to actually like putting it, you know, in a more forward-facing role where like patients could, you know, access these sorts of things, yeah, themselves. If it's a tool for the professionals, then we maybe are a little bit more, you know, put the responsibility on the professional and can use some earlier. But yeah, I would probably advocate for going for it before it gets to 10x better, but nevertheless, that does seem like the sort of mentality that we have.

So just honestly, for me, maybe for the audience, but for my benefit, how are you managing those trade-offs between like needing to go out and search? Because if you wanted to use 001 Pro, you'd have to go do your own search, copy-paste in, let it do its thing; GPT-4 can do its own web search. So in like the very nitty-gritty, what model do you go to and how do you set it up for success?

I'm not using the web search right now; I'm more using 001. I think I, I mean, I'm playing with 40; um, it's moving so quickly that it's more—there's no sophisticated reason for that; it's just kind of like what I'm comfortable with, and then um moving from there. And I'm really impressed with 40's reasoning. I think web search still kind of scares me a little bit, just because there's a lot of garbage on the web, and I have to really be confident in any answer I'm getting out. I think like checking everything is still paramount here, but hopefully it won't be that way in the near future.

How long does it tend to think on the questions that you're giving it? Long? Like at first, I think it's shortening too, like by the day. Like at first, I remember it would just be kind of hanging there for a while; I'm like, are you okay? Then now it's just really fast, so like under a minute in most cases. It sounds like also I think I'm getting better at the prompts, like as you said, you have to learn how to ask it stuff too for it to come out with the right answer right away.

Yeah, I, I would love to, you know, we'll trade—one thing that I imagine you've probably also found, but I've definitely found even in low-stakes situations is I try to be really neutral in the way that I ask questions because I find one of the most common failure modes, at least from what I experience, is the model running with a preconception, which might have been a misconception on my part, and kind of mirroring that back to me. And I'm not doing genetic analysis, but even in terms of like how to solve a programming problem or how I should think about architecting my application or whatever, a lot of times if I give it a sense of where I'm leaning, it will lean that direction too, perhaps without good reason. So that's one—what else have you found to be important in prompting?

I found that it actually thinks a little too much, where I'm like, is this gene related to this phenotype? And I'll bring up a study, and I'll get—I'll look at the study, and it's the gene related to it, like it's a step too far down the chain, which is really impressive because like it made that intellectual leap, but I need something simpler—like I need a paper that's just linking that gene into that phenotype. I was like, no, too far, too far. Be like how to ask it so it's not thinking as much, and it's just—and when you describe that it's bringing up a study out of its pre-training knowledge—like you give it a question—yeah, it says, so and so et al found this, and then it sounds like it is also marshalling knowledge of the graph of interactions and sort of saying, well, this paper showed this, and then I know this, and I'm like, I'm not making that case in the paper; I just want you to say like, it's upregulated and cancer—like that's all I want. And that was actually kind of wild because then you're like, okay, it's thinking—it's really thinking and making conclusions—that is quite interesting.

Do you think that those things are real and we're just not there yet, or is it going off in a direction that is just you think like just fundamentally not super productive when it does that?

That's the million-dollar question; I don't know. Maybe I'm not smart enough to understand it, and it's right; I don't know; we find out; we just need to keep playing with it and keep working with it and keep using it. Like this sort of is like the move 37 equivalent in—yeah, of course, I'm sure you're familiar with the AlphaGo championship from years ago where move 37 is like, you know, AI shorthand for a output from an AI system that is like very surprising to the human experts and nevertheless proves to be, you know, a genius move—like it was one of these moves where was like, oh wow, this thing is playing Go in a way that we never thought Go could or should be played, and like we actually have something to learn from this system. They initially thought it made a mistake, and then it turned out was like a genius move.

Do you think is there—it sounds like you at least have some allowance for the possibility that some of these weird analyses you're getting back like might be move 37-like brilliance, but we just don't have easy ways to even resolve yet whether it's like going in the right direction or not?

I have faith in it. I think that we just have to keep an open mind and again, just keep playing with it and see what it can do. How much data do you have to—do you actually throw like whole cases in?

It's kind of too hard to do that right now. I'm not throwing in like a medical record, even if it's behind the firewall. I'm just—I'm summarizing; we're building up to see the limits. So that's actually part of the grant that I have with OpenAI is to see how far we can take this and how much it can handle and how much it can replace me. And the barrier to doing that right now is that a context—like I could imagine multiple different reasons that it might not work to just take, you know, the simplest thing I would try, but then I'm sure there's going to be a barrier. But if I just said, okay, here's my whole medical record and here's my whole, you know, genetic summary of of all the, you know, strange variations whatever, maybe take the top, you know, chunk of that file, copy-paste, analyze this—what makes that not um viable today? I mean, the amount of compute needed for that and then also the opportunity for tangents and all the utility—like not not everyone who smokes gets cancer, you know—like we have to figure out what we're asking and what's relevant and what's meaningful and what's a good use of resources. I think there's forgetting that every question takes energy, so we can't just throw everyone's medical record in there and everyone's genome and see what comes out. We have to be kind of thoughtful about it and see in the cases where it does have utility and what information is necessary. And also, though to your point, I think it'll be really interesting about trajectory and predictions and, you know, all the medical record mining that we, we're doing now—can it do it on steroids and like come up with predictive models and and interject like, okay, I know normally you would want to uh see a a colonoscopy at 40, maybe you need one at 24, you know, just based on genetics and everything that it's able to see that we're not smart enough to see yet. I mean, I think the possibilities are really—it's a really exciting time.

Yeah, I mean, Sam Altman has recently said that they're losing money on the 001 Pro subscriptions even at $200 a month, so it sounds like people are maybe not in general being so conscious of the compute that they're consuming and just kind of throwing a lot at it. The budget's got to be pretty high, right?

Yeah, I mean, in terms of like from the medical system, like what people would pay, willingness to pay or like what insurance is prepared to pay compared to like a $200 a month 001 Pro subscription, I assume, you know, that looks very cheap, right, by comparison to hiring professionals and teams of people like yourself.

Yeah, yeah, and people like using it. I mean, well, everyone I know loves using it. I don't know if that's sample—so what else is going on with this grant and your relationship with OpenAI? Are you like working with them closely and iterating on use cases and giving them feedback, or what is the dynamic there?

Yeah, exactly. They've been just wonderful and super cool, and it's fun meeting—working with smart, motivated people—like I get emails at 11 at night on a weekend—like they're working on—they work hard; yeah, there's no doubt about that. So we're just getting back into it after the holiday, so hopefully next in a few months I'll have something really exciting to talk about. I, I'm just blown away that they're so forward-thinking and able to support this type of work. It's happened a lot faster than I would have thought, even just a couple years ago, that's for sure.

Maybe in terms of wrapping up, what it—it sounds like we have—if I try to summarize everything here, we have a data-sharing problem, we have a limited capacity for analysis problem as humans, mhm, and one of those is going to require a non-AI solution; the other one AI is increasingly ready and able to do a lot of analysis, but then you still have a number of kind of practical issues around like, well, there's the knowledge cutoff and the search, and you do your own kind of curating the search, and you can't throw everything into it because maybe it's a little too big for the context window, and you don't have all the workflows you might like because it's like all in a browser, and so you kind of have to like paste stuff in and get stuff back, and it, it seems like it's still like fairly manual. Could you—this is almost like your OpenAI customer uh interview, but like what do you imagine the experience being like, say a year from now, maybe when we like refine a little bit more and like do a lot more integration of these systems? What do you think that could be for you and for patients?

That's a great question. I think they might wrap things so it's less free-form, like, or be able to guide patients—make it user-friendly—with a point—like you're not just with a prompt, and then you're meant to come up with the correct question to get it to answer what you're thinking about. I think that's pretty low-hanging fruit and will be very useful for patients. I think for researchers also, the prompt engineering and use cases—I think we're still figuring out, at least at our institution, where it will be most impactful and what people want to use it for, so that's going to be clearer in a year—just getting people in the door and communicating, and we have like surveys—like, okay, what would you use it for? What have you been using it for?—and clearing that up and then making that better. I, I'm trying to be realistic here; I mean, I would love to say that we're using it to solve cases; I don't know if that'll be true, but I hope it is. I think it's just going to be more intertwined in our day-to-day existence.

What do you think?

All bets are off. I don't know. I mean, 003—that's another question I had—are you on the review team for 003 at this point? If not, I assume it'll be coming your way before too long.

Hope so. Yeah, I mean, it looks like that is another significant step up in just raw reasoning ability—the math—the frontier math results in particular—everybody was kind of citing the like 25% success rate, but that's like the very high compute level that costs maybe thousands of dollars a problem or whatever, but it was also really notable that the low compute setting was still like 10%, which was like five times better than anything that had come before, which was maxed out at 2%. So it seems like the 003 series is going to be another pretty serious step change in terms of just how hard of a problem these things can solve. We might start to get into context window limits being binding if there's just too much information in a medical history or in a genetic file, but I suspect that those can both be like filtered and kind of summarized and, you know, kind of boiled down to what matters most, such that even in a couple hundred thousand tokens, which is what they currently have—I would honestly take the—we probably will be solving cases with a 003 in a year's time—side of that bet, or for that matter, a 004, because, you know, the the gap in time between a 001 and 003 was so small, and the the signals that we're getting are like—we don't really see this slowing down—like there's going to be, you know, more progress on this front. So I'm always kind of wondering what's missing, and increasingly it's like harder and harder to look or harder and harder to find, you know, or pinpoint the things that are really missing, and that doesn't mean there's nothing missing—I'm sure there's still some things—but it is increasingly hard to say what they are. And my best guess would be that you probably see at least some cases that you could just throw in a 003 and get like meaningful, insightful conclusions back in a year.

Yeah, maybe we should get together again in a year and review the progress.

We'd love that.

Cool. Well, anything else on your mind today? I really appreciate the introduction to all your work and how you're using AI in it, but anything else on your mind before we break?

Just thank you so much. I think next year this time might be kind of different, so yeah, hopefully that seems to be the new normal—change is the only thing we can really count on.

Well, I'll look forward to—put it on my calendar now—to get back together in a year and see where we're at. But for now, Kath Brownstein, MPH, PhD, assistant professor at Boston Children's Hospital and Harvard Medical School, and a recent recipient of the Chat GPT Pro Grant, thank you for being part of the Cognitive Revolution.

Thank you. It is both energizing and enlightening to hear why people listen and learn what they value about the show, so please don't hesitate to reach out via email at TCR@turpentine.co, or you can DM me on the social media platform of your choice. [Music]