📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Global Namespace for Precision Medicine: The Data Breakthrough w/ William Baird

Hammerspace12:02

Transcription

Hi, I'm Molly Presley, the host of the Data and Chain podcast. And as you can see, we're broadcasting live from a trade show today. We're at Supercomputing 25 in St. Louis. And really excited to be talking this week, not just about high performance, but also about data access and the whole concept of gain access to more data globally.

So, if you're new to the Data Unchained podcast, let me tell you a little bit about the podcast. We founded this talking about what are some of the challenges of gaining access to this decentralized data that's now being created in edge devices, different instruments, different clouds, and maybe needing to be shared with different applications than it was originally created for. We're really excited about today's guest because he has driven a thought leadership initiative uh around some of the challenges in the life sciences community. in this exact area. So without further ado, let me introduce William Barrett. William, thank you for joining today.

>> Thank you for having me.

>> So William, um before we talk about data and decentralized data, maybe tell us what is your role? You're with Garden Health and tell us a little bit about Garden Health as well.

>> Sure. So uh William Baird, I am the associate director for Hyperborns computing at Garden Health. We are a liquid biopsy company and our focus today and I stress today um is on cancer. So basically oncology in all its different flavors and forms. We specifically do detection of cancer and its various mutations for advanced cancers or we do monitoring for whether or not how your cancer treatments are progressing or another one of our products is for early detection. So basically for the screening that we do so that instead of having to go and do more uncomfortable uh regular checkups we can actually go off and just do a blood test. Everything do is based on blood.

>> So it goes through a a wet lab process. It ends up in um going through a genome sequencer and then from the genome sequencer it goes onto our high performance computing environment and there it produces quite a bit of data and today we're now up to curating about 60 pabytes. Um, and all of it is genome. Um, we we have we're either at or past a million patients.

>> And so all of this data is it actively being used for research? Is it a big archive? Maybe just talk a little bit about how it's being used.

>> It is both used to reproduce a report for the doctor so the doctor has better results for the patient. It's a precision medicine. Um, so they not not all cancers are created equal. Some are more resistant to treatment. Some have require specific treatments and others are you know they change over time. So people need to have a regular checkups and checkouts for those sort of thing. So that is the first step. But when we take that data and use that data for improving our products um more so that we can tell doctors better what um you know detect it easier, detect it more precisely and so on so forth. It turns out 80% of our data is actually in play um that we have on prem today.

>> That's a lot.

>> It is well it's very weird relative to what most people have. Most people have like 5 10% of their data is actually hot and the rest of it is cold or cool or at least at most warm.

>> 80% we found in the last 6 months has been accessed and used in the various analyses that our scientists use and if I try to push that data off into archive they pull it back down again.

>> Interesting. So you have personally put a lot of effort and garden has put a lot of resources into a global namespace initiative. maybe you could tell us what is that, why are you tackling it, like what's the pain point and maybe talk a little bit about the initiative itself.

>> So there are kind of three problems our scientists have. One, they need all the data. They need all that data when they need it, not when it's available when you download it from say AWS because if something's kept most cheaply in Deep Glacier, if you it's going to take 48 hours or more for it to transfer down and into the system. Um, so they don't ha a lot of times like they have a request from the FDA. We need to have this data in the next 24 hours. We need this in the next we have oh wait we found something interesting and we have a submission that's coming up. We need this other samples that from that are out there as well. So that was our first pain point. The second pain point is we are diversifying not just to onrem HPC we're also you starting to use the clouds and more than just AWS but others as well and the scientists moving data management is not what they're great at they like to produce lots of data they like to hold on to that data but sorting it putting it archiving it taking care of it um is not something they're really good at so then they get frustrated as Well, because they also, oh wait, now we have a new environment. We have to rewrite our stuff to where the data actually lives. And so, you know, they're very very stuck. The by doing the global namespace that we have planned um and have been testing and going forward with. We want to have it such that wherever they are, the same data is available everywhere and it looks the same wherever that particular environment is. So the same file tree will they can search it in AWS on an instance there as you can on our HPC cluster in the R&D environment as you can in prod or in um often GCP or any other uh particular cloud like maybe Azure in the future that we're using. So our intent here is that we make them stop worrying about where the data is and archiving it because we'll take that care of it for them.

>> And are you talking just about the HPC researchers or is this other types of users of data as well?

>> So today we're working with the HPC users. One of the the plans is to share our data with other entities in a curated way >> um so that they can take advantage and use it for their own um research and whatnot. As the discussion this morning, data is currency. We happen to have one of the largest, if not the largest repository of gen cancer genomics data in the world.

>> So very valuable for both human health as well as financially.

>> Exactly. And for being able to find new ways of figuring out health issues. It may be for example some of our patients have other conditions as well. And if we can tag we know that and we can tag that then other types of research could benefit from this as well. So that's excellent thought leadership as far as traditional HPC architectures that were designed a little bit differently but made some of these other challenges difficult and so adding in this name global namespace solves a lot of that. That's great thought leadership. I know you're working to drive this into some standards to make it a little bit more accessible too.

>> Yes. So one of the issues is that today we're producing well over a pabyte a week.

>> Wow. This means that compounded with the problem that this is all hot and then made worse by the fact I have to retain it for 30 years.

>> So in the near future it's a high potentiality. I'm going to be talking about 21 exabytes of data.

>> Well that's great but and a challenge that we all love. But on the other hand what if one of my vendors decides they're not going to do this product I'm required I need to do this anymore. Right? What if I need to have a second replica on a different vendor for um in order to make backend vendor um in order to make sure that if there's an outage like say what certain cloud providers have had recently >> all very real situations >> all real situations um that I am not impacting my cancer patients because in many cases they are very much this is a life or death for the advanced oncology patients. Yeah, they don't want to wait and they certainly don't want data.

>> It's not a case of wanting to wait. They can't wait. They're literally >> in some cases days away from passing if we don't get our data the information to them.

>> That's a burden to carry.

>> It's a burden and it's honor.

>> Yeah, exactly. So, you had an announcement this week that you're broadening your reach a bit. Tell us just a little bit about your announcement. Um the intent with this whole uh announcement and this all this work we've been doing is to build a set of standards so that if I with the vendors I have today we use uh GPFS with Seagate um in order to Seagate live in order to do this namespace. if we don't um if Seagate decides they don't want to do object storage anymore or IBM because everything is very sensitive to price if their pricing structure changes um that we can then say okay we have other alternatives in the future and so this set of standards we've been doing is they'll allow us to ingest the data written by one um and be useful and maybe not as quickly or using all the secret sauce that another vendor has allow us to ingest it and allow it to run um and do continuity of business.

>> So you have a pretty good path at that point.

>> Very good path. And one of the things I want to highlight is this group we've been doing. There's about 34 if I recall entities that participated whether they're national labs like Livermore or there's the various founding members from the storage side such as uh as I mentioned IBM but also DDN WA and Hammerspace. Hammer Space has been highly helpful in participating and pointing out some of the things that have to be tackled in all of this.

>> Excellent. Well, William, I know you have a busy week. Um, the show has everybody back toback. Really appreciate you stopping by to take time with us. If people want to learn more about this initiative, do we point them to the Garden Health website or is there somewhere else?

>> Uh, we do have a GitHub.

>> Okay.

>> And you can find a single namespace and we have the presentations from and the minutes from all our meetings. We're very open about what we're doing and we are transitioning this over to Oasis, the standards body.

>> Sounds good.

>> And that standards body will be meeting in January. The call for participation should be going out momentarily and it should be very much a case that you guys, anyone who who's an Oasis member can help shape the standard. What we produced so far is a document that's meant to be a seed that because a lot of the different folks got into the room and had the the food fight already to sort out what's what. This allows us to save time to actually moving into a proper stamp.

>> Okay. So, great way for you all to go out and do some of your own research about the initiative. And then I think William, you have a meeting every 3 or four months. Is that approximately the cadence?

>> Today is the LA the capstone meeting. Right. So, we're going to that particular working group will won't be moving forward. But today is the 18th um at 2 p.m. at the National Blues Museum. Anyone who wants to come and participate is more than welcome to.

>> Excellent. Thank you so much for joining the show.

>> Thank you for having me.

>> Thanks for listening to Data Unchained powered by Hammerspace. To learn more, visit hammerspace.com. If you have a guest you would like to hear on the show, email me at molly@amspace.com.