Transcription
So today we're going to be taking that fresh look at focus zoom. Just to talk about some of the less known interactive features. And before I begin, this is a tool with a fair bit of history, which means it's just here we go, sorry about that. Quite a lot of people to thank for their work over the years. I don't think we could fit them all in one slide. Here is just a few names. To collaboration between both, um, and our partners at the Brera. A lot of great people helped make this possible. Solo Kazoom itself may look familiar from the complications. It's a tool that's been around for more than 10 years. It's widely established as a way to show jiwa summary statistics in biological context of information in that region or a geo a summary statistic region. Plus, recently we've been working on a new interactive version so that you're no longer limited in data sharing to the, you know, handful of static snapshots that you could publish in the paper. And this is an interactive version, an embeddable focus zoom jas widget. The key to making this pot is not just the visualization, but the data behind it. So we put some effort into bringing together information like the recombination rate track from the HapMap project, the, let's just make sure this, recombination rate from the HapMap project, genes in the region overlaid with genetic coordinates from Genco, and disequilibrium from the Thousand Genomes Project and a selection of possible LD populations. So together, these things can be overlaid in a single region plot and provide biological context for your results. Because this is a web-based tool, these region plots are interactive. You can drag to pan, you can scroll to zoom in the region, and if you click on a point, you can see tooltips with more information. These can include, as you see on the right side, links to external databases or pages, as well as interactive buttons that toggle specific features like changing the only reference variant in order to recover the plot appropriately.
The nice thing about the portal then is it's not just showing a single region plot. It brings together an enormous range of different data sets. So you can interactively select through the portal's drop-down menus and stack panels vertically to make custom pet comparisons between different phenotypes or different analyses. Within today's, the last numbers I've seen suggest there are 80 data sets and 189 traits to choose from, many possible combinations of options. In addition to summary statistics, the portal also provides an interval track plot. So you can see chromatin state from the Roadmap Project and pronation mem analysis. You can see chromatin state either for specifically for a single barrage or just across the region, as well as the taxi tracks and other things. And just like the data sense, you can interactively add these other annotations to choose which ones are most suitable for your research questions. So as we've added more and more data sets, the challenge becomes how to summarize them. And for this, a genome-wide association study property loss becomes useful. This represents all the archive al use for one single variant across a wide range of phenotypes. And using some clever method development from colleagues, there is a bottom line analysis option. You can choose to show most eval use, you can show information from many data sets, or you can run this bottom line analysis that corrects the sample overlap to provide the best available information for each given phenotype. You can also compare that datasets in the portal to a few us of a large data sets like the UK Biobank, and you can put that in context as well. There are multiple ways to visualize this data available. And again, you can see this on me single variant information page. This is the forest pot that shows things in terms of effect size, confidence intervals, and negative log ten feet across all the bottom line analysis phenotypes available. And we've been really happy about how many features the portal has been able to bring to you, bring in with us, and we're continuing to add more new features. So I thought I'd talk about either pending work or things in progress or some things that the portal may already have integrated, like various adjusted in ways that you might not recognize. But these are local scene methods that can interact with local Sam in the future as well.
The first most frequent feature requested ASHG was user selectable LD population. So now you can choose one of six different options sourced from the Thousand Genomes. This is powered by some great work from our colleague Daniel year with the new Michigan LD server and an IU three usable piece of software. So now you can click on this and dynamically recover the plot to choose the options most suited to your data or analysis. On that will be coming soon in Locust Zoom version 0.10. We also want to add more biological context from other known information. So a new feature allows you to see individual hits in your study on compare them to the known claimed significant results from the EBI G wasp catalog. One thing that's not clear in the static snapshot is if you click on one of those tick marks on the top track, it will now highlight it points on the other panels in a bright color that for that, um, step. This allows you to really bridge connections very nicely between different data sets and make more clear which snips are the same one rather than just lining up an imaginary vertical line, hoping for the best. So the new G wasp catalog feature allows you to see how the significant findings in your study relate to claims of biological significance from all other. And for the next section then, we've really been working it not just comparing to known research, but adding connections, adding new information on the fly. So for that, under turned over to my colleague Ryan, who's going to tell you about some of our interactive analysis features.
So hopefully you guys can hear me. If you can't, just say something in chat. I'll see it. I'm going to talk real briefly about some of the statistical work that we've done for Locust Zoom in the portal. The first one here is this idea of calculating credible sets. So, you know, the goal would be we want to show which variants in the region would be in the I 5% incredible set. And so we sort of implemented this existing method from Malar at all 2012, the links there at the bottom, that lets us calculate this posterior probability of inclusion in the credible set just from the p-values that are available in the region. So we don't need any extra information, we can use what we have here. So we go through encountering these factors and have to wait the final posterior probability and then and then we can have Locust Zoom display that in the regions. Let's see if I can. So you can show which variants are in the credible set, yes/no, or you can color them by their contribution, they're basically their posterior probability. This is an example of a region where, you know, the signal is kind of spread out. So in this example here, you see, you know, the inclusion probabilities contained within only a few variants. And then in the next example, you know, you can see Wanda where we don't really know, there's quite a few variants that have a fairly even probability of being in the set. And I think there's one more example here at the end. Yeah, so here we have a very fine, no inclusion probability, they're so low. Consume, as you saw here, has all the interactive kind of features here, and you can change how you want to display. When you pan and zoom, the credible set calculation will happen, you know, instantly for the region you're looking at, and then, you know, color appropriately. And the last point here, this important was to make this kind of method when we were working on it be modular. So, you know, this library that we developed for doing this, you can plug it into Locust Zoom and Locust Zoom and take advantage of and show these things in the plot. You could also separately take that functionality and use it in other places. And in one sample of that on the portal, and what before, so for knowledge, we there was already existing infrastructure for showing this kind of thing, show incredible sense, you know, from pre-calculated results. And so they were able to take this and actually plug it in and use those, you know, routines to calculate credible set posterior probabilities for variants in a region and show them in table form, as you see here. So this is on the gene age and for everything with TTP. Yeah, so, so yeah, this is one of the methods we worked on for the portal.
The other thing that we were working on recently was aggregation test. So these are tests where you have a number of rare variants that you're not powered to detect association for, and you want to group them together in order to test them. So, kind of similarly, credible sets, we developed a library that can calculate these aggregation tests like and scat from summary statistics. And this library can be used and looking soon to display the results, and it can also be used in other places. And we'll show you an example of that on the portal. So here, if you're using a 10 Locust Zoom, we can have Locust Zoom showing you the results of the all the aggregation tests that were run in a region. So on the right area, you see, you know, one gene has been highlighted in red, that means at least one of the aggregation tests that we ran was significant. And an updates, you know, this table here that shows you the genes in the region and your mask, which is just a grouping of how those variants were grouped together, and shows you, you know, what test was run and what was the p-value for that for that test. And this can happen interactively as you get that, as you're panning and zooming and dragging around. Also, not showing here, we have another table or you can see single variant statistics and sort of look in see, is my test driven by a likely a single variant, or was it really a, you know, significantly an aggregation test where it's, it's a few rare variants that are all collectively contributing to the result? And so this all this functionality for calculating aggregation tests is in our library, it's called wear a metal about, yes, thank you. And you can also see so that some how that works and one of our look is doing demos here at the link below. Found the portal again, there's already, you know, existing infrastructure for actually computing or running burden tests, doing custom aggregation tests. And so they were able to take this library and plug it in there to allow people to try, you know, running different tests. So they have an ey is just a little part of it that's available, but there's a UI for selecting what variants do I want to put into my aggregation tests, a protein truncating, maybe they're ants, I don't, when we have certain minor allele frequency, and I will show you when variants it's going to use, and then run the run the test and show you the p-values and results. So the library that we built over on burden scat Beatty and scat o on summary statistics with score statistics and coherence matrices. And these are, you know, summary statistics that are easy to share and, you know, are pretty common and many meta-analysis kinds of studies. In the future, we're hoping also add into the same libraries conditional analysis of conditioning on many variant region and in the future possibly also allowing for meta-analysis. So, you know, if you had multiple sites contributing statistics, we could run the meta-analysis right there in the web browser. I've shown you that, show you the result. So this is also on the gene page. There's a high impact variance tab. If you click on that and see the custom aggregation test button, and you can, you can try winning all of these different tests.
One last piece here of new functionality that we've been working on is this ability for Locust Zoom in the browser to actually update other elements of the page as your viewing Locust Zoom. So you could be panning around in Locust Zoom and have a table on one page that updates with what you're seeing. So as an example here, maybe, you know, you have a table of p-values and effect sizes, credible set posterior probabilities, and as you move around, the table will automatically update based on your scene. And then you can download those statistics to your computer. And we've had this kind of functionality for a while to download things in particular, that loading SVG images. So you could download theirs, you hover over the plot, you can download an SVG of what you're seeing to capture both the plot and the data that's being displayed on the fly as well. That's that's actually Andy's gonna talk again now about how to plot in low resume with your own data. Thanks, friend. Also, we've been really happy with how many datasets Locust Zoom has been used with so far. And as you've seen, there's a lot of new features come in. Some of these we want to try out before sending to the portal, and some of them we just want other people to be able to use Locust Zoom with their own early stage research as well. So this is a modular and reusable I heard you can drop it into any webpage. It's built on standard technologies like D3, SVG, HTML, CSS, JavaScript. That means it will work on many different kinds of websites and many different use cases. It's a configuration driven library, which means you can build new kinds of visualizations with control over point size, shape, color. All the interactive features you've seen so far can be customized for your use case. We also have on the backend standard API servers that provide common data like genes for builds GRCh37 and build 38, which means you can use this on a wide range of data sets. This is a permissive open-source license, so it's used in both industry and academia across a number of different sites and in-house analysis tools. We always welcome questions or contributions on getting started. We also recognize that not everyone wants to build their own web server to visualize their Givas. And so we've been working on lowering that barrier to entry, or how much is needed to visualize data. Again, this is a test bed for new Locust Zoom features, as well as using some ideas inspired by our work portal. Um, the first of those tools, it's called local zoo, which is a way without uploading your data to quickly make region plots from Gavi summary statistics that live on your hard drive. This uses a lot of the stuff we've shown you in the last section, including credible sex, selectable LV population. You can connect to a fee wasp lock, the UK Biobank, you can export your results as a table for analysis. This is all done in the browser and supporting comparisons to multiple data sets. Sometimes those for very large data set you want a little bit of help interpreting. So if you upload our your data to my Locust Zoom or a new hosted upload service that was released today, HTE last month, we offer needing two topics or summarize your file. You can get information annotations like topless eye nearest genes, Manhattan plot, and summary some for your entire study. You can click on any of those points and get that same view that you just saw in Locust Zoom. Can really explore with a wide range of analysis tools and dig into your data. This new website allows you to share your own analysis on these trusted collaborators, or you can click a button and make that data set public to foster comparisons with other research published work. We have fifty or sixty, I want to say, sample studies up there right now to see how this works. And again, we're continuing to develop these tools. We actively solicit feedback and we really, we really try to build this with modularity and reduce and mod the nature of research is we don't know how people will use our tools and we're trying to support a broad range of possibilities. So here are some places on get up where our code is available again under permissive open-source licenses, uh, and with that, we're happy to take any questions you might have, suggestions, etc. Thank you for your time.
Okay, if you can have a few quick though, this chat already. So your honor Annabella, okay, from personal legend Charles, is it possible seamless to add community contributors plug-in track? We do have some mechanisms, yes. The software engineering answer is, look at this, it maintains a registry of features. So if you look at the library, you can load an extension. We actually have several extension modules in our repo, most notably credible sets, ah, that demonstrate the principle. You can create like pre-packaged custom analysis tracks and data sources, and then a JSON serializable configuration object that finds that plug-in and automatically uses it to render the plot. So there is extensibility built in, and we even have demonstrations on how to get started with it. Okay, the next question from Alisa, I can answer. Can these slides or resources be available in the portal or through email? Yes, we are recording the webinar. I will make the recording available, and we also convicted signs available and resources. I think most of them below consume resources that are not on the portal, is that? And you talked about are available through Michigan websites. All right, well, for those who have not taken a survey, I think encourage it. It, you know, because I think it, we really have some valuable work thinking and using online site as we revive it and relevant all year. We are the next webinar is scheduled for January, mid-January, so that'll be in the new year. When he 20, he will have a December release and somehow probably broke it down as a mention before on the software stack, and then they're talking about how you might interact with the portal, create guy. I think that is a nice area, the different community people, as well as futures that we have here, community. So look to that, and you can see the resources we mentioned. This will be posted to our resource page or video for you data at any time. I think with that, thank you for your attention.