Transcription
Dat nerds, welcome to this full course tutorial on Excel for data analytics. This is the course I wish I would have had when I first started as a data analyst. You're going to be working right alongside me as we master how to use a spreadsheet, starting with the basics of functions, charts, and tables, working our way up to our first portfolio project. We'll then shift gears into advanced features like pivot tables, power query, and power pivot, ultimately building our second and final project, analyzing real-world data.
Now, to master this tool, we're not going to go straight for 11 hours. Instead, we're going to break it down into 10- to 20-minute lessons. During this, we'll have exercises for you to learn while doing, not just watching, followed by practice problems to reinforce your newly learned skills.
Now, Excel is the most popular spreadsheet tool in the world. It's estimated to have over 1 billion users—that's one in eight people in the world—and for data nerds, it's one of the most popular skills for data analysts, coming only behind SQL. Oh, and the same can be said for business analysts in this tool's popularity. Truth be told, Excel was one of the only skills that I knew when I landed my first role in data analytics, but it was able to handle everything thrown at me. And so I've been cataloging, over the years, all of the most important features to perform data analytics, and I compiled it in this course.
This video is for absolute beginners. You don't need any analytic or spreadsheet experience. We'll be starting with the first half on the basic chapters, which will build up your knowledge on the fundamentals, with covering which versions of Excel you can use for the course along with installing it. Then we'll get you familiar with working around how to manipulate a spreadsheet. From there, we'll shift into practical exercises, analyzing data using formulas and functions and then visualizing it using common charts and statistical analysis. At the end of the basics chapters, we'll put your skills to the test to build an interactive dashboard to predict one salary based on job and location.
For the second half of the course, we're going to ramp up our learnings, diving into advanced analytical features, focusing on using pivot tables and add-ins to dive quickly into data insights. We'll learn power query to connect to a variety of data sets and perform ETL, or extract, transform, and load. Finally, we'll learn data modeling with power pivot and perform advanced calculations with the DAX language. By the end of the advanced chapters, we'll have built a full data analytics project analyzing the data science job market, which you'll be able to share—this and the previous project—in order to showcase your experience with analyzing data in Excel.
Now, I'm a big believer in open-sourcing education, so this course and all the content required to complete the course is completely free. I not only get you set up with Excel, but I also provide all the different Excel workbooks and sheets needed to complete this course. With this, you'll get access to the data sets needed to make those final projects and even how to share them.
Now, unfortunately, the AdSense revenue alone from this course isn't enough in order to support all the different costs associated with building this, so I have an option for those that want to support and help out. For those that purchase my supporter resources, you're—you're going to get access to a lot of features that are going to help speed up your learning, all provided through this custom dashboard to track your progress. You'll get guided practice problems to perform after each lesson that will not only provide the solution but also walk you through how to get it if you get stuck along the way. You'll have access to a community of others in order to jump in and comment and ask for help. Additionally, you'll be getting my step-by-step instructions that walk through each of the lessons as I perform it. And finally, when you complete the course, I'll email you a certificate of completion that you can upload to LinkedIn.
Now, one quick shout-out before we jump in, and that's to Kelly Adams. She helped me plan out a lot of the different lessons for this course, along with being the brains behind a lot of the different practice problems. And frankly, if I didn't have help, I probably couldn't have completed this course.
So before we go any further with what we need to and actually diving into this course, we need to first understand what is Excel and where the heck it came from. So in order to understand this, we need to go back—oh, a little too far back—ah, just right—ancient Babylon, when we used to trade livestock like it was crypto. Now, it's during this time that we started record-keeping, and we didn't have paper, so we used stone, and we partitioned it into rows and columns. During the time of the Romans, they began to perfect this even further with accounting. Eventually, we get some advancements in technology; we start getting this on paper. This is when the term spreadsheets gets the introduction. This maintained that familiar row and column format in order to catalog different things, spread across different sheets—spreadsheet. Fast forward to the 1900s, and we pack rooms full of underpaid people in order to maintain and keep track of all the different transactions on paper spreadsheets. With the advent of computers in the late '70s, we started to see our first spreadsheet softwares—VisiCalc and Lotus 1-2-3. Then our boy here decided to revolutionize the world—a little bit—okay, not with that, but with this: I'm Bill Gates, chairman of my Microsoft. In this video, you're going to see the future. Since its launch in 1985, it's been wreaking havoc in the spreadsheet software community, dominating market share and to continue to dominate over the years. Microsoft has added more and more features. It initially started out to where you'd only be using it for the cells of entering different formulas and forming quick calculations, along with getting different charts and analysis. Shortly thereafter, it was upgraded with pivot tables, and that's my secret weapon to quickly analyzing data, as I no longer have to remember which comes first—index and match.
Now, VBA, or Visual Basic for Applications, was included in the mid-90s, and it's a programming language in order for you to automate tasks in Microsoft applications. Now, we're not going to waste any time in this course learning VBA; frankly, I feel it's outdated. You should learn Python instead, and there's newer tools that actually automate the process of data analysis, like Power Query. This was first introduced as 2010 and then rebranded to Get & Transform and then rebranded again to Power Query—sort of similar to what Google does with renaming products. Anyway, this bad boy is like washing down a couple caffeine pills with a shot of espresso; it can ingest and clean so much data in the blink of an eye. Hardcore data nerds call this ETL, or extract, transform, and load. Power Pivot was also introduced during this time of Power Query, and it's like putting your spreadsheets on steroids. This allows us to perform data modeling on data sets greater than a million rows—greater than what Excel actually holds in the spreadsheets—and combined with the power of DAX, or Data Analysis Expressions, we can supercharge our calculations.
Fast forward to today, and there's been two other major features added to Excel: Co-pilot, which is basically ChatGPT inside of Microsoft Excel, and Python in Excel, which is basically Python inside of Excel. Anyway, Co-pilot is great—wait, that's a lie. So I do believe AI chatbots are great at helping us out when we get stuck, but I don't want you to rely on that to actually learn this technology of Excel. And for Python in Excel, you need to know—well, Python—if you don't know this yet, it's completely useless.
Now, with all these features, it can make it seem like Excel is overwhelming, which I completely get that, but when you focus on the basics and work from there, I think it makes a lot easier to learn it. It's also why this course is almost 11 hours long. All right, enough with the history lesson. Let's actually get into the course material and what you're going to need for this. Also, we're going to be going over what data set—or what data—we're going to be analyzing for the project for this, with the link provided below. You can navigate to this, which is the GitHub repo that has all the different folders and files needed to take the course. Now, don't understand if you're not familiar with GitHub; we're going to walk through this. This pane here basically outlines all the different folders that you have access to, and if I navigate into something like Resources, I can see I have a Data Sets folder, Images folder, and even a Problems folder. So for those that purchase the course practice problems, you have access to the problems inside of here, and they're broken down by chapter along with the lesson. In addition to that Resources folder, you can see numbered here; we have each of those eight chapters. And if we navigate into something like Spreadsheets Intro, we have a workbook for each one of the lessons. So you want to download this file; you just navigate to it, click the three dots, and click Download. But I have an alternate method coming up in a bit. Inside the workbooks, I provide a blank template for you to go through and actually fill in, and we'll be getting to what's in this final sheet of actually being filled in. Now, as we move into the advanced chapters, they're going to have something like the Data sheet, or you're going to use the data from the Data sheets in order to do different operations, and we'll put those in different sheets as well.
So how do we get these files? Well, the easiest way is to come up here to this code and go to Download ZIP. With the file downloaded, all you need is to unzip it, and then from there it has all the different folders with the appropriate workbooks inside of them. Now, after going through a lesson, I then have practice problems for those that purchase the course perks to go through. Here's the course dashboard that you'll get access to that breaks it all down for the problems based on the chapter itself and then by the lesson. And inside of each of these lessons is multiple—multiple different problems for you to go through and work. The other perk that you'll receive with those practice problems are the course notes. These break down the concepts in a similar format of all the different chapters and lessons. Here's the one on Excel Install, which is going to be what we're covering next, but it provides all the different background on all the different material that we'll be covering this, and it's in the same format that I'm covering it in the video, so you can follow right along. Just as a reminder, there's no requirement to purchase these practice problems or course notes; it just helps support me.
Anyway, what are we actually going to be covering in this data analysis that we're going to be doing inside of Excel? Well, you're going to be taking the role of a job seeker in exploring what are some of the top-paying roles along with skills of data nerds. For this, we're going to use the data from my app, datnerd.tech, that is collected to this point—up to 3 million jobs—it tells, based on a job title and also on a location, what are the top skills, and it not only tells us the salary of these skills for a particular job but also the salaries of the jobs themselves. Now, the main data set we're going to be using for the majority of this course is this one here inside the Data Sets folder—Data Job Salary. All this data set includes over 30,000 job postings from 2023, and it includes a wealth of information such as company name, salary, and location. As we go through these examples, I'm going to be doing it from the perspective of a data analyst, which is their top job in the data set, but as shown here, there's a lot of different other job titles that you can check out and use as well, so feel free to deviate. Additionally, I'll be primarily focusing on the United States, but there's a lot of different countries in there as well, so feel free to plug in your home country and analyze this instead.
Now, with any course, you're probably going to get stuck along the way, and so how do you get help for this? Well, I don't recommend just jumping into the comment section and waiting for somebody to help you out. Instead, I recommend using a chatbot like ChatGPT. In it, you can provide whatever error you're seeing, and it will help you out and guide you along the way on what to do. And there's other great options as well, such as Gemini or even Claude, so feel free to use whichever one you're most comfortable with.
All right, if you haven't done so already, it's your turn now to go in and download that GitHub repo with all the different workbooks needed for this course. In the next lesson, we're going to be getting into installing Excel and mainly understanding what are the different versions that you can actually get with Excel and which one you need for the course. With it, I'll see you there.
Let's now actually get into working with Excel. So in this lesson, we're going to be going through how to actually install Excel onto your computer, assuming you don't have it. But before we get to that, for those that maybe have Excel or an older version of Excel or have different computers, we're going to actually go through what are the preliminary requirements you need to have or set up in order to be able to have the Excel you need for this course.
Now, here's a breakdown of the different chapters within this course—that is the rows here—and then for the columns are the different Microsoft products that you can get in order to have Excel. Now, if you're running Excel on a Windows machine, either through Microsoft 365, Microsoft Office Home and Student, or even an older version of Excel up to about 2010, you're going to be fine with completing all the different course content. However, if you have the Mac version or Mac operating system and Excel is installed directly on that operating system, you're not going to be able to complete the advanced chapter, specifically on Power Query and on Power Pivot, along with the project. And it's similar as well for Microsoft 365 online, as you won't also be able to complete the advanced data analysis section. Now, if you have any of these first three versions of Excel installed on your computer, you can skip to the next lesson if you want; I'm just going to be going through, before the install process, of breaking down each of these different versions so you understand your options—what you can get.
So let's get into breaking down all these different versions available. First up is Microsoft 365. Now, with Microsoft 365, you're going to get a host of different Microsoft applications—not only Excel but also things like Word, PowerPoint, and even Outlook. And there's two major plans I'm going to recommend for this: either the Family plan, which allows you to give out these keys for these different services to up to six people, or a Personal plan, which allows you to give it to—well, yourself. Now, I do want to call out that if you're a college student or maybe you work for a big corporation, you may have access to a free Microsoft 365 plan. So if you're in college, check with your college, and if you're working for a business, check for your business if you have access to this, so you don't have to pay money for it. But regardless of that, if money is an issue, Microsoft 365 Family offers this free one-month trial, which I think you can complete this course within a month, so technically you could do this for free. If you don't want to get charged, you will need to actually cancel before the end of that 30 days, and at that point you'll still have Microsoft Excel installed on your computer; just everything will be in view-only mode; you won't actually be able to edit any of the different spreadsheets that we've operated on during this course.
Let's now move into Microsoft Office Home and Student. Now, this bad boy is the alternate recommendation I'm going to give you if you don't want to pay for a Microsoft 365 subscription. This is only a one-time purchase, and it gives you keys to Microsoft Office, so you can install all the different Microsoft products—of Excel, Word, and PowerPoint—onto your computer for the low, low price of $150. Similar to Microsoft 365 subscription, this will not only work on a Windows machine but it will also work on a Mac machine.
Let's now move to this last option because it's sort of in the bundle of it—of Microsoft 365 online. Now, this version of Microsoft 365 is completely free, but sort of a catch to this. Here I am on my web browser, logged into Microsoft 365 online, and I have access to all the different apps within the browser, including something like Excel, so we can go to it. Now, this version looks very similar to the version that you can actually install the applications on your Windows or Mac machine. There are limitations, like I discussed before, about Power Query and Power Pivot, so you're going to be limited if you're trying to follow along in this course when we get to those advanced chapters. Also, the layout on the web browser version of this app is much different from that that's installing your computer, so I'm not going to be providing any support on this course on actually—actually how to navigate this, so you're going to have to figure that out yourself.
So we've discussed everything except for these Mac versions of Microsoft 365 and Office. So here's a quick recap of all the different features and cost of the three major versions of Microsoft that you can get in order to get Excel on your computer. For this, personally, I'm using the Microsoft 365 Family plan because it includes all the different features that I need, and it also—I save cost because I'm splitting with my brother, who—now that I think of it—is actually paying for it, but it provides everything that I need, and so it's the one I'm recommending for this course.
Now, before we get into the install, I want to briefly show what are the differences between using Mac with Excel installed vice Windows and Excel installed on it. Anyway, here's Excel installed on my Windows operating system, and Excel on this operating system is, in my opinion, the flagship product from Microsoft, so they're investing all of their effort and resources into designing this application to make it the best possible, and then from there, Excel online and then Excel for Mac are really just copycats of this. Anyway, the two main differences and the problems I've run into in the past that Excel for Mac doesn't have are in this Data tab. I have a lot of different data sources I can choose from, and that's specifically related to our Power Query lesson. And then finally, it has Power Pivot, which is just completely non-existent on Excel for Mac.
Now, here I am on a Mac machine, and we can see that it looks very similar to before, but there's a lot of limitations that we're going to find with this, specifically going back to that Power Query—not a lot of different sources you can choose from—and then, yeah, Power Pivot is just completely non-existent. You may be like, "Luke, I have a Mac machine; what do I need to do in order to have the most premier version of Excel and use for this?" Well, for that, I recommend installing a virtual machine. And virtual machines, like Parallels shown here, allows you to host a different operating system on your Mac machine. This Windows example that I was showing earlier—if I actually expanded out, you can see in the background here I'm running this on a Mac machine, and I have full capabilities enabled to carry out and running Windows on this. Now, I've been paying for and using Parallels over the past 3 years, and I can tell you the support and the offers from it are perfectly fine, and I love using it. Now, personally, I'm using the Parallels Desktop Pro Edition, but you can get by with just using the standard edition. Now, they also have this one-time purchase that you could do, which is $129, but it doesn't get any further updates, and I really like how it actually updates and fixes any bugs that may run into. Now, the other reason why I like Parallels is because it has this Coherence mode. I have this blue little icon that I can click up at the top to go into Coherence mode, and then—wait for it—it allows me to access any of those windows inside of my Windows—virtual machine—inside of Mac. So here is Excel running right here inside my Mac, and this is not only limited to Microsoft Excel but also products like Power BI, which I'm using pretty frequently as a data analyst. I can also run this into Coherence mode. But enough about that. Now that they got that out of the way, let's actually get into installing Excel via—in your Windows.
Machine or on your Windows Virtual Machine. So the first thing we need to do is navigate over to microsoft.com. And I'm going to click up here to Microsoft 365. We're going to be going through setting up the free 30-day version. So I'm going to click this "Try for free" and from there start my one-month trial. It's going to ask me to sync my data; I'm assume you don't have it. I'm also going to assume you don't have an account, so we're going to create one. I'm going to put in my email address, and then from there create a password. After providing some personal information, you're going to need to verify your email with the code they send you.
Now, to be clear, this is the Microsoft 365 family plan, which, after that 1-month trial, it's going to be charging you at $99 every year. So if you're just one person and you're trying to switch to the personal plan after this, you'll need to do that at the end or near the end of those 30 days. From there, like any company, they're going to ask for some payment methods. I'm going to just go ahead with PayPal. PayPal's all set; go ahead and do more paperwork of adding billing address, and with that I can start trial and pay later.
So now that I'm logged in, I want to install the desktop app so it gives me access to—right here—it's going to go ahead and begin this. It's going to ask if I want to allow this app to make changes to your device; yeah, I trust them. So only took a few minutes, and all the different Microsoft 365 office apps were installed. So I just come down to the search bar down here, type in Excel, let's pop it open, make sure it's working. And in order to get started, you need to sign in, in order to verify that it's your subscription. So I put in my email and password, and already forgot my password. Now I'm resetting my password, and now I'm all set up. All right, and we got to agree to some lawyer talk of accepting licensing agreements. At this point, I'm pretty worn out of going through this process, so I'm just going to click through everything. I'm not going to send any optional data; personally, I don't like to do that. I don't want to personalize right now, and it looks like I'm finally done. All right, I'm into it. And now that we're into Excel, we can see up here it should have your name or your account that you're going into, and go in here into the blank workbook. All right, so that basically concludes this lesson on installing Excel.
I do want to show real quick how easy it is to actually cancel your membership, should you want to go about just getting the free version or the free 30-day trial, and you want to cancel it before any—if I go back to my account, I can go in here to manage subscriptions, and here I'm inside my Microsoft account, which tells me I'm subscribed to Microsoft 365 family. I can share it with up to zero to five people, and for that I just click on it, and I can copy a link and provide it to whoever I want to share it with. We're going to cancel it, so we can go to manage subscriptions right here, and all we got to do is click "Cancel subscriptions." It's going to have me confirm that I do want to cancel this family plan; makes me scroll all the way to the bottom after showing me all these different prices that I could get instead, and I'm going to say, yeah, I don't want my subscription. And as I'm filming this on August 27th, it basically says, hey, you still have access this for 30 days until September 26th, so still technically have access to it. So if you haven't done it already, it's your turn to now go and install Microsoft Excel, the one of the options that I've shown here. In the next chapter, we're going to get into a spreadsheets intro to get you familiar with how to actually use all the different functionality or graphical unit or interface GUI of Excel. With that, see you in the next one.
Welcome to this chapter on an intro to spreadsheets. And this chapter has three different lessons. In order to understand what we're covering in those three different lessons, we need to explore some vocabulary with it. So let's jump into Excel. For this lesson, we're going to be focusing on worksheets, and that is basically, as you can see, this tab here called "Sheet 1," that is how to manipulate these different cells within this worksheet, or also known as a sheet. In the next lesson, we're going to be going into workbooks. So workbooks basically captures either one sheet like this, "Sheet 1"—if I add another one, "Sheet 2"—so it encapsulates multiple different sheets within this program of Excel. And then finally, in the third lesson of this chapter, we're going to be moving into the ribbon, which is up here at the top and has a bunch of different functionality to extend into those spreadsheets, along with using this file tab up here that has a whole bunch of features within it as well.
Now, this chapter was designed for those that may not have experience with using Microsoft Excel before. So if you don't fall in that category, as in you've used Excel in your job and you're pretty familiar with all those different features I just shown you, can feel free to skip this chapter and then move into the next one on functions along with all those different practice problems. But if you're not comfortable with that, stick around; we're going to get into it. All right, so the first thing you need to do is open up that first Excel sheet in the files you should have downloaded from GitHub on onecore worksheets. Inside of here, I have an original sheet that allows you to actually go in and fill in everything we're going to be doing and manipulating during the course of this lesson. Then if you get lost along the way or want to peek ahead to see what we're actually going to do, you can actually scroll over here or select the final sheet to see that. Now I want to make this as big as possible for you to see, so I'm going to go ahead and close out this ribbon up here, and you can just do that by double-clicking on any one of these different items up here. And then from there, I also want to zoom in, so I'm going to come down here to the bottom right, and I'm going to just zoom in to about 200% and scroll on over.
Now, inside the spreadsheet, it has all these different cells, and it's organized in a manner where it has rows, and the rows are labeled with numbers 1, 2, 3 all the way down to about a million. And then we have the columns, and the columns are alphabetical, and they all go all the way to where they start duplicating, where they'll put another letter in front of the other, and it'll go all the way through XFD. So let's practice some data entry here. I have a table we're going to be filling in for this lesson; basically has all the different skills associated with it, and then I want you to actually go through—while we're going through this, and you don't have to provide the values I do, you can if you want—we're going to be filling it in based on our difficulty when we may have started it or level, and then filling out some other self-formulas as we go. So we're going to start first with Excel, and then the difficulty. So I'm going to select right here, and I can see which cell is selected because it's sort of highlighted here on this B and also two, but also right up here next to this formula bar—I just call that formula bar—we can see that we're calling out the name of B2. So anytime we reference any cells, it first references the column letter and then the row number. So in this case, I'm selected in C7, so I'm going to go ahead and give this a number; I'm going to say four for myself. As you notice, I—I just put it right in the box. Alternatively, I can also select the cell I want to go to and then come up here into the formula bar, press what I want. So I want five for Python and go from there. Whenever I press Enter, it then goes down to the next cell. So technically, I could just go through and enter this all in using my keyboard, and I don't have to click or move manipulate at all except to select the cell that I wanted.
So those were all numerical values. When we move into the skill—known on whether we know it or not—we want to put in whether it's known or not; we want to put true or false. This is known as a Boolean value. So typing in something like "true," I can see when I press Enter it actually updates to be all caps for this "TRUE," so it recognizes the data type of this as Boolean. Now, if you're taking this course, you probably don't know Excel, so we're going to put in "false" instead. Now, say I want to update the rest of these for "false," "false," I can—yeah—go through and actually type it up, or I can select this lower right-hand corner of cell C2, and now I can drag these values down, and it will autofill it in. Now, autofill's not just limited to Boolean values; let's say I had something like "Luke," I could put that here and just drag it down; it's going to fill in "Luke" all the way through here. A cool feature about Excel is, say I have something like one and then two, I could select both of these cells, and then when I drag it down, it's going to actually fill in three or four. Now, autofill can also throw you off, especially for dates. So let's say we're filling in when we're starting Excel, which is—we'll put in for the today's date—in my case, it's August 27, 2024. I'm going to go ahead, press Enter to save that, and it automatically updates to this formatting here in America. If in Europe, you may see the month in a different location. Anyway, if I select this and actually drag down, what you'll see is is it will do that autofill in, but it's not going to keep that same day per se; it assumes we want to increment by one day. Now, specifically with dates, if I want to change the format, I can actually come up here, and I'll expand out this home ribbon again, and right now it's recognizing that the number is of date, and for date I have a few different options. I can do short date, which is shown here, or even something like long date. I can also go even further, which we'll explore as we get further into this course, into this more number formats, and date actually has a whole bunch of other different options that we can choose from, but for right now we're just going to keep it this simple date format, and I'm going to click OK.
Now, assuming you haven't started any of these, I'm going to go ahead and actually just select all the different cells that I want, and if you were to press Delete, it's only going to delete that top cell, and that's sort of annoying because I want to delete all these different cells. Instead, what I'm going to do—if I'm on a Windows machine, I'm going to press Delete, or in my case I'm using a Mac Windows VM, I'm going to press function Delete, and it's going to delete all the different content. Right, I'm also going to go ahead while I'm here, delete all that different content down there; we don't need it. Now we're going to move on to level type of diet; we're going to put into this is text. So in the case of Excel, you're probably a beginner, so I'll put in "beginner," and then if I want to, I can go through and fill out different levels for each of these: so Python, Advanced, R, BI, Advanced, and so on for all these. Now, one thing to notice real quick is, for the date, it does specify in here under this home ribbon that it is a date, but all these other ones it just characterizes as general, which is perfectly fine. Now, for these other options down here, let's go ahead and say I wanted to put in "beginner" for all the rest of these. Can't necessarily drag and drop this, but what I can do is I can actually copy it; specifically, I could right-click the cell and come up here and copy it, but I don't recommend that. Also over here on the home menu, they have an option as well to copy or even cut something, so I can select something like copy as well, and it's going to put these marching ants, as they call it, around the cell to tell you that, hey, it's actually selected. And then if I wanted to paste it, I go ahead and select down here, and I could paste it down below. That's not what we want to do; I don't like going through and actually selecting all these different buttons; I want to minimize it as much as possible, and I want to use shortcuts. So in order to stop these marching ants, I can go ahead and press Escape, and I'll select the cell that I want to copy, and from there I'll press Ctrl+C, and that copies it, and then I can go ahead and paste it below by selecting the cell that I want and pressing Ctrl+V. Now you'll be noticing that when I'm going through this, I have these shortcuts appearing right here next to me on the screen, so you'll be able to follow along as well as I'm using these shortcuts. The other option is I could cut this, so I could press Ctrl+X and then paste it in here, Ctrl+V, but this is going to go ahead and take this value out of here; we don't want to necessarily do that, so I'll just copy this again, Ctrl+C, and then paste it right above here, Ctrl+V. Shortcuts are going to be a big timesaver, and we're going to be using them a lot throughout this course in order to save you time and having you to go back to your mouse in order to manipulate it and select the different cells. All right, so let's step this up a notch, and we're now going to get into using formulas. And formulas are denoted by whenever we go into a cell like difficulty here, which we want it to be on a 1 to 10 scale; we denote formulas by an equal sign. And in this case, we want the difficulty to be on a 10-point scale, basically transition from that 5-point scale, so we need to multiply it times two. So we could do something like 4 * 2, and I press Enter, and it's going to give me, as I expect, eight, but I actually don't recommend hardcoding values that are already inside of Excel here; specifically this four. So instead of this, I'm going to remove this, and I can either type in the cell coordinates of the cell, so I could type in B2, and as you notice it's highlighting—one, the B2 is blue, but then the cell B2 is highlighted in blue. Alternatively, I can have an equal sign here and just go over and actually select it as well. Whenever I press Enter, it's going to go ahead and say, a, it's four now. Now that I'm referencing that four, I want to say that this is 4 * 2. Pressing Enter, we have eight. Once again, we're going to use that power of autofill, so I can select that cell of F2 and now drag it down, and what's going to be pretty interesting about this is the two, as denoted in the formula bar and actually whenever I click into it as well, the two remains the same, but autofill automatically knows to adjust the formula or the cell coordinates for the next cell down based on how I did that autofill. Just to show this as well, I could say, hey, let's equal this to B6 right below it, and then if I were to drag this over, it's going—going to then put in C6, D6, E6, then F6. So pretty cool. I'm going to go ahead and delete this.
Now the last column we're going to be filling in is skill and level; we're also going to be using a formula for this, and we'll set this equal to this skill thing and also this level. So I'll start by putting in an equal sign, and then it's not on the screen right now, but I know it's in B—or sorry, A2, and I can see that selected by scrolling over here. Now, how am I going to get in that F2? Well, I can do an ampersand, now, and from there I'll put in F2, and it has this selected as well. Pressing Enter ended up in the wrong one; sorry about that; should have been E2, and now I have "Excel beginner," but there's no space in between there; this is sort of hard to read. So what I can do is actually manipulate this to include another ampersand, and then in between this I'm going to put quotes, and this is—hey, insert this text character in between it; specifically I want to have a space, then a dash, and then another space, and then press Enter. Now, if I tried to do this without the quote, if I just did this and press Enter, I'm going to get a typo in my formula; you have to actually put those quotes around to show that it's text, and it's trying to correct it for some minus sign; I don't really like how it's doing it. Oh my gosh, it's freaking out. Now, anyway, I put the quotes back in there; pressing Enter, boom, we have it. And like before, I'm going to just do autofill to fill all those in. So let's zoom out a little bit, cuz we're going to be now be working with ranges, which is a collection of cells. Now, if you notice whenever I select—in this case I'm selecting B2—it says B2 up the top, but if I go to select more of this, it will actually call out that five r or five rows by two c or two columns, and then when I let go, it just goes back to B2. Anyway, ranges are a selection of multiple of cells. So if I come over here to I1, put it in equal sign, and then if I want to say copy this entire range, I can go ahead and select this all, so it's saying it's A1:G6, so start the upper left-hand corner of A1 and the bottom right-hand corner of G6. Now this is pretty cool; there's a new feature of Excel of dynamic ranges; it's going to go ahead and fill this in; there's only one formula in here of that A1 through J6, but you see that has this shadow border around here that's showing that this dynamic range is now filling in for all these different things, and if we look at the formula bar, it's sort of grayed out here too, for only at the very beginning does it show that A1 and G6, and then you could manipulate it. So if I wanted to, I could change it to G5, and it would just go down a row. Now we're not limited to just that; we could in fact select an entire column. So in this case, I'll put an equal sign, and let's say I want to do the—the full column of column A right here; I can select up here A; it's going to select all the way down, and if we go over to the formula bar itself, we can see that it's saying A:A, that means all the contents of column A are going to be included in this, and from there it's putting a copy—putting all these different things, and then when there's not a value in it because it's a copy—similar to over here for these dates of zero—we're going to see zero in all these different values all the way down. Now, similarly, I can also do a copy of a row. So in this case, if I wanted to—or multiple rows—if I wanted to do rows five and six, I could press Enter. Going to get an error with this though, and that has to do with this Q column right here that we're copying and pasting here, so I'm going to go ahead and delete that real quick, get rid of it, and now we have that rows five and six duplicated below, along with that shadow around it and all there. Now, these ranges are going to save us a lot of time later, so I'm going to go ahead and delete this right now; I don't want any of that. As later on when we get into actually using functions within formulas, I can use something like the average function, put in a range in here, so it selects all of it, and then get the average of it in this case. Now, one last thing to note on this before we wrap up here on how to save this is, you may have noticed that
This date started over here; it is a number, and that's because that's how Excel stores dates within this spreadsheet right here. So if I actually click on it, go back up to home right now, it's storing it under the format of General right now. So if I were to make this into an actual date, we can see that it is, in fact, 8/27/2024. Now, just some fun little trivia: if I were to put in number one and transition it to a date, so coming up here and selecting date, that first date starts at January 1st, 1900, and then they move on the numbers from there. All right.
Last thing we need to do is now save the work that you just completed with this. You can do this multiple different ways. We can come up here to the top of your Excel workbook right here and click save. You can also, as shown, you can use Ctrl+S. Alternatively, you can come over here to the file menu and then come on down to save or save as, and then if you wanted to, you can specify the location where you actually want to save your file and save it there. Now you do have the option, which I highly recommend if you're working with real-world files, you want to actually save them to save this autosave feature. The one caveat to this is that your files have to be stored on OneDrive right now. With the plan that I have, I can store about one terabyte of files on there. So if you'd like to do that, feel free to transition your files there. I'm not going to, um, and I won't have autosave on for this, but for very important files, definitely do have autosave set up. All right.
For those that have purchased the practice problems and notes, you have some practice problems to go through and get even more familiar with manipulating cells inside of a spreadsheet. After that, we're going to be going into manipulating a workbook. With that, see you in the next one. All right. We're going to be continuing on with this spreadsheets intro, focusing now on workbooks. So previously we were focusing on worksheets, which are a sheet inside of a workbook. Now we're going to be focusing on manipulating and moving data between workbooks. Now, for this, I don't want you immediately jumping into that 2 or workbooks Excel file. This really just has all the answers in it; it doesn't have really what we need for it. Instead, we're going to be starting with a new notebook and instead importing in some data. So specifically, if we go into this folder of Zore resource into data sets, we have this one Excel file called Data job salary monthly. Now, this is similar to the data that we're going to be using for the remainder of the course. We're actually going to use another Excel sheet, but this one here is pretty neat because it's broken up by months into different sheets. So all the job postings for January are in this sheet called Jan, and so on for February, and so on for March. So what we're going to be doing in this lesson is moving; we want to just evaluate the January data, move that into a new workbook.
To get a new workbook as easy as possible, we're going to come over here to the file menu. I'm just going to go to new and click blank workbook. Now here I have that new notebook. Right now, it's titled Book2 because it hasn't been saved. Anyway, going back to that file menu just to show you, I have different options. I can get a new notebook, so we went into new and just selected a blank workbook. Also, we could use this Home tab and select a blank workbook. Based on that, also have a bunch of different tutorials you can check out. Also, we have this open tab right here, which allows you on the left-hand side to select a location like This PC or even browse different locations in your file system. But frankly, I'm using more often than not over here on the right-hand side, this right here where this shows a past history of Excel files I've worked with, so I can go through and actually select an Excel file pretty easily. We're going to explore more about this file menu more in a bit. Let's get moving some data first.
Now, before we get into copying this data into the new workbook itself, I want to actually just copy it within its own workbook. So if we notice some controls down here at the bottom, we have all the different sheets. If we want to add another sheet, which I want to copy it to, I'm just going to add this in right here, and I'm going to call this Jan copy. Press Enter, and that's a new sheet, and I added that by just double-clicking in there and then allowing it to add; addition. I can right-click it, and I can do things like rename it, and that will do the same thing. Now there's also some controls around here. You notice there's some arrows on right here, and what that does is just scrolls all the way over or incrementally over, so I can see all the different sheets. In this case, there's more sheets than I'd actually see in one view. Then we have the scroll bar over on the right-hand side. This is actually just controlling the scroll area within our new sheet of Jan copy. So previously, we saw how we can copy ranges using a formula. In this case, I'm entering equal to, and then I'm just going to select this range right here. Press Enter, and I can get it inserted in, and then actually looking at the formula, it's just =J1:P8, and this has its range right there. All right.
So I want to get the contents into this sheet, so I'm going to start by putting an equal sign, and then I'm; I go over to that Jan sheet, and when I go over here, you're going to notice that now next to that equal sign I have Jan, the name of the sheet, and an exclamation point. This is identifying the sheet, and I want all this different items. So as I go to select it all, you can see that it's updating in the formula bar. Right now, I have A1 through P2 selected, but I actually want to select everything in this sheet, and we're about at 3,000 rows, and right now I'm only about 500 of those. This is going to take forever, so I don't recommend necessarily doing this type of method to try to select all your data. So I'm going to go ahead and Escape out of this and go back to where we were at the Jan copy. Instead, once again, I'm going to press that equal sign, go back to that Jan sheet right up in the formula bar. Once again, I can see that it has the Jan and the exclamation point. I'm going to select A1 to start with, and I'm going to press the shortcut Ctrl+Shift, and then the right arrow key, and now all the top row is selected. From here, I'm going to continue to hold Ctrl+Shift and press Ctrl+Shift down, and it's going to select all the different arrows. So as we can see up here, A1 to P3103, scrolling down, we don't have any more data. Now all I have to do is press Enter, and I did this to basically show the nomenclature now. So now we're not only selecting a range, but we're also selecting a range from a different sheet, and this is how Excel does the nomenclature or the formula necessary to make this work. And once again, this is a dynamic range appearing inside of here, but we really want to put it inside of here into this new workbook. So what I'm going to do is I'm going to actually delete this sheet right here because we don't need this copy sheet in here. I don't want to actually manipulate my data at all. Going to right-click it and select delete. It's going to prompt me anytime that, hey, you're going to permanently delete a sheet; do you want to continue? Yeah, I want to continue. Now, once again, I'm going to go back to that original blank sheet that we have. I want to put it into here, so I'm actually going to name this one Jan, and then we'll call this one Formula; technically it was a formula, not a copy. I don't know why I did copy before. Anyway, back into A1, once again, I'll press that equal sign, and then going back to that other workbook, I will select it, the first cell in there, which is actually A1, and now we can see we have in the formula of the bar, which is actually the front of the bar, which is sort of strange in the other sheet that our other workbook that we work with, we have inside of brackets the Excel file name, the sheet that we're in, and then the actual cell range of A1. We have dollar signs around this; this locks the references of it, which we're going to go into more detail on, but the main thing to understand is this has A1 selector right now, but we want to select all this data. So that shortcut of Ctrl+Shift right, select all the different columns, and then Ctrl+Shift down. Okay, it's all selected. I'm going to go ahead and press Enter, and it's going to take me back to my original workbook that I was trying to work with. This now that was using formulas to copy this data. We're going to explore two more options. The second one is going to be somewhat familiar, using copy and paste. So I'm going to create this new sheet; I'm going to call it Jan copy and paste. From here, I'm going to go back to our original data that we have, and since we're at the bottom of the sheet, I'm just going to select the bottom right-hand corner, press Ctrl+Shift+Left. Now, if you noticed, it went and stopped at this blank cell right here, which isn't a big deal. I'll press it one more time; it'll go to the next cell over that actually has a value in it, and then once again, it's going to go all the way to the end of A3103. So basically, if there's any blanks while you're trying to do this, it's going to stop at those values there. Okay, and then from there, I'm going to press Ctrl+Shift+Up, and as we're saying, it's going to stop at every different blank cell along the way. This is going to take forever, unfortunately. I don't recommend you actually do that ever again. Instead, start up at the top left and do the Ctrl+Shift over to the right and then all the way down in order to select all the cells. Now, like we did before, we want to copy it. I could either use this up at the top in the Home ribbon right here; I could actually select copy or the shortcut, which I'm going to recommend, of Ctrl+C, and from there, going back into our new workbook, selecting cell A1 and then using Ctrl+V and pasting all this data in.
Now moving on to the third example, which is actually the one I recommend you do anytime you need to move sheets of data. Basically, in both of those previous approaches, you could go about missing getting data to move over, so I don't really recommend doing that. Instead, I would come down here to the Jan sheet, right-click it, and select Move or Copy. So we have this new window that pops up, and it has two books right now. It has this Excel sheet selected of Data job salary monthly. We don't want to move to that; we want to move to Book2. We also move to a new book, but Book2 is open; that's what we've been working in; that's what we're going to move to. Okay, we can see we have the different sheets that we've already made in there, and it says in this dialogue, this is where you want to put this before this sheet, and we want at the end, so we'll select Move to end. Now we don't want to take this sheet Jan out of here; we just want a copy of it, so we're going to select this Create a copy and then click OK. Now Jan has moved over here, but I do want to actually differentiate this, so I'm going to put Move or Copy. Now, in the next lesson, we're going to be exploring more about the ribbon, but we're going to be exploring now more about the file menu, or also known as backstage view. We've gone through this Home, New, and Open. We also have this here for Share. This is available for, well, if you're sharing it via OneDrive, this makes it super easy to share with your coworkers. We're not going to go into a lot of detail, but this is a great option if you're working in OneDrive and you want to actually collaborate with other coworkers; you can work on Excel files at the same time. Moving down to the list here, we also have Get Add-ins, and we're going to be actually looking at different add-ins we can use in the advanced chapters whenever we get to that, so we're working with some add-ins with that. Next up is Info, which has over here on the right-hand side some key metadata about our Excel file itself. Then if we want to get into actually protecting our workbook, which we're going to cover in a few chapters down the road, you can get into actually doing that. The only other thing that I find myself doing from time to time in this section is on Version History. Once again, this requires you to be using OneDrive for it, but you could go back and revert back into a previous version that you work with, so it's great for that. Now moving into Save or even Save As, since we haven't saved; yes, they're both the same right here. I'm going to go ahead and save this, but I don't want to save this on OneDrive Personal; I'm just going to save this on my desktop, so I'll come and select Desktop, and then I'll name this Two Workbooks and save it.
Now, beyond Save As, we also have things like Print, which I really don't find myself doing that too often; should be sending an electronic version. Export, if I wanted a PDF version of something, and then finally Close, as well. Same thing as this X up here, just an X out of it. And there's two more areas down here that I want to call out, and that's Account, and that allows you to actually see behind the scenes of what's going on with your Microsoft account. And this is generic to all the different Microsoft products that you have, so not just Microsoft Excel. As you can see from my information, I'm actually inside the Microsoft 365 Insider program, so I get a lot of access to Insider features, get to experiment with new stuff before any other people do. Anyway, this is where you want to come anytime you want to make sure that you have your Microsoft products up to date. I have automatic updates available, so even if I check to update now, it's going to tell me, hey, I'm up to date. The other thing to note on this is the different Office themes that you have on this. I'm actually going to change this right now to use system settings, which on my Mac I use dark theme, so it's going to go to that. Last two options are hiding down here behind More. I have Feedback, so if I wanted to give feedback to this product, I'd probably go to something like X or Twitter instead, and then finally Options. We'll be getting to Options later on in this, but this allows a very much more advanced features that we can actually go in and customize using this menu, especially whenever we get into add-ins; we're going to be doing that from here. All right.
So now you become an expert at how to manipulate different spreadsheets or sheets along with manipulating them between different workbooks. In the next lesson, we're going to be going into this ribbon up here and actually exploring everything a little bit further and getting a sneak peek into each one of these. For those that purchase the practice problems and course notes, you have some practice problems to go through now and experiment working with different workbooks. With that, see you in the next one where we get into the ribbon. See you there. All right. This final lesson of the spreadsheets intro, we're going to be getting into the ribbon inside of Excel and better understanding what are all the different tabs and what are the capabilities by doing some simple exercises. For this, we're going to continue to be analyzing that January data set that we worked from the last lesson, and we're going to actually get into actually performing some data analysis with it. So for this lesson, you can open and use that ribbon menu Excel file, which I have right here, and all the data that we're going to be working with are that January data is in this Data tab along with all the examples and all the different tabs, but I don't need this; I'm not going to work with this, so I'm going to close this out. Instead, I'm going to be working off where we left from last time in that Two Workbooks where we actually moved over that January data set. Now, quick disclaimer for any of these files that you're opening up: if you're noticing the security warning of automatic updates of links have been disabled, you can go ahead and just enable the content and then click right here on Do not ask me again for network files and select Yes, because I want to make it a trusted document. Now, if you're getting any of these areas that the file has been moved, renamed, or deleted, because mainly you have it in a different location of what I had it, here's actually the address of the file that I'm using. I open it up anyway; this is the actual address of where the file is. Anyway, you can come down here and select these three dots on the file in question and just select Change Source, go into Browse, and then from there inside the actual file itself, select where this is. So in this case, it's looking for that data set file with the Data job salary monthly. I'm going to select it, select OK, and then it's prompting me now that this linked workbook hasn't been refreshed; want to, and go ahead and refresh it, and it's going to update it. All right, close out of this. Now, anyway, that was all silly because I'm going to go ahead and delete this Formula one right here and also this copy and paste tab right here. We only want to keep the Move or Copy, which is the actual sheet that we moved over that has all the data for this lesson. Okay, I'm just going to rename the sheet Data. So let's dive into this Home tab, and this thing has a lot to do with formatting the text and how things appear within the spreadsheet. For example, I can select all these top rows right here, so basically A1 all the way to P1. I can change this font size to something like 12. For the fill color or the background color, I can change it to something like a light gray. Right now, it looks like it's already bold; I could turn it off or turn it back on. Inspecting all these different columns, I can see that some of it is hidden, especially here, this Date column. I can see inside of here; this is the actual value, but whenever we actually look at it from afar, like it, it has these ampersand signs. So double-clicking on the edge of that H column right here, it actually expands out and moves it where it needs to go. You can actually do this for all the columns by just selecting all of them and then double-clicking that last one, and then that expands it all the way. We can see that that last column is well super long, so it has all the different skills. Typically these titles up the top, I'd like to maintain centered, so that way I know that it's a title, but I could move it to either side also, so I can move it up or down if I wanted to, but we'll leave it right there in the center as well. Getting into the number formatting itself, I can actually go and select something like Job Post Date; it's going to select that whole column. If I wanted to, I can turn this into a date. So in our case, I want to do a short date. Now, other columns I would want to format are these Salary Year Average and also Salary Hour Average. So besides just clicking here, I can also just select that, hey, I want to use this as an accounting number format, and it's going to automatically put these decimal places at the end, two decimal places. Since we're in the 100,000s, I don't really care about, so I'm actually going to remove them by saying Decrease Decimal; I'm going to do that twice. Now, for something like Salary Hour Average, I'm going to also convert this to a currency, but for these, these may have
Two decimal places of values included in it, so I'm going to leave it now. So for the Styles and cells portion, we're going to be getting into this more, especially into conditional formatting in the spreadsheets Advanced chapter and chapter 4. So we'll save that for then.
The next thing I want to do is get into this editing, and this is a pretty powerful feature. We can actually sort and filter our data if we wanted to. So what I'm going to do is actually select all these cells from P1 all the way to A1 and then come in here inside of editing, select sord and filter, and apply this filter.
Let's actually get into filtering this data specifically. I'm wanting to investigate jobs or data analyst jobs in the United States and specifically full-time jobs. We're going to be looking at the salary data for this, so I want to filter it down for it. So I'm going to select here; I'm going to unclick select all and select data analyst. Now it's going to filter for all the different data analyst roles there; nothing else that's not there. Additionally, that job schedule type I want to be looking at full-time roles only. I don't want to include any other ones, so I'll select full-time. I want the country; I don't want to be skewed by any other countries. I live in the United States, so I'm going to then select United States. And then finally, I only want to look at the salary or the yearly salary data, so I can actually come over here to the salary rate and select here: I only want to look at the year data. Okay, so now this has everything in it that I want.
We're going to get to analyzing and visualizing this in a second. Before that, I want to talk about two other features: add-ins, which we talked about before on how you access to the file menu; you can get to add-in via this. And finally, analyze data, which in my opinion isn't that strong of a feature. This tab uses a little bit of artificial intelligence behind the scenes for you to investigate, so it'll actually provide you different visualizations that you could actually visualize out of your data and/or even you can go as far as asking a question about maybe you want to see, hey, the distribution of salary rate or something like that. All you have to do is come down here and then insert in the chart that you want to insert in. I'm going to close out of this. Now we can see that we've made this salary distribution um that we maybe want to visualize. Overall, though, I find that this analyzed data is pretty hit or miss, so I'm not using it very often.
Now the insert tab is where I spend the second most of my time after the Home tab. They conveniently put in the correct order. There's three major use cases that I'm using. Out of this, in chapter 4 on the advanced use of spreadsheets, we're going to be going into tables, and then in chapter 5 we're going to be going into pivot tables. But even closer to that, in chapter 3, we're going to be going all in depth on how to use these charts. But let's get a sneak peek into this specifically. Remember, we filtered this table down to data analyst jobs in the United States and specifically full-time roles. We want to visualize this salary year average column. So with column M selected, I come up here to recommend charts, and it's going to give me a visualization of some well-recommended charts. Now there's only four here. I can also select this other tab up here on all chart and actually try to see, hey, what would this look like maybe in a pie chart or a bar chart. Anyway, I want this in a histogram, which we're going to go into more detail on how to read this later. What all have to do is just come in here, double click it; it'll insert it in. Now notice how whenever this was created, we now have new tabs appear inside of here, specifically with this selected; we have this chart design and format tab. If I select off of it, those tabs disappear, and select it again, they reappear. This tab allows me to dive in and actually further customize these visualizations to how I want them to appear. I can even move them to, let's say, a new sheet, and I can title this something like histogram and then move it. The charts don't always necessarily appear just like that.
Let's actually do a deeper analysis to see what are the different job title short columns available. I want to clear all these different filters on here, so I'm going to come back up here with this one row selected, come into editing, sort in filter, and I'm going to say, hey, clear all the different filters. Now selecting column A, going into insert and into recommended charts, it's recommended this clustered bar chart, which is actually what I want to view, so double clicking on this, this provides me a breakdown of all the different counts of the different job titles within our data set, and we can see things like data scientist, engineer, and analyst are some of the highest amount of job postings in this data set. Now, unlike our histogram example, this actually provides this data in a pivot table, which we're going to be going into in the pivot table chapter, which allows me to further manipulate the data. So say I want to actually sort this, I could right-click the values right here and click, hey, sort smallest to largest, and then closing out this pivot table tab right here, I can actually see what is the highest amount of job compared to the lowest, which is cloud engineer.
Now there's remaining tabs we're going to be going, and hopefully rapid fire, in order to cover these as I find I'm using these less frequently than these other tabs that we previously talked about. The draw tab allows you to, well, draw on your spreadsheet, so I can just write on it if I wanted to, but I don't really find myself doing that except for maybe being I'm building dashboards. Besides that use case, it's pretty rare. If I want to undo this drawing right here, I can come up here and click undo, or I can select Ctrl Z, and it'll remove it. Page layout tab is great if you're having to print out any data for those co-workers that are living in the past and don't know how to accept things digitally. You can do everything from adjusting your page layout to adjusting the scale that you're actually viewing things. Now personally, I find myself more using these sheet options right here. So if I go to this job count tab right here, if I wanted to, I could turn off the grid lines on here, as you can can see it got white on the background. I really like that. Now if I wanted to make sure they had actual grid lines around my table, I come back to the Home tab and for here I can select borders, and from there I want to put all borders on there, so now I look like I have this table right here along with my graph, super fancy.
Next up is formulas. This is where you need to go if you can't remember a function that maybe you want to use. If it's a text function, you come in here, select something like text, you can scroll through and actually see even a description of the different functions that are available. So in this case, replace, it tells you, hey, replace this part of a text string with a different text string. Depending on what version of Excel you have, and the newer ones, you'll have this insert Python to insert Python functions, and then finally they have more advanced features with maintaining and updating and formatting your different formulas and functions, which we'll be diving to in the next chapter.
Now besides the home and insert tab, the data tab is the next tab that I find myself using all the time. In chapter 7, we'll be diving into Power Query, and we're going to be focusing heavily on this getting transform data and also queries and connections. And then in chapter 8, when we get to power pivot, we're going to be going into managing our data model with power pivot. In chapter 4, we're going to be going into this forecasting, and we're also going to be adding in some extra add-ins that are going to appear in this data tab. Now I sort of skipped over the data types and sort and filter because we've saw them on the Home tab; they're just conveniently located here in bigger format for you use also all right. This tab on review is probably the least likely for me to actually use. I can actually go through and check things like spelling and add comments or even protect my sheet. Besides that, I'm not finding I'm using that this often. View tab is similar to the review tab and that I'm using it a little bit more. You can change the format of how you actually want to view things, but mainly I'm finding myself using this the most of freeze panes. Let's say you see I'm scrolling down here, and I don't know what the job or what the headers are right here. So going over to this data tab, I can actually come in here to freeze panes and select freeze top row or even freeze First Column. So in this case, that top row actually stays up there, and I really like it like that. Now let's say I want to freeze both the top row and that First Column; there's not really a selection for that. So here's what you can do: you can come over here to freeze panes and select unfreeze panes and then select something like a cell like B2. That means I want everything above this and to the left of it to freeze. So now when I select freeze panes, this upper or top row is actually Frozen, and then the actual First Column is Frozen as well.
All right, final tab is help, and I'll be honest, I think this is pretty useless. If I get stuck with anything along the way, I'm finding myself navigating to something like chat GPT, and it's helping me a lot quicker than trying to navigate through this help box that it provides, and I'm already getting an error message with even accessing it, so you can see how often I even use it then.
Now we've been doing a lot of manual clicking with using the ribbon, and I think a good resource that goes with this is shortcuts. So if you come inside of the resources folder, we have an Excel file here called Excel shortcuts, and what this has in it is a list of all the different shortcuts that I find myself using anytime I'm inside of Excel. So it's worth having all of these—I'm not going to lie—committed to memory. It looks like a long list, but I'm telling you by the end of this you're going to have all of these basically committed to memory; they're going to be timesavers. Now although I shed on people that print out stuff, this would be something that I do recommend actually printing out and having next to you so that way you can reference really quickly while going through this course. All right. Now I know we moved fast through that, but we're really going to be diving into, as I called out during this lesson, all of these different tabs even more as we advance through all the different chapters. That was more of a sneak peek into what you're going to be exposed to coming up in this course. All right, for those that purchased the practice problems, you have some problems to go through and actually experiment more with the tabs. In the next chapter, we're going to be jumping into functions and also more specifically formulas in order to build them out and form data analysis on that data science job posting data set. With that, I'll see you in the next one.
All right, welcome to this chapter on formulas and functions. In this lesson, we're going to be focusing specifically on going a deep dive and understanding formulas. Then in all the follow-on lessons, this we're going to spend the majority of our time working on functions. For that, we'll be exploring the entire function library, focusing on the key functions within this library that I find that I'm using time and time again in data analytics.
So what are we going to be doing in this lesson? Well, we're going to be focusing on a fictitious data set. We're going to keep it small in order for us to get more familiar with operating with formulas and operating on this data set specifically. By the end of this, we're going to be able to input into this worksheet a number of years of experience or total salary and be able to see whether these jobs meet those conditions, specifically me, that I meet both of those conditions. So for this, you can follow along by opening that formulas intro workbook. In this workbook, we'll be staying in this data sheet right here. All the different answers when we get to the math operators, comparison operators, or cell referencing are shown via that sheet, but we'll just be sticking for data for now.
First, as math operators and as shown by this table here, you can use a variety of different symbols for to conduct different multiplication, subtraction, division operations that you want to do. So let's dive into testing some of these out. We're going to be filling in each of these columns that correlate with the associated job title as we go through this. So the first one's going to be experience; pretty simple, right? We talked about before, in order to reference another cell, we would use an equal sign, and then from there we can either type or select a cell. I'm going to recommend just typing it to make it go faster: C3. It's highlighted blue because that's the cell that's highlighted. Then we'll be using the autofill feature of this to fill in all the cells below, and we notice that it updates to here. This one's equal to C12, which correlates to this one right to the left of it. So let's calculate our total salary, and this is going to be taking our annual salary in column D and adding it to our bonus Max in column E. So we can do this by specifying D3 + E3, and from there, there, pressing enter. Once again, to autofill it, I select that cell that I want and drag it on down. Now if I want to calculate what is the rate of bonus or the bonus rate, that is going to be the bonus divided by that salary. So in this case, E3 / D3. Once again, going to use autofill, drag and drop it all the way down. Now for all these values, I don't like what it's formatted as right now. I'm actually going to change this to a percentage, and I want to see one decimal place, so I'll press this one to expand out one. Now anytime I do any type of mathematical operation in Excel, I always want to try to confirm it that it's correct; I did the operation correctly. So in the case of this bonus rate, I can do this by confirming what we got for total salary previously. So if we took that bonus rate is which we want to confirm, right? So we're going to take that and multiply it times our annual salary, right? So that should give us that bonus rate right there. Then if we wanted to, like we said, we want to confirm total salary right here, so I can just add in that we want to also add in that annual salary itself, and we do have that total salary right here to actually confirm what's going on. Dragging it down and doing an autofill, all these values look like they correlate to what it should be for total salary, so I feel we calculate a bonus rate correctly.
Now going back into the formula itself, you can see we have multiple operations in here. How do we know whether multiplication, addition, subtraction, what comes first? Well, really, if you know the order of operations, it really is the same here. Here the different operators listed in their order of Precedence: exponentiation comes first, multiplication, division are second, then addition and subtraction are third. It's then followed by concatenation, which we did in one of the previous lessons, followed by the comparison operators, which we're about to get to. So with that segue, here we are: comparison operators. For this, you probably are familiar with the first three; the last three are something that get a little bit more complicated whenever you have a greater than or equal to, less than or equal to, or in this case, a not equal to. So previously, I just sort of did a cursor check to make sure this confirmed total salary column equals this other total salary column, but imagine you have hundreds of thousands of rows; how can we actually compare this and find these values? Well, what we can do is we can say, hey, is G3 equal to I3? This looks a little bit confusing, right? You have two equal signs in there, but everything to the right of the equal sign it's basically a comparison, and from there it either ends up as a true or a false, and we can drag and autofile this in, and everything is true. Similarly, if we want to find something like is the bonus Max greater than the annual salary, we can do, hey, is bonus Max at E3 greater than that at D3, and the typical of any data science job, none of these really exceed that at all.
All right. Now that we're familiar with math operators and also comparison operators, let's dive deeper into cell referencing, and we've been doing this previously whenever we reference another cell like A2, but we're going to add a little twist to this. I'm going to go ahead and hide some of these columns that way we clear up the clutter. Going to hide column F by right-clicking it and selecting hide. Then I'm also going to select all the columns H through K and also hide them. Want everything to appear on the same sheet, so we're going to be referencing this table down here for this portion of the exercise, and this is potentially goals that you may have when you're trying to land a job. You may know how many years of experience or you should have, know how many years of experience you have along with a goal total salary that you want to achieve, and so we're going to be building out formulas with this in order to be able to find out which of these jobs actually meet our conditions of the expected years of experience and total salary for. So for this, we'll go with that I have five years of experience, then I'm looking at $90,000. The first we want to calculate in column L is whether it meets our experience. So for this, we'll say, hey, is C15 right here less than or equal to the value right here in our experience, and as expected, five is less than or equal to basically equal to 5; it's true. Now we're going to run a problem now when we try to autofill this. If I try to autofill this down, I'm getting this one is false, and then these all is true, but I would expect, especially this AI specialist at three, it would be false. And so let's actually inspect this. Well, as we can see from this, this is referencing well C23, which is way down here, but it's still referencing the correct C11 right here. The problem is we didn't really want this value up here, this C15, to actually change whenever we went to do the autofill down below it. So what we can do here is provide a fixed reference of that cell. In order to do this, we're going to insert those dollar signs that we saw previously before the column and then also the row. So in this case, I have C locked, and I have 15 locked. Now the formula itself doesn't change at all, but now when I drag and drop this down, all of these are updating correctly as expected. AI specialist is going to be false whenever I actually click on it to inspect it; it's still referencing that C15 C11. Next, we're going to move on to column M of seeing if it meets our salary requirements. So for this one, we'll be seeing, hey, is the salary or total salary in G3 greater than or equal to our total salary down here of 90,000. Now we already know we need to lock C16 of this 90,000 because we're going to be autofilling it down. I can manually type in the dollar signs, but a shortcut to this is just pressing F4. If you're on a Mac, you'll need to press function F4. Anyway, this locks this in, so now whenever I drag and drop this down, as expected, the only other one that's less than 90,000 is this data analyst rule right here. Now I want to play with this just a little bit more. So we talked about this right here, putting a dollar sign in front of the column and then a dollar sign some of the row is a fixed reference; they also have what is called a mixed reference. So I'm going to go ahead and put my cursor right there next to G3; I'm
To press F4, and it's going to do the absolute reference. But if I press it one more time, it's going to do a mixed reference. If you notice, there's only a dollar sign in front of the three; or if I press it again, there's only a dollar sign in front of the G. Now, technically, this is going to work, but fine, because we're going to now lock this G column for this, but it's going to allow the three to update. So I'm going to show you this now by actually dragging and dropping this down. And from there, inspecting that last cell contents, we can see that that G is locked as expected, but it moved down. Now, instead of locking just the column, we could also lock the rows. So I could also do change up C16 now instead and lock the rows of C16, cuz we're going to still stay in that C column right there. Pressing enter now, autofill, we don't have to just go down; we can also go up. So inspecting it, locking it didn't really change by only locking the row of 16.
So let's wrap this all up by actually def finding out which of these actually meet both of our conditions of 5 years and 990,000. Well, it turns out that behind the scenes, true is equal to 1, and 0 is equal to false. So if actually were to take this and add this true to this true right here, we should get two. Autofilling it all the way down, we have 2, 1, 2, 1. So basically, confirm that, hey, zero, yeah, false is zero, because 0 + 0 is 0. Now, I recommend instead we're going to be going through and doing L3 * M3, so that way anytime either one of these are true, they will return a one. And now, in order to get a true or false back on whether it meets both, we can select that N3 and see, hey, is it equal to one? Type over there, equal to one, and it evaluates to true. So now I'm going to go ahead and just hide these columns so we can actually see this a little bit better, but we can find values in here that meet our conditions of the 90,000 or 5 years. And let's say we're doing job searching and it lasts over a year, um, we have to change this to six. This will automatically update the formulas that we've used here, as shown here. So that's our intro to formulas, and for me, the hardest thing to wrap my head around when I was first tackling this was around absolute and mixed references. So we have some practice problems for those that purchased the course; practice problems in order to go through and test this out and understanding what happens whenever you lock the row or lock the column. All right. And after that, we'll next be diving into an intro into formulas, which I'll be covering for the remainder of this chapter. With that, see you in the next one.
For this lesson, we're going to be focusing on an intro into functions. Specifically, we're going to be going over all the different functions that we're going to be deep diving within this chapter itself, along with some common problems you may run into and errors and how to troubleshoot it. To do this, we'll be continuing on from that data set that we used in the last lesson. Specifically, we'll be calculating things like averages and counts and how many jobs actually meet our goals, and we'll be using functions for this. So you can continue working in that workbook that you had from last time or open this function intros workbook. In this function intros workbook, I've gone ahead and moved our job goals over here to that column RNs, and then added in this bottom portion right here for the averages and total counts. Really, you can do and manipulate as you want. So why use functions? Let's look at a couple quick examples on the importance of these things. Let's say we wanted to get the average of each one of these Columns of experience, annual salary, and bonus. Max. Previously, we know we can actually reference each one of these cells to calculate the average. We wanted to do that; we would have to actually add up all the values. So I have to go through, select C3, C4, all the way down to C12, and we would need to divide it by that total number of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10. In that case, we get the average. Also, that me count that 10 wasn't necessarily perfect, so I don't really recommend doing this, but anyway, nonetheless, we can actually do autofill to calculate the averages as the is as well as it automatically update the referencing correctly to it, but I don't recommend doing that. Instead, I recommend using functions. Specifically, we can use something like the average function. As soon as I start typing a function, a in this case, all the functions that have the a name pop up. If I wanted to—well, I do know I want average right here—I can select it. It provides a brief statement of what it's actually going to do, and then I can double-click it to insert it. Below here, it actually specifies what's going on with this function here, and specifically, to provides me to, hey, provide in these numbers. Now, I could select these number by number, as we can see that there's in Brackets here, this number two, that means it's an optional parameter, but instead, what we'll do is we'll just provide a range, providing it from C3 all the way to C12. In that case, I got 5.3, similar to above, and then dragging this over, we can get all the other values as well. As a quick example, also previously, we had made this sort of convoluted formula in order to calculate, calate whether we met both conditions of mean our experience and also our salary, which we're specified over here. Well, there's actually a formula for that, and it's called the AND formula, and what it takes for its arguments are logical values. So it can take a logical one for the first parameter. I can specify L3, and then for the second parameter, I can specify M3, and notice how this second parameter now highlights or becomes more bold as I put it in, so you can keep track of where you are in the formula. Any, I'm going to close the parenthesis, press enter, and it evaluates to True. Dragging it all down, these should match these other ones, and yeah, this is definitely something I'd use over these formulas that I've used before. So let's dive into this formula tab more and understand the capabilities that we're going to be carrying out the next lessons in this chapter. The most powerful of these, especially for those new to Excel, is this insert function. Anytime you're looking for a function and maybe can't, can't recall the name, and you're not sure what even starts with, you can put something in here. So say I wanted maybe the average, I can type in average, and then everything that basically calculates a different average off of it, even if they're closely related like this rank average, will pop up in here, along with a description below explaining it. If you've used a formula recently, you can come in here under recently used, and I frequently find myself just going back to this in order to select something I may have used recently. Now, in the next seven lessons, we're going to be diving into each one of these, all the way through it, from logical and text to look up and also math and trick. Now, one note, we won't be going into detail on this financial functions because I find they're sort of nuanced, but we will be going into all the different ones that I'm using on a daily basis as a data analyst that aren't specific to financial applications. So let's get into understanding the basics about formulas by calculating these different counts and especially counts around whether any of these jobs meet our goals. For this, I know I want to use a count function, so I'm going to go to this insert function. I'm going to type in count. Now, there's a bunch of different ones that pop up. Count itself just counts the number of cells in a Range that contain numbers; it has to have numbers in it. If I wanted to do something more around text, I would say, hey, count the number of cells in range that are not empty. I could do even do something conversely of counting the number of blank cells. For us, we want to actually do count. So as we showed before, I'm just going to come in, type count; it's going to prompt me that I need to at least put at minimum a value, and I want to count all these cells here. So using autofill to fill it over, um, we can see that all the different values are 10; nothing really spectacular here, but now let's get into a pretty unique use case of count. So in this scenario that I'm count trying to calculate in cell C16, I'm trying to find out how many jobs above here in these 10 right here, how many meet our goal of less than or equal to 5 years, and I want to count the number of these. So I know I want a type of count; I can go into insert function; I know it's here inside these different statistical functions. Specifically, I have these different counts right here, and I'm going to scroll over this COUNTIF right here, and it's going to provide me a description. It says, Hey, counts the number of cells within range that meet the given condition, and that's what we want to do; we want to meet a condition of a certain amount of experience. Now, it provides this box in order to help me input in these values. So for the range here, what I can do is specify, hey, I want to count inside of here if they meet a certain criteria, and just going back to that range right here, we can see that it already input all those different values into an array-like object. Okay, so the criteria right now is NX; I want to put something in here. I can also press this box, and it'll make it disappear, and I want to compare it to this experience, but I want it to be less than or equal to five. So I can press enter to accept it, but the problem is it's going to evaluate whether five is any of these columns here, and right now we see that there are two. I'm going to go ahead and close up so we can see this better. Right now, we can see that there's two fives in here; that's not what we, we want. We want to see everything that is less than or equal to 5. So instead, what we need to put in here is less than or equal to 5. Now, I'm going to press enter, and we're going to get an error. This is pretty common whenever you are manipulating different formulas and you have—in this case, I have this less than or equal to right here—so Excel is confused by this. What we need to do is actually put parentheses around this, which basically sort of makes it into a string or text if you will, but now it knows, hey, I want you to look for less than or equal to 5; I want you to evaluate this entire thing. Pressing enter, bam, we have six values here that are less than or equal to five. Now, similarly, I can drag this over because we want to also do this for experience, but I don't want to do less than or equal to five; I want to do greater than or equal to 90,000, and in this case, we have nine, cuz we only have, have one that's less than this. But as you find out on this course, I don't like hardcoding values into my formulas. In this case, I have five inside of here, but I'm already having five right here. What happens if I want to change this maybe to say something like three? Well, it's not going to actually update these values right here. So I'm going to go ahead and actually change that back to five, and we're going to make another formula that actually fixed this. So I want to drag these down, but we actually didn't lock either one of the these cells, and it will cause errors if we do. So I'll just select right next to it, press F4 next to C3; I'll do the same of F4, doing the same in this cell as well. All right. Now I'll take this and I'll drag this down. So now let's actually fix this to be more Dynamic. We don't want it to be less than or equal to this five right now. What we can do is that ampersand operator, and then from there, put in reference to S3, which contains our five. Pressing enter, bam, we got six. Same thing here; I can delete that 90,000, put in an ampersand, and then from there, we're going to be basically putting it to mashing it together with that 90,000, and it evaluates. Now, when we change this experience to say something like two, we can see that it actually updates appropriately to see that, oh, only one job meets this requirement. So pretty cool. I'm going to change that back to five. Now, frequently, you're going to run into errors with your formulas. Let's say I wanted to divide one by zero; not a good thing that we need to do. Anyway, I'm going to get this error. Right now, you can notice it because it has this green check on the upper left-hand corner, but also it starts with this hashtag, and it's saying, hey, you have a divide by zero error. I can even come down into here, and it tells me even more on this, provides help on this; or if I wanted to even ignore it. Now, in this sheet of this workbook, I have a bunch of different errors in here that you may run into from time to time again, and we're going to be running into these errors as we go through the rest of this chapter. So if you get stuck along the way while we're going through this, I feel like this is a good reference for you to maybe save somewhere in order to understand what is going on with the different errors you may encounter. Now, the biggest time saver I've found with any of these errors is using some sort of chatbot. Specifically, me, I'm going to go to something like ChatGPT or even Claude; they're going to be able to provide really quick help in understanding what an error is and what I need to do to fix it. All right. So now it's your turn to dive into and test out these intro into functions and play with them and experience some of the errors of your own. After that, we'll be diving into logical functions, a major type of function that you need to be aware of. With that, I'll see you in the next one.
Now that we have the basics down on formulas and also functions, we're going to be moving into one of the most important type of functions to know, logical ones. The most popular of these are an IF condition, basically looking at something and then providing a response based on it. So for this analysis, we're going to be jumping into our data science job salary data set, but we're only going to focus on the first 20 rows of it here and on the next few lessons as well, as I don't want to overwhelm you with the all the data just yet. Now, for the final results, we're going to be doing two major things. The first is determining within this list of jobs whether they meet our conditions of finding the job we want of a data analyst or business analyst and will Market not desired or Role desired. Additionally, we're going to do a common practice and analytics of bucketing, basically taking those salaries and depending on the amount value, putting it into a certain bucket. For us, we're going to be looking at whether they have salary data in this data set, or more specifically, if they are greater than our goal of 85,000. So why are these logical functions needed? Well, let's jump into that last data set real quick and simplify how we can actually use these as a quick example. Previously, in this P column, we were evaluating whether they met both of our conditions of experience or salary. We can use an IF statement in order to clarify this. So I can specifically call out with an IF statement, saying if it has the Logical test that we want to actually evaluate. So I'm going to put in P3 in this case, as it's going to return true or false, and then from there, the next value in there is value if true, which what do we want to return if it is true? Well, that our goal is met, and then if it's not met, we want to have, well, not met. Okay, and then this whenever we drag this down will provide not met or goal met depending on if this is true or false. And so that's the power of these IF statements in helping us actually provide this value. So that was just a quick example of IF; let's actually jump into some more examples so you get more familiar with how to use this. So here we are in this data set, and I don't need all the columns of this data set, so I'm just going to select the columns that I don't need. I'm going to select B through G, and then hide it. Additionally, I'm not going to need I or J, so I'll hide these as well. So our first goal is to identify whether these jobs meet our conditions of either a data analyst or a business analyst. We're going to start simple by just finding out which one is a data analyst first, and then which one is a business analyst and meets those conditions. So once again, we'll start with that IF condition, and for this, we're going to put in that logical test. Remember, pretty the example, we need to have a return either true or false. So we're wanting to check whether senior data engineer in A2 is equal to data analyst in K1. Now, we're going to be autofilling this down, so we need to make sure that the A2 we're fine with it actually adjusting as necessary; K1, we want it to lock, at least lock on the row value of one. Then, if it's true, we'll be Role desired, and if it's not, it's not desired. As expected, senior data engineer is not desired. Let's drag this all the way down, and just double-checking it, we see that the data analyst roles are Role desired. Okay, so I can drag this over now, and just to double check it, shifted over to B2, but it's still, but it's selecting a right one of L1. So actually, what I'm going to do is I'm going to delete this, go back up here; I'm sort of a perfectionist; I'm going to end up locking that A value so it stays in that A column; none my values are going to change here, and then when I actually drag it over, I can check that, okay, A2 is the correct one I once selected to compare it to business analyst in L1. Okay, then I'm going to autofill all the way down. Looks like there's only two business analyst roles here. So now, how can we identify that it meets both of those conditions, both data analyst and a business analyst? Well, we're going to do one approach first, and it's called a nested IF statement, and it's not really the approach I'm going to recommend, but it's something that you should be aware of. So what I'm going to do is I'm going to select cell K2; I'm going to go ahead and copy this formula, plugging it in here; we have it here, and making sure that it operates correctly; yep, it does. So how does this nested IF statement work? Well, we're going to still evaluate our first condition; is the first role evaluated as data analyst? Does it meet that? If it is, we want to mark it as Role desired. Now, we get into what happens if it's not a data analyst. Well, now we want to now check if it's a business analyst. So I'm going to close out this, and what we can do is I'm going to take this business analyst formula right here, everything up to the IF, and I'm going to go back in here, and I'm going to drop it in right here inside of the value if false. So it's an nested IF statement, an IF inside of another IF. So now, if we don't meet this first condition of the value IF isn't true, it will go into the nested IF statement and start checking this condition; is now the software data engineer equal to data analyst? If it is, it's Role desired; if not, it's not desired. So let's now drag and drop this all the way down. I'm going to expand this out a little bit, and now we can see if it's data analyst, we get Role desired, along now with if it's business analyst, also Role desired. But I'm not a fan of nested IFs as they're hard to read. Instead, I like using the functions of AND and OR, and should be a little bit familiar because we saw it from the intro lessons that we did previously.
With "and," it evaluates whether both conditions are true. So, in this case, I'll put in condition one of B3 and then condition two of C E3. Both conditions are true, so it satisfies as true. Dragging this all down, in all the following condition cases, they're not true for both conditions, so therefore it evaluates as false.
With "or," it checks whether condition one or condition two is true and then will return true. So, inputting in the conditions of B3 and C3, one of the conditions here is true—actually, both are. Dragging it down, I expect—yeah, the second and third rows are also true, where the final one, both are false, so therefore it is false.
So let's run the same AND/OR logic that we've run before in order to determine which one we actually use. So, in this one, we're checking whether both of these jobs of data analyst and business analyst are equal to this one here, senior data engineer. As expected, false. And what should we—we should expect for all of these, all of them are false because none of these are going to be both data analyst and business analyst.
So, as you can probably guess, "or" is probably going to be the one that's going to work for us. We're evaluating whether either data analyst or business analyst are going to match up to that value of senior data engineer. In this case, we're getting those TRUE values for data analyst and TRUE values for business analyst. So now we're going to put that "or" function inside of that IF for the logical test, and from there we can determine whether it's rule desired, not desired. Dragging all down, all of it's matching as expected. Okay, I'm going to go ahead and hide these rows.
So now, what happens if we don't want to just evaluate for a TRUE or FALSE condition? Basically, we want to evaluate for multiple different conditions. Well, that's going to be something that comes up if you need to ever bucket data, which we're going to be doing with salaries.
Now, for this first one, we're going to just use a simple IF statement. We want to determine whether a salary is greater than 85,000, or if it's not, we want to just specify that the salary is low. So, for this, we're going to be evaluating IF H2—we're going to go ahead and lock that H column—is greater than that 85,000, which will lock that completely for the 85,000. Then we want to say the salary is greater than 85,000. Conversely, if it doesn't meet this, we want to say that the salary is low. I'm going to expand this out a little bit, and then we're going to drag this down. As expected, we have the values returning those are 85,000, and then this one at 35,000, it is marked as low.
Now, the problem we're running into, and why we need multiple conditions, is this: the salary is low, but there's actually no data there. We need to specify in these conditions that, well, there's no data. So, for this, we can use an IFS formula. And what happens with this is you provide a test and then a value if true, and that's just the first one. We can then provide another logical test and the value if true. So, the first thing I'm going to test is if there is no value there. I'm going to go ahead and lock that H column as well, and when I'm looking for a blank, I'm just going to put in two quotation marks, signifying that it's blank, and the value if true is "no data." Okay, put another comma. We can see we now—we're on to logical test number two. The next thing we want to test is if it's greater than 85,000. So we'll see H2 again, locking that H, and we want to send it—if it's greater than or equal to that 85,000, which will lock—if it is, we want to return back that "salary is greater than 85k." And finally, we're on to the final logical test, and basically, we want all of them to pass this condition. So, instead of providing, "hey, salary less than 85,000," we're just going to pass in TRUE because we want it to be true, and we would expect this to be any values between a number that are between 0 and 85,000. So, like before, we're going to specify "salary low." Running this, we're going to expand this out and then drag this down. We have when it returns "no data," "no data," "salary less than 85,000," return "salary low," and then whenever it's greater than 85, the correct results.
Now, IFS functions are one of the more complex functions to work with, so you do need some practice with this. Like, for those that purchased the course, practice problems—you have some now to go into and actually try this out, manipulate, and better understand how to work with this. With that, in the next one, we're going to be jumping into my next favorite type of functions: math functions, which are heavily used in data analytics. All right, with that, I'll see you in the next one.
Now, in this lesson, we're going to be using math functions and also some statistical functions in order to perform EDA, or exploratory data analysis, on our job posting data set. And for this, we're going to be focusing on the five major functions of COUNT, SUM, AVERAGE, and also MIN and MAX. And we're not only going to focus on the core versions, such as just COUNT, but also the IF and IFS versions. So, they have multiple different versions that we're going to get to.
Now, for our analysis, we're going to be diving into the full data set of the data science job postings, which has over 30,000 different job postings, and in it, we're going to be specifically diving into data jobs that are in the United States for data analyst. And we're going to be able to use these sort of different functions that incorporate IF and IFS in order to fine-tune in what we're looking for. One quick note: you're not limited to using United States and data analyst. You can use the scenario that you're in, of what country you're in and what job title you're most interested in instead.
So, we're going to be filling out this table right here, and we're going to start on row three, focusing on those COUNT functions first. Now, the data set is actually much larger than this three columns. I actually—I'll unhide between A through K, but we're not using any of these columns in between here, so I'm just hiding them and making it easier for us to work with. For this, we're going to focus on the core function of only COUNT, and we're going to be looking at those that have all the yearly salary data in it. As you can see over here, that there's missing blanks in here, so we don't want to count those that are missing. Anyway, what I'm going to do here is select column M, and as you knew it, selects the range of M:M, and then from there, press Enter. So, what we're finding is that around 22,000 jobs out of these 30,000—we're going to find out—have salary data. And how do I know about that 30,000? Well, let's actually see. We can actually use—instead, we can use a COUNTA function, which stands for "count all," and it counts the number of cells in a range that are not empty. Specifically, I want to capture those in the "Job Title Short" column right here. So, I'll do A:A. Running this, we get to see that it's around 32,000 jobs.
One technical note before we continue: these are—since we're doing the columns themselves, in this case, the COUNT(M), it's also counting that column header in this case. So, if we want to be exactly accurate—which, in this case, I just need roundabout numbers—if we want to be exactly accurate, technically we would want to go in and say subtract one to get what the actual value is. But frankly, I'm just trying to look at general numbers right now; I'm not too concerned about one or two off. So now let's dive into analyzing this further on my needs, looking for specifically focusing on the United States first. So, we're going to find those that have in the "Job Country" here, "United States." And for this, we're going to use the COUNTIF function, and this counts the number of cells within a range that meets the given condition. So, you provide a range; in this case, we're going to provide the range of that column K, and then the criteria itself. We want to filter for "United States," which I conveniently typed above, so we'll select it right there. I'm also going to lock it by pressing F4, and then running this, we get that about 25,000 jobs contain "United States."
So now let's evaluate those data analyst jobs using that same thing of COUNTIF. Once again, we provide the range; in this case, we're looking at that "Job Title Short" column, and for this, we want to look for "data analyst." Locking this cell, we get about 9,600 jobs for data analyst.
Now, next up, we're going to be using COUNTIFS. Specifically, we're doing this because we want to find jobs that contain not only "data analyst" but also contain that they're from the United States. Now, we can't just add these two columns together because, one, it's going to—as we—once we add it up, we see that's even greater than all the jobs there; that's not what we actually want. We want conditions like here on row 16, where it's a data analyst and United States, whereas something here on row 223, where it's a data analyst in—that's not going to meet our condition, so we wouldn't count it. So, using COUNTIFS, this counts the number of cells specified by a given set of conditions or criteria. For this, we need to specify a range and then the criteria. First, we'll focus on the range of A for "Job Title Short," and we're looking to match that of "data analyst," which I'll lock by pressing F4. Then now we're moving on to criteria range number two, where for this one, we're looking at "Job Country" now, and for that, we want to look for the criteria of "United States," locking this with F4, closing this with parentheses, and then running it, we get around 8,000 jobs. And this makes sense, right, because it would be less than that 9,000 data analyst because some of these aren't going to be from the United States.
Now, with how this is flowing, we could actually make a visualization out of this data right here. So, going into Insert and then Recommended Charts, we have here a funnel chart. So, I'm going to go ahead and insert that in, and this basically shows the funnel, if you will, of jobs we have. We started with almost 32,000 jobs, and we got towards the end of the jobs that we actually care about—US data analyst—at around 8,000. I'll go ahead and move this off to the side for now.
All right, next, moving into the SUM function, and the core one itself of actual SUM itself, it's pretty simple. We have to just—we're going to obviously be using the "Salary Year Average" column for this because we want to sum up the numbers in them, and I'm putting in that column of M, and we get the sum of values there. Now, unlike COUNT, where a COUNT has a COUNTA or COUNTALL where we're trying to find if there's blanks or not, that's not really applicable in SUM and AVERAGE and also in MIN or MAX. So, I'm actually going to go ahead and just gray these out because we're not going to need them.
Now, moving into SUMIF, which adds the cells specified by a given condition or criteria, this one is a little bit more complex than we dealt with with COUNT because we first want to provide the range that we're going to be evaluating for a certain criteria, which in our case, the range we want to evaluate is "Job Country" because we're evaluating for if it contains "United States," which I'll lock with that four, but we're not summing the countries because there are text columns, so we have to provide this sum range, which is column M. Similarly, once again, we can do that SUMIF looking for "data analyst." So, in this case, we're going to be looking at column A to evaluate if it has "data analyst" in it, and then from there, the sum range once again is going to be that column M.
Now, the SUMIFS, similar to that COUNTIFS, adds the cells specified by a given set of conditions or criteria. For this one, we provide the sum range first, so it gets a little bit confusing; you've got to make sure that you're actually reading the formulas. In this case, we're going to use M because that's the sum range we want to use, and then we're first going to evaluate for that "Job Title Short," that column A, which we're going to evaluate for "data analyst," and then we'll evaluate for the "Job Country," evaluating for "United States," closing the parentheses, and then running this—bam—as expected, this value is less than that of the data analyst.
Now, moving into the last three of AVERAGE, MIN, and MAX, which I think are actually more valuable than that SUM one we did, I'm not going to walk through actually typing in all these in because now you've had a familiarity with how I did the SUM, which follows the same example for AVERAGE, MIN, and MAX. Feel free to—if you want to—you can go through and type it out on your own to get more experience doing it, but overall, I think this has some very unique insights from it—from this analysis we did in it. We can see that salaries in the United States are around 125,000, where the data analyst is only around 93, and specifically, US data analyst is around 94. So, data analysts in general are lower salaries than the other jobs in the data science industry. As far as MIN and MAX go, we're having as low as 25,000, but we're having as high as—well, at least for a data analyst—up to 650,000, and apparently there's a job in here around $960,000. And you may be wondering what jobs correlate to this $155,000 or $960,000. Well, we're going to be diving into that further when we get to those lookup functions.
One last note on errors before we go: I commonly find the most common error with these functions is a #VALUE! error, and that usually occurs whenever—in this case, we had column A selected initially for criteria range number one—let's say we accidentally selected multiple different columns for this. Obviously, we're not trying to evaluate all the different columns; we only want to evaluate one column for that criteria of if "data analyst" falls in it. Anyway, when I run this, I get a #VALUE! error. Anyway, this is a common one that I see come up time and time again. So, anytime you're going through this, any of these or the practice problems themselves, make sure you're investigating to see that you've actually input in the correct ranges to evaluate, because it's commonly causing those #VALUE! errors. All right, with that, you have some practice problems to dive into, and next we'll be diving into even more statistical functions in order to really dive into how deep you can go with EDA, or exploratory data analysis. All right, with that, I'll see you in the next one.
We're now going to be taking this up a notch, shifting gears from focusing on math functions now to statistical functions. For this, we're going to be using our job posting data set and analyzing the salaries in this, specifically looking at common statistical functions like MEDIAN, STANDARD DEVIATION, and even QUARTILES. Once we have the basics, we're going to shift into an actual analysis looking at what is the average salary of different job titles, and we'll even get a sneak peek of visualizing it. For this lesson, you can start by opening this statistical functions workbook. We're going to be starting by filling in this table here on the different statistical functions we're going to be filling out, and we're still working with that data set we did previously. If you noticed, I've hidden a lot of the columns that we won't be using for this. So, we've done a few of these different types of functions already. Let's go ahead and fill these in. For COUNT, we'll be using the COUNT function specifically on that M column of "Salary or Average," and like before, we have around 22,000 values. For AVERAGE, we'll be doing the same on that M column; we find that's around 123,000. For MIN, we'll also run this on the M column, and that's around 15,000. For MAX, that's going to be around 960,000.
So let's move on to our first true statistical function. We're actually going to go into this to actually see what it does, and that's MEDIAN. It returns the median, or the number in the middle of the set of given numbers. So let's go ahead and type that out: MEDIAN, and in there we need to specify number or numbers. We can specify a range; we're just going to keep it simple right now to actually show what this function is actually doing. It's selecting the middle of numbers. So I'm just going to select these top three numbers right now, and what I expect for this function to do is to provide basically—in a set of numbers given—provide the middle number. So it should provide us 140,000, which is the center number of these three. We don't care about the center of just three values; we care about the center of basically all of our different values. So I'm going to place the entire M column into it, and that is around 115,000.
Now, why is this average higher than this median? Well, let's actually visualize it. I'm going to select this M column and go to the Insert tab, going to Histograms. I'm going to insert a histogram, and what this is showing is the distribution of salaries from 15,000 all the way to 950,000. The bottom x-axis is a little confusing to read, but it's basically a range. So, in this case, 87,000–93,000—how many counts of salaries are falling in between that, and that's how large the bar is next to it. Anyway, getting back to that original question: why is the average higher than the median itself? If you call back from the definition, a median is the middle number in our set of our list, but our average, however, is taking all the different values and, well, averaging it out. And as we can see from it, we have a large amount of salaries around, well, $100,000, but we do have some up here that are getting close to a million dollars. These basically outliers are causing us to have a higher average. So, basically, those values that are near 960,000 are dragging that average way higher. So that's why I prefer to use something like the median when I can in order to analyze these salaries because they're not skewed by these outlier salaries that are just something that you're probably not going to get.
All right, next up is STANDARD DEVIATION, and for this, you have two options: STDEV.P and STDEV.S. The P stands for population, and the S stands for sample. This data set is around 30,000 salaries, and there's way more than 30,000 data science jobs available, so that's a sample of the actual population. So we're going to be using STDEV.S, and for this, we can insert a range into it. So, what does this value actually mean? Well, if we had something like a normal distribution, which our salary data is somewhat close to, we'll find that one standard deviation from something like the average has—in this case right here—34,000. So, if we went above and below the average by one standard deviation, around 68%, which is a heck of a lot of data, is within this one standard deviation. So, in our case, if I was to take the average and then subtract this standard deviation, along with taking the average and then adding the standard deviation, around 70% of the salaries are going to be between 75,000 and 170,000. But what if we wanted to be more precise about finding, say, something like where does 50% of the data actually fall? Well, we can use QUARTILES.
In this case, specifically calculating the first and third quartile, here's a graph that I did from my Python course, which when you get done with this course, feel free to check it out, but anyway, it looks at the salary distribution of data analyst, United States, has this histogram right here, very similar to what we plotted previously in Excel, but in it, I'm able to plot out quartile one where the quartile one starts and then quartile three where that one starts. So, between this quartile one and quartile three marker lines, 50% of the data falls here, with this red dotted line being the median again. Which—let's actually get to calculating this. So, if we want to do something like the QUARTILE, we're going to see that there's a few different functions available for this. We have exclusive and inclusive; we're going to do inclusive first, and then I'll show the exclusive after to basically show how it's different. So, this takes two arguments: the first is the array, so I'll put in that range of the M column, and then lastly, it takes the quartile, and we have one for the first quartile, two for the median, three for the third quartile.
Anyway, I have these values over in the U column, so I'll just select that and use that for this. And for the second quartile, we're seeing that basically, as a just red, that's also equal to the median. Now I'm going to go ahead and get rid of these Min and Max CU. We can also use that by with our quartile function. And I'm going to go ahead and drag and drop this up and then also below. So what we can see from this, with this first and third quartile, is that around 50% of the data falls between 90,000 and 150,000.
So frankly, when it comes to using quartiles like here and standard deviation, I find myself more gravitating towards quartiles. Anyway, what about that other quartile function, specifically that one around exclusive values? Well, once again, I can select the array that we're going to use; we're going to use M, and then finally, the quartile itself. Now notice for this one, this one doesn't have a value of 0 and 4 that you actually can put in for the Min and Max; it's exclusive, so it excludes those outliers basically of the Min and Max. So specifying that column next to it, when I actually drag this down, we can see that the Min and Max are provided in this, but it's the same values for that, say, first, second, and third quartile. If you notice here, we get this NUM error, and as we inspected when going through this formula, zero and four were not available to actually input into the formula. So any time you're inputting things into a formula that doesn't necessarily exist, you're going to get this NUM error.
All right, the last function to investigate is the MODE, and this returns a vertical array of the most frequently occurring or repetitive values in any array. In our case, we'll once again provide column M, and surprisingly, we find that 90,000 is one of the most repetitive values. If we go back to that histogram we plotted earlier, we can see that the largest line right here, with a value of 1925 occurrences, occurs between 87,000 and 93,000. So this makes sense on the 90,000 being the mode.
So let's get into some data analysis now by actually ranking the average salaries of these different job titles. I'm going to go ahead and hide the columns V through R. Now, in order to rank the salaries of the different job titles, I have this list here for you where you need to first calculate the average salaries of each of these job titles. So for this, we're going to be using, as last time, AVERAGEIF. First, we need to specify the range that we're going to be basically running that IF on, not necessarily the values, but the range of the job titles. Next, we need to provide the criteria for this; we'll provide it of "Data Analyst," which is in W2, and then finally, the actual average range of column M. Dragging this all the way down, we have our different averages for all the job titles. One note real quick: in future lessons, we're going to be jumping into using MEDIAN to evaluate these job titles, because personally, I like that more, but that's a slightly more complex problem, so we're going to stick simple for now. Anyway, with these averages, we can now actually rank it. And this returns the rank of a number in a list of numbers; it sizes relative to other values in the list. So first, we need to put in the number that we want to rank; in this case, we want to do that of "Data Analyst," and then from there, they have the REF, or the reference array; in this case, we're going to provide it right here from X2 all the way down to X11. Now I can change this from descending to ascending, but I'm going to keep it how it is. Now I'm going to drag and drop all the rest of these, and we had a little bit of an issue CU; we have repeating numbers right here. It's obviously because I didn't lock my cells appropriately. So selecting this range that I want to actually lock and pressing F4, go ahead and lock that, and then we'll drag and drop this again. Hopefully, this works this time, and boom, we have all these ranked from highest to lowest. We can see Business Analyst is some of the lowest; Data Analyst not far behind; and Senior Data Scientist has the highest.
I'm going to take this one step further; I'm going to highlight everything from "Job Title" down to the bottom salary for Software Engineers, going to go into INSERT in here and go to Recommended Charts, and basically, the first one that pops up, this clustered bar chart, I'm going to insert in, and I can just change the salary up here by double-clicking in here, and I put "Average Salary of Data Science Jobs," and there we have some data analysis is actually viewing these. One minor touch to this: I really don't like how these are unordered right now, so I could actually go up here, select these three titles right here, and then under the Home tab, select I want to actually filter it and then order this rank from—well, we'll say—largest to smallest. One note: you may not have been able to see it, but it actually rearranged the data inside of our data set. That's not a big deal for me; I'm not caring too much, but that is something that will be effective whenever you do this. Anyway, with this, we can see things like senior roles are getting paid the most, and things like Analyst are sometimes getting paid the least compared to these.
All right, you now have some practice problems to go into and thus practice your skills with these statistical functions. After that, we're going to be jumping in the next lesson into arrays, which is a super powerful feature, sort of new to Excel in the past few years. All right, with that, I'll see you in the next one.
We're going to be now shifting gears and jumping into a more advanced topic of arrays. And with arrays, what you can do is, typing a formula in a single cell, we can use this to fill in cells below it or cells to the side of it, all with one single formula. So we're going to be slowly working up to an easy, then a medium, and then a hard problem, and how to use these. First up, with the easy one, we're going to go through and basically identify all the unique job titles and then go through and actually sort it alphabetically using arrays. Next, we're going to move into our median problem of calculating the median salary. If you recall back to our last lesson, we were calculating the average salary based on a job title. Well, we can use arrays to calculate the median. And then, finally, one of the most hardest problems we're going to get into, actually looking at, based on the month, how many different jobs were submitted during that month. And before this, we'll be using the SUMPRODUCT formula and a combination of other ones using arrays. For this, we'll be using the arrays formula Excel workbook.
Now, before we jump into those problems, we need to first understand that there's actually two different types of arrays. We're going to start with the first one of modern dynamic arrays, which we've seen before. And with this, what we can do is, using a formula, we can specify a range to identify, and then whenever we press ENTER, B2 to B5, it's going to actually fill in with all these. We can see that it's modern or dynamic because it has this shadow around the edge. If I select any of the other ones and not the core one, where this one's actually highlighted, when the other ones, these are grayed out. Taking this a step further with array multiplication, we can actually go in and multiply this column, one of A2 to A5, and multiply it times B2 to B5. Anyway, in this sequence, you can see that it goes down: 1 * 1 is 1, whereas 4 * 4 is well 16. Anyway, that's modern dynamic arrays. Classical arrays: let's say we want to do the same thing. In this case, well, we're going to have to go about it a little bit differently. We need to select all the cells that we want to fill in first; this is a very key concept to get for right first. Then, from there, we can start entering our formula. So I put in equals; in this case, we want to do the same array multiplication. I'll take A2 to A5 times it times B2 to B5. Now, whenever I am done with this and I want to actually execute this, I don't just press ENTER; I have to press CTRL+SHIFT+ENTER, and then it fills in the array. Notice it's not grayed out around the edges like this as a shadow; this one does not do that, and all of the different formulas are now filled in below this. And you'll notice that there's a curly bracket around this. This was used prior to around 2020, and so you may come into contact with Excel spreadsheets that have this, and if you don't know about it, if you come into here and say you want to like mess with this formula and you press ENTER, you're going to get an error message. But now let's say we have some additional values in it; we'll say we'll add five to each to the bottom of these. If I wanted to adjust this array, if I came in here and then change this to six for both the bottom and the top and press CTRL+SHIFT+ENTER, it's only going to adjust the ones that were previously selected. So now if I want to include this bottom row right here, for modern dynamic arrays, it's pretty easy; I can just come in here and adjust this to six, and this is done. However, for classical arrays, or classic arrays, not classical, I have to actually select all these different cells and then go in and actually enter the formula that I want to enter. If I try and press ENTER, it's going to give me an error message, and I realize, okay, I have to press CTRL+SHIFT+ENTER, and it'll actually fill in. Anyway, the main point of this is classical arrays are a mess; we're going to be focusing on modern and or dynamic arrays for the remainder of this course, but you need to be aware of classical arrays in case you encounter them in the wild.
So jumping into our data analysis, we're going to be focusing with the data set that we've been focusing on before, and I've hidden any columns that I don't feel are relevant for our future analysis that we're going to do. Anyway, the first thing we're going to do is find the unique job titles, and for this, we can use the UNIQUE function, and this returns the unique values from a range or array. So the first thing we need to do is actually put in the array itself. I don't want to actually select this column A because I don't want this "Job Title Short" to appear, so I'm going to select A2 and then press CTRL+SHIFT+DOWN to select all the way to the bottom. I'm going to close this parenthesis, and we have all the different unique titles in there. Now I want to get the sorted job titles out of this, so as you guess, we're going to use the SORT function, and for this, all we really need to do is specify the array. Now, if you notice from this one, whenever I went ahead and selected it, it specifies that R2# and that basically says that, hey, there's an array basically formula inside side of R2; we want to extract all the contents of that using R2#, and so that's going to work to be able to provide us all those values, and then it's going to sort it. In this case, we have it sorted in alphabetical order. One thing I haven't called out both times is these are once again dynamic or modern arrays; you can see that gray box around each of these, but just to show this also works by specifying R2 to R11, it's going to provide us the exact same results, but I really like the shorthand nomenclature of the R2# sign.
Now we're going to get into calculating the median salary, and if you recall back to our last lesson on statistical functions, we went through and calculated the average salary for each of these job titles using an AVERAGEIF function, but as it discussed last time, when comparing something like the average to the median, the average in this data set is slightly higher due to those basically outliers of those high salaries around 960,000. So we want to use MEDIAN. So what are we going to be eventually calculating? And now that's this table right here where we sorted our business or job titles themselves, and then we go into actually calculating the median salary based on these different job titles from our data set. Now there's a pretty complex formula going into here, so because of this, we're actually going to break it down step by step by step, going through each column, explaining how this actual process works in order for us to get to this final value. For this, we're going to be doing it for "Data Analyst" only, as we can see the final value we're going to get to is 990,000, which over here, which our final results, 90,000. So I'm going to go ahead and delete this to actually start with. Now we need to look for two separate conditions: the first one, we need to look to find, do the job titles here actually match up to this value here of "Data Analyst," and this provides Boolean values back where we get to this value down here for TRUE as expected. In row 16, we have "Data Analyst." Now, if we scroll down further, we can see that our next "Data Analyst" job doesn't have a salary for IT; these type of things will throw off our final MEDIAN function that we're going to actually be calculating, and so we need to basically filter it out as well. Well, so with the salary data set selected, I'm going to then go through and filter this basically not equal to a blank value, and as expected, we're getting FALSE values for these blank ones. Now, similar to what we saw in the intro in arrays where we were multiplying different arrays together, we're going to do the same thing here with these Boolean values. For this, I'm taking that formula and wrapping it in parenthesis; it needs to be in parenthesis in order to execute properly for the we contains that "Analyst," and then the second condition that the salary can't be blank. Whenever we multiply these two Boolean values together, we get returned back either a zero or one, and the only way we get a one back is if both these values are TRUE, which is the condition we want to meet. Now, for zero or one values, we can actually see if we did an IF statement here, if we did a logical test of zero, what is it going to return, whether TRUE or FALSE? So for zero, it returns FALSE, and for one, we'd expect to turn TRUE. Anyway, we don't want to necessarily return TRUE in this case; we want to return the salary that corresponds to that row in the data set. So I'm going to go ahead and delete this. So for this, we're going to start with that IF function itself, then I want to place all the different contents that we saw in that previous V column. Now we want to return the salary, which are these contents right here, so I'll be our value if TRUE, and then if FALSE, we just want it to be FALSE, which we can just leave blank. So now scrolling down, we can see that we have nothing but those values for "Data Analyst." Scroll over just to confirm: 129, yep, that "Analyst." All right, the last step: we need to go ahead and put inside of our MEDIAN formula all those contents that we had before, that entire IF statement itself to evaluate, so that array that it's going to basically find out for all those salary for "Data Analyst," and it's return back the median salary. Now this also works for other functions, so let's say we wanted to use the MODE; we want to use a MODEIF condition; they don't have this available, so we could just plug this inside of MODE, and then running this, we can see that well, the most common value for "Data Analyst" apparently also the median of 90,000.
So going back to our data sheet, let's actually go through and step by step calculate it for each of these different sorted unique job titles that we did previously, and we're going to be building this step by step, how I normally build a formula. So the first thing: we're going to look for the job titles itself; do they match to that "Business Analyst"? So selecting column A2 and then CTRL+SHIFT+DOWN to select all the contents on the cell, we want to see if that's equal to this "Business Analyst" row right here. And now remember, we're going to be dragging these down; do an autofill, so we need to be particular about how we lock these cells; specifically, we do need to lock these values right here, and just for safe measure, I'm going to lock the column of this one. Okay, pressing ENTER. All right, we, we have our array back, looking for "Business Analyst," and we can see that it's working by what we see down here in row 84. So let's actually do that array multiplication by now filtering out salary that doesn't have values or blanks. So we're going to put another set of parentheses next to it; we'll put in our salary data and that it's not equal to blank. Running this, we confirm that the first value of "Business Analyst" that has a salary has a one. Now we need to wrap this all inside of an IF to basically return, instead of that one, we want it to return the salary itself. So for the value of TRUE, I'm going to put in the selection of the "Salary Yearly." Running this, we confirm this is again correct, looking at row 180. Almost done; just now need to wrap this all inside of a MEDIAN function, and bam, 85,000. And hopefully, we actually locked all the cells properly; dragging it down, bam, looks like we got all our things, and we slightly messed up our formatting here, so I'm going to go ahead and put a thick outside border on again to make that right again. All right, so that's how you basically transform any function in Excel that doesn't have that, you know, COUNTIF or AVERAGEIF function or capability into other functions.
Now moving into probably the most complex example that we're going to be, be using, not only this lesson, probably in the entire course, so if you get around this, you're going to be good to go for the rest of the course. Anyway, what we're trying to look at here is the count of job postings based on the month that it was posted in, and we're going to be using the SUMPRODUCT function for this. Now SUMPRODUCT is not anything that you should be afraid of; basically, before we were doing, whenever we were doing the intro and we were talking about array multiplication, how it went through line by line based on this, and we have our values of 1, 4, 9, 16, and 25 line by line. Well, if we were to do the SUMPRODUCT of the values in column A along with the values in column B, we're going to get 55, which when we look at the sum of these values here, we can see that it is 55, so it's a sum of the product of the arrays. So getting back to our example that we're going to be solving, we're trying to aggregate it by these names of these months. If we actually scroll over to the data set itself, the "Job Posted Date" is in a date time format. So similar to the last example, I'm going to be walking you through column by column by column on how we get to this final value that we're going to be eventually putting into our table here to thus calculate these values for the counts per month. So we go ahead and clear these cells to start, and we're going to start first by: we want to extract out the month from this "Job Posted Date" column. So for this, we can use the TEXT function, which we're sort of jumping ahead because we'll be doing text functions in upcoming lessons, but there's a good little sneak peek. Anyway, we can plug in here something like a date time value, and then from there, we wanted to output what is the format text for? Well, I know that if we do three Ms, it's going to provide me the shorthand month of this. Additionally, if I do four Ms, it's going to provide me the long month of this, and there's a host of different format codes that you can provide by this table here when I'm looking it up in something like Perplexity that says that, hey, if you provide certain things, like if I provided a double Y, it's going to provide the two-digit year, and so on for other values. You look this up in something like ChatGPT. Anyway, get back to this example itself; I want to actually autofill this all the way through; it's not around any other.
Columns that I can actually autofill all the way down, and I don't want to sit here and drag it all the way. So what I can do is select the column itself, and then when it has these basically four arrows, I can then drag it where I want. I'm going to drag it right next to here, and then now actually autofill it all the way down. Now that I have it complete, I'm just going to select this column again, make sure I have those 4 arrows again, and drag it back to the column it needs to be.
Now, seeing what you did here, you probably like, "Luke, can't you use something like a COUNTIF in order to calculate the months?" Now, using this, and you'd be correct with that. Remember, call back for the COUNTIFS, we can provide a criteria range; in this case, we're going to provide it column V, and then for the criteria itself, we'll provide the actual month, and then actually dragging and dropping this all the way down. Once again, my formatting got messed up, so I'm going to put that thick outside border back on there. Anyway, these values here for what we're going to get finally are the same, and so you really could stop this lesson right here. And if you want to do this of creating a new column and then just using COUNTIFS, you can do that, but this is a lesson on arrays, so we're going to get more complex with this in order how to use the arrays in order to actually calculate this without having to create these extra columns. So I'm going to go ahead and hide this because we're not going to use it.
So before we can actually summing up, we need to get an array of all the values that we'll say equal to January. So, so we'll start by creating that TEXT function. It's going to be slightly different before, because we're going to be making it out of an array. We want to actually select all the values from H2 all the way down to the bottom. We want to then go ahead and lock it; we want it to be evaluating for that long month name, so four lowercase M's, and when I want to check if it's equal to, in our case we're looking for January, we'll look up here at this U2, or U1. I got a typo up there; update that to U1. Anyway, we now have, okay, that this value is true right here, and we can tell from row 11 that this is in January; it is true, so it's working out just fine. So now if I tried to actually run a SUMPRODUCT, which is what we're finally trying to do on all the contents of this array itself, we'll do W2, uh, #, we're going to get back zero because this isn't in the format that we want. We actually need to convert this, unfortunately, although it is on the back end is zero, and on the actual functions themselves can't actually calculate it. So we can do this by basically converting it, and the first thing we can do actually is just, we'll put one negative sign, and then I'll put in that W2#, and what this does is it negates the Boolean values. So basically, true, which is normally a one, it negates it and makes it negative one; zero a negative zero is negative. Anyway, we need to actually apply two negative signs, because we don't want it to be negative one; we want it to be positive one. So doing this one more time, we now have positive ones in there. So now we are using SUMPRODUCT, because SUMPRODUCT I feel are better with arrays, but we could use, in this case where it's a single array, we could use actual just SUM itself. I didn't want to show that, and we get that value of 3102, which correlates to what I expect as the value, but we're going to use SUMPRODUCT because, as you'll find out in future lessons, we're actually going to be modifying it even further, what's inside of here, and so we need this SUMPRODUCT in order to do those. Anyway, we get the same value of 3,102.
So going back to our data tab, let's actually calculate this fully for all of these different values, walking through it step by step by step as we do previously. We're going to start with our TEXT function, and we want to look at that job posted date column. I'm going to go ahead and lock all those cells; it's very important for this, going to be dragging and dropping that down, and remember for the format text to this, we want it to be four lowercase M's, and in this we're checking whether it's equal to this value here of V2, which is January, and I'm going to go ahead and actually lock just that column. Pressing enter to make sure it goes correctly, yep, we got True Value here for our row 15 value. First thing we want to do is do that double negation, which we need to actually wrap these in, this whole formula itself in parentheses in order to get our zero and one values, and then finally we're going to wrap this all once again in SUMPRODUCT, putting that closing parenthesis on there, pressing enter, get 3102, and then doing autofill all the way down, we have all our values. Once again, format is messed up; I'm going to put that thick outside border. Now, the other reason why we're using SUMPRODUCT in this case is because in older versions of Excel, before we had these uh modern dynamic arrays, SUM is not going to be able to work over arrays, and you actually have to use SUMPRODUCT. So this allows us also to have a safe way to calculate using arrays and then give it to people that may be archaic and have older versions of Excel. All right, it's your turn now to jump into some practice problems to get more familiar with working with arrays inside of formulas. In the next lesson, we're going to be getting into, probably, I think, one of the most funnest types of functions, lookup functions like VLOOKUP and XLOOKUP and things like that, which are super helpful for data analysis. All right, with that, I'll see you in the next one.
Lookup functions are one of the most, I'd say, funnest functions whenever you're learning to be a freak in the sheets; specifically, we're going to be focusing on three different lookup functions: VLOOKUP, HLOOKUP, and XLOOKUP. V and VLOOKUP stands for vertical; H and HLOOKUP stands for horizontal; and X and XLOOKUP just, uh, they wouldn't be different. In order to learn about these functions, we're going to be performing some data analysis, and if you recall back from our math and statistical functions lessons, we found out what the median, Min, and Max salaries were, but for the things like the Min and Max, what were those different job postings that correlated to that? Well, based on the structure of our data set, we can use the VLOOKUP and also XLOOKUP functions in order to find this out. Now, because of the structure of our data, we're going to have to do something different in order to implement HLOOKUPs, and for this we're going to be able to get out or extract out horizontal type data. We're going to basically transpose it into a vertical format using HLOOKUP, but if there's anything you remember from this lesson, it's that of XLOOKUP; this one is the most dynamic and flexible and how it can be used, and we're going to be doing in a final example using this in order to bucket our salary data set, allowing us to categorize it into different ranges and whether it has data or not, all using XLOOKUP. For this, we can start using the lookup functions workbook. We have two main tabs in this: data and data or 2. Data one's where we're going to start in first for this section on VLOOKUPs.
So for this, we're going to be using that job posting data set. I've hidden any unnecessary columns, and we're going to be filling in this table right here. So what I'm trying to do with this is fill in based on this Min, as you can see the formula for Min, the formula for Max, and the formula for median, where we actually calculate this from the Sal year average column. We want to then extract out based on these values the company name, a job title associated with it, and then the country associated with it. So we're going to start with VLOOKUP first, and VLOOKUP looks in a vertical type format; specifically, it says it looks for a value in the leftmost column of a table, and then returns a value in the same Row from a column you specify. So for the first value of this, we want to provide the lookup value; in this case, we want to look up 15,000 from that salary year average column. Then from there, we need to provide the table array. Now, remember for this, it needs to be the leftmost column of the table, and we want to get columns M and O. I'm going to select column O because if we start at M and try to go down, it's going to mess up because there's blank in it, so I'm just going to do control shift over, and then control shift down to select all the data, and then change this a column to M instead. The next thing we need to specify is the column index number, and right now we're in column M, so that would be the First Column, so M, N, O; we're in the third column. You can imagine if we have a buttload of columns, what kind of problems are going to run into, so we'll get to that when we get to it. Okay, now they have a range lookup; we're going to leave that blank for the time being; we're just going to execute this formula as is, and for this we're getting an NA error; if we actually click into it, value not available error. And why is that? Well, if we actually go back to that VLOOKUP function in the definition that it provides for it, the last statement is, by default, the table must be sorted in ascending order. Right now our salary values are not sorted, so it's having issues going through it and actually finding that 15,000 because it's unsorted. Anyway, we're not going to actually sort that table; that's going to be too much work. We can actually now go into that fourth parameter of range lookup, and instead of doing an approximate match, which was the default, we're going to do an exact match by providing false. In that case, we find that Net2Source Inc is the company name of the job with 15,000. Now I want to autofill for this, but we need to actually lock some cells real quick, so I'm going to lock this right now by pressing F4. Then from there, we'll drag it down. Now, one thing to note on VLOOKUP, XLOOKUP, and also HLOOKUP, this is just going to return the first value. So in this case of this 115,000, it says it's Volt Technical Resources; however, I do a control F of 115,000, we'll find that yes, it's at row 19 for Volt Technical Resources, the first one that provides, but it's also in row 42 with Northrop Grumman. So it's only providing that first match. Now what happens if we wanted to next get things like the job title or the country itself? Well, if I were to put in the first two values, the lookup value and then the table array, what will we put for the column index number? Remember, in VLOOKUP, the leftmost column of the table itself is what we're going to be looking up, but however, columns A and K are even well more left of that table, so unfortunately we can't use VLOOKUP for this, but we will be using XLOOKUP for this; that's why I'm going to recommend it over VLOOKUP, but I think you guys start at the basics first. However, before we get into that, we're going to now shift gears and cover HLOOKUP in order to look up values in a horizontally oriented table.
This case, this is horizontally oriented because we have things like the months across the horizon, if you will, and then we have in the columns, in the column standpoint, we have the job titles of the different ones of data analyst and your data analyst, so on. Now, the data in this table is calculated using the data from the data tab in order to get the counts of months, and you've previously seen this in the last lesson where we went in that hard example of SUMPRODUCT where we now go through and do some array multiplication in order to find out the different counts for the job titles based on a month. Anyway, for this HLOOKUP, we want to look up based on a month what is the associated job count for a specific job type. So let's say we want to just look at that May column. Well, we can put in HLOOKUP, and this looks for a value in the top row of a table or array of values and Returns the value in the same column from a row you specify. So only selects from that top row. For this, we provide a lookup value; in this case, let's say we're looking up January. Then from there, we provide the table array itself; we can go and just select this data. Now, I could technically, I could select all this data because it's just going to go to the associated column associated with this, so that we included row A doesn't really matter. Then from there, we want the row index number; what value do we want from this January? Do we want data analyst, senior data analyst, senior data scientist? So we can just count down what we want; we'll start with data analyst first, so we'll put in that's the second row in this, so let's try to enter this, and for this we get 753, which if you go back to this, we're doing Jan A1 through M7, and then the second one. So why are we getting 753? Well, once again, this has to do with the range lookup; we're doing an approximate match; similar to VLOOKUP, it expects that these values for that top row are in, in this case, alphabetical order in order to perform that approximate match; these aren't in alphabetical order; they're actually in chronological order. So instead, we need to specify false. Now running it, we get the correct value of 982. Now we can also apply this to a situation where maybe we want to transpose these values into this new table that we have here on month and count, and then up here I'm going to also just specify what we're looking at; we're going to look at data analyst. Now with our HLOOKUP, we're going to be providing that lookup value, the table array, and then the row index number. Say if we wanted to go in here instead of data analysts, we wanted to look at data engineer instead; how can we get this to update? Well, we can use another function for this; specifically, we can use the MATCH function for this, and this Returns the relative position of an item in an Array that matches a specified value in a specified order. So in this case, I want to look up data engineers in the array from A2 to A7; it's providing me a one because it's not, it's doing the approximate match again; once again, they're not in alphabetic, so I have to specify exact match using zero. Okay, and now I get data engineers in the fifth place. I'm also going to move this column over and make this a little bit bigger. Going back into that HLOOKUP that we're going to use for this, we're going to provide that lookup value, which we want to actually lock by pressing F4. Then we're going to provide the table; once again, I said you can select that A column if you want or not; we're going to lock all these values as well because we'll be dragging it down. From there, we'll be providing the row index number, which we've calculated right here in this P based on that match that we're performing; want to lock this as well, and as far as the range lookup, well, we want to do exact match. Running this, we get an NA error because I was silly, and the lookup value we want to actually do is for the month of January, not the data engineer actual lookup; confusing this with HLOOKUP, sorry about that. So we'll put in O3 for this instead, and then running it, and now we're getting back to 236, which is not that engineer's thing; we're one off, and this has to do with how we did our match up here, which specified A2 to A7; basically, we're counting down from the second one, where in HLOOKUP we included all the way up to that first row. So this is just a simple fix by changing this one up here to A1, and now our values update appropriately, and then I can go ahead and just drag and drop this all the way down, and once again going to get into some troubleshooting because this is all the same values, and that's because I fully locked this actual month number, and instead I wanted to press F4 and only lock the column of O. Now, finally getting to the final answer, we have it, and we can confirm this that data Engineers should have 396 on the December value; that's correct, and now we can do things like this where I can go in and say, hey, instead I want to look at data analyst, and it will update for this instead. Now, once again with HLOOKUP, we run into issues like VLOOKUP; if there's values Above This top row, I can't really think of that any applications that that's applicable in this, but it is a limitation. Anyway, this is why we're going to be shifting to the next topic, and that is using XLOOKUP to now, based on these salaries that we were previously trying to identify, identifying a job title and a country associated with it.
So what is the definition of XLOOKUP? And this searches a range or an array for a match and Returns the corresponding item from a second range or array. By default, an exact match is used; that's pretty awesome considering all the issues we ran into with HLOOKUP and VLOOKUP. Anyway, instead of using a single table, we're going to be using multiple ranges for this. Let's get into it. First, we're going to provide the lookup value, which in this case is 15,000, and then we want to provide the lookup array, so we need to select this entire M column here for what we want to actually look up, but we have these blanks in here, so I'm going to just do a trick of selecting the O column, selecting all the way down, and then from here I'm going to just go in and actually change these values to M instead. Now we want this to remain the same, so I'm going to press F4 to actually lock this. Now that was our lookup array; now we want to get into what return array, or where we want actually look to see, and that's to the left of this in this job title short column. These arrays have to match up in where they are, uh, where you're selecting them. So in this case, I selected over here in the second row; I need to do the same for the job title. Then from there, pressing control shift down, I select all of them. Once again, I'm going to lock all of these by pressing F4. Now let's close the parentheses and go ahead and execute it. We can see, see that data engineer is the lowest paid salary with this 15,000. Now we can also add in this default parameter; in case you can't find a value, you can put "not found," but in our case we made or we calculated this min, max, and median from our data set, so technically this isn't really necessary. Anyway, let's see what the other job titles are for these. Max looks like it's data scientist, and then the data engineer for the median, which is that first one that appears right over here in row 19. Now doing the same for the country, I'm going to go ahead and just copy and paste that formula in that we had from the other cell, and I'm going to just adjust this now to use column K instead of column A for the actual return array. Okay, with that updated, press enter, and we can see, see that Brazil has the lowest one, and what is the highest one? United States, and also the median, United States.
All right, we're going to crank this up a notch, and now we're going to jump into actually bucketing our salary using XLOOKUP; specifically, I want to use this table that I've created in order to properly categorize different values based on this. So in this case, we have this value of 140,000; it's going to fall into our bucket of 125,000 to 200,000. There's no data in this one, so I want to say "no data." This one's greater than 200,000, so I want to say "greater than 200,000." So for this, we're going to be creating a new column, column Q, and we're going to call it "Salary Year Bucket." I'm going to go ahead also and hide this column O for the time being; we don't
Really need this for this now. Technically, you already have the requisite knowledge in order to bucket it. I could put in a nested IF function similar to below, and it has 1, 2, 3, four, five (if you will) nested IFs to go through and basically check each of the different values as it's going through in order to bucket it appropriately. In this case, it correctly categorizes it, and then if I wanted to, I can drag and drop it all the way down. But now this sheet is filled with all of these nested IF statements; this is really going to slow your spreadsheet down, so I don't recommend doing this. Also, building something like this, you've now hardcoded in your values into it, and what if you want to change this later? You'd have to update all your formulas; it's a mess; don't recommend doing it. So I'm going to select all this, control-shift down, and then just delete it all.
Instead, we're going to be using XLOOKUP for this. Specifically, we need to look up the lookup value, which is going to be the same one that we did before, that M2. And then we want to look up the lookup array. Now I conveniently made this table here that it's providing values at (if you will) the higher end of the bucket, so we're not going to necessarily do an exact match for this; we'll get to that in a second. Anyway, now we want to look at what do we want to return: the return array, which is on the left side of this table; that's the values I actually want to return back in that column. If not found is not necessarily applicable.
So now, getting into how we're actually going to match based on these salary buckets based on these values highlighted in this T column right here, well we need to do not exact match; we need to do exact match or next larger item, and this is the value of one. Basically, in this case of this 12850, it's going to look for initially an exact match of 12850, and it's going to see that nothing's there, so then it's going to look for the next larger item, which is that 200,000. So therefore, it's going to return, as we're going to find out, the 125,000 to 200,000. Now I can try to drag and drop this down, but I'm going to run into errors because I didn't lock my formulas correctly. So I need to go back in, lock that S column with F4, and lock that T column with F4, and then I'm just going to autofill all the way down. And now we have all of our different job postings bucketed into these different salaries. So instead, I wanted to go through and actually change this to be 150k, and then match this to 150,000; I go do it, and it would update appropriately. I also need to update this column as well, but now it all updates, and it's in one single location. So this is really the power of using that XLOOKUP over the IFs in order to perform this type of bucketing. All right, you now got some practice problems; go through and get more familiar with using these different lookup functions. As I said before, make sure you're prioritizing understanding that XLOOKUP; it's the most powerful. But the one caveat to XLOOKUP is that it was introduced around the 2020s, so anybody using, once again, an archaic version of Excel beyond or before this year, they're going to have compatibility issues using this. So that's why you need to also be familiar with that VLOOKUP and also HLOOKUP; you're going to encounter them in the while. All right, with that, I'll see you in the next one, where we're jumping into text functions.
Now I know this is a course on data analysis, but text functions are actually imperative for performing analysis on text data. And for this, we're going to be working in this lesson on a data set of job applicants, and we're going to take it a step further using text functions in order to analyze. Specifically, for our final analysis, we have information on the different skills that each one of these job applicants knows, so we're going to be able to perform an analysis to see what are the most common skills from these applicants. But before we get to that final analysis, we first need to beef up our knowledge; we're going to focus on three main areas. The first one is text combination; we're going to be working to combine different columns into a single column. From there, we'll move into the second one of text extraction, being able to, out of a single column, extract multiple values. And finally, in the third one, performing some sort of text search in order to also extract out—in this case, we're going to be extracting out the state name from an address that contains a city, state, and area code. So for this, you can start up by opening up the text functions workbook, and in the data tab, we have this data set which you haven't seen before; it's only about 20 rows and includes a list of job applicants. Now we're not using the full data science job posting data set because a lot of the examples we're going to do in this it would be basically bogged down your Excel spreadsheet, so especially how we're going to be implementing these, it's really meant to be used for smaller data sets. You may be like, "Luke, what if I have a bigger data set and need to clean up the text?" Well, that's where Power Query comes in, which we'll be covering in the advanced chapters, so stick around for that. Anyway, moving into text combination, we want to target these columns right here, F and G; we want to combine them into one line to have a single address. So I'm going to go ahead and hide this column H for the time being; we're going to be putting that full address in column J. And this one's pretty simple; all we're going to do is TEXTJOIN, which concatenates a list or range of text strings using a delimiter. The first thing I need to specify is the delimiter: how am I going to separate that street and the city state? All I want to do is a space, so I'll do that, enclosing it in double quotes. Next is ignore empty: basically, if there was an empty cell in here, it would just ignore this, and it's not going to input multiple different spaces between it; just ignore it. So we want to, in that case, we're just going to put in TRUE. The final one is text, and we can specify—you could do text and then comma and then text to—um, that's really verbose; I don't really like doing that. Instead, I'm just going to select the range of F2 to G2. Now we can see that the address is fully concatenated, and we can drag it on down, and it works for all of it.
Now the opposite of combo is extraction, which we're going to get into next, and in this case, we're just going to use a single column and extract out multiple values. In this case, we have this full name column; we want to extract out the first name and the last name. Go ahead and hide these other columns; we're not using them. In this case, we're going to specify the TEXTSPLIT, and it says it splits text into rows or columns using delimiters. So we'll first start by specifying the text, which is B2 in this case, and then the column delimiter, which in our case is going to be that space. Once again, we're going to use that double quotes for that space and then end double quotes, and then this is going to be a dynamic array, and it has these two values here. Now, dragging this all down, down, we see that it fills in for all these different names. Now we just split text there. Also, could be cases where we maybe want to extract out a certain amount of values or a certain amount of text from a column. In this case, we also have our application ID number, which is a combination of letters and numbers, but as you can see from this, there's some values in here that are actually repeating. Sometimes we want to refer to this, the shorthand of this, and let's say we only want to get the last three digits of the applicant ID because we know that's always different. Well, in this case, we can specify the RIGHT function, and it returns the specified number of characters from the end of the text string. We specify the text itself and then the number of characters. In this case, we can just say three, and it's going to provide back that 548. We could also just change that to include all the text numbers in case this number gets bigger than that, and then go ahead and drag it all the way down.
Now one last one before we get into actually performing that analysis: we want to, we want to go through and extract out the state from this city, state, and zip. And as you notice from all these, they have a common format in that the city has a comma, and then the state starts, and then there's another comma following that. So we're going to be using those basically delimiters (if you will) in order to identify where we should potentially extract out this state value from, where this state, these two-letter value. So the approach we're going to use for this is, as we go through this, is we're going to find the location first of that first comma, space. Before this, the next, we'll find where it actually ends, and then finally, well, using those values, will actually extract out using the MID function that state abbreviation. So the first thing we need to do is find that comma, and this returns the starting position of one text string within another text string. So in this, I'm going to specify that the find text we're going to be looking for is the comma itself, and we're going to be looking at within text, obviously G2. Now we also need to find the second comma. In this, we can use that FIND function again. Specifically, we're finding that comma, specifying that within text of G2, and now we have the second optional parameter of start number; we want to start from nine, which is the first one we found. This in running this, we get nine. Now the problem here is because we're starting as the exact number that the comma actually starts, that's why we're getting that back, that value of nine. We need slightly actually bigger than nine, but anyway, we'll fix that in a bit. Instead, let's actually get into extracting out that, or at least trying to extract out that CA of this value, and then we'll fix that issue in cell R2. So for this, we're going to be using the MID function, which, which similar to that RIGHT function, is returns the characters from the middle of a text string given a starting position and length. So in this case, we want to extract out G2, and we'll provide it the start number of, well, what's valuable in Q2, and the number of characters, and for right now, we'll just put in—we know we want to extract out two, so we're going to put in two. Now we're running into issues; we're only getting back a comma, if you will, and if we actually make this longer to actually zoom in on here, we get comma space CA. Now when providing four, and this has to do with right here, this value on the start number isn't correct; this nine right here is exactly at the comma; we need to actually specify for that start number of where the C is, and these are all two spaces over. So I'm going to come in here, and I'm just going to modify this shortly and add two to this. This is also going to fix our previous one that we had when finding this of 13 because 13 now has all the way over, and then finally that MID is fixed; we can change this now to back to two. Now you know me, I don't like hardcoding values something like this two, and really what we're doing here is we're doing adding two based on the length of the comma and then the space after it, so there's two characters in there. Two. This is still that 11 value that we saw before. Similarly, inside of our MID function, I don't like doing this two here because states maybe could be more than two, so I don't want to hold it necessarily to that. So instead, I'm going to do R2 minus Q2, which in our case is going to be two, and we have California. All right. Now we can take all the different values, actually drag it on down, and we get all of our states extracted from this. All right.
Diving into our final analysis, we're actually combining all of these different functions we just learned about, specifically with this data set. We have this column H right here, and it's a list of different skills that each one of these job applicants have. We want to combine this and aggregate this in order to analyze the most common skills. For this, we're going to have to walk through four different steps in order to get this into our final visualization that we can actually visualize and see here. So I'm going to go ahead and clear all these values so we can get started actually doing this. The first thing I want to do is actually combine all of these values into a single long text string, and we're already having the separator of a comma and space between each skill, so we're going to use that same separator to continue separating this. So using TEXTJOIN, we're going to first specify the delimiter of that comma and a space; it's asking if I want to ignore those hidden or empty cells; I do. And then finally, we need to provide the actual text itself, so we'll go down through and select in our data tab H2 to H21. Going back up into the formula bar, closing this parenthesis and then pressing Enter. Look, I have a like slight typo in here; I need to actually put double quotes around both of them; you can't mix double and single quotes. Now we have this super long list uh that has all of our different skills in it; it looks like it's properly delimited. Now that we have all these values in one cell, we can then use the TEXTSPLIT function to now separate this into different cells because we're going to want to then move into transposing it next. And for this, once again, the delimiter we're using is that comma and space. Running this, we have all the different values separated out by different cells. So now, almost there, we need to get into making a table right here, basically having skills in the left-hand column and then the counts of those skills from what's above here. So first thing we need to do is get the unique values of this, but if we just run UNIQUE on that row six, we're going to run into an issue to where it actually goes out to the right and actually doesn't get the unique values for all these. So the first thing we need to do is actually transpose, which moving it from horizontal to vertical of that row six. Okay, so it's now up and down all the way. Now, in this case, we want to run the UNIQUE function on this to extract out all those unique values, and scrolling down, looks like we have all the unique values. It does have a zero because that we did that row six, and so when we get to these empty cells over here, keep on scrolling over here, it records as zero. I'm fine with that for the time being, and we'll continue. Last thing we had to do is use basically a COUNTIF to count these different skills based on whether they appear or how often they appear in this row of six. So I'll type in COUNTIF; we need to specify the range first, and we'll do six; I want it to stay there uh as we're because we're going to autofill down, so I'm going to F4 that, and then from there, specify the criteria, which is going to be A1. Okay, so three values for that one, and then dragging this all the way down, bam, got this all filled in. All right. The last thing we need to do is actually visualize this because we want to visualize these skill counts. Select the area that we want; we're going to go in and insert in under recommended charts; you can do a bar chart, but I'm more a fan of horizontal bar charts, especially when we have text values and we need to be able to see all the different names. So I'm going to have to expand that out a bit, and I'm going to change this title up here just to something like "Skill Count of Applicants," and bam, now we can see things, some trends out of this, that a lot of people are claiming to have experience with Databricks, which that's unusually high; there's probably something I want to investigate for this, but a good little thing that we actually can analyze and see from this analysis that we did. One minor note: I would normally go through and actually sort this from high to low, and you can definitely do this; you'd have to copy and paste the values over; you wouldn't be able to use these values right here and S sort and filter them because we're using the modern or dynamic array to find these unique values. So that's definitely an option if you want to do, and I definitely would recommend you do something like that before sharing some sort of visualization like this. All right, you now got some practice problems to go through and get more familiar with these text functions, which, like I said, are imperative for data analysis. In the next lesson, we're going to be moving into our last one in this chapter on formulas and functions on date and time functions. With that, I'll see you in the next one.
All right, saving the shortest lesson for last, we're going to be focusing on date and time functions. And for this, we're going to be using that same data set from that last lesson, which is about 20 rows of job applicants. Now, similar to text functions, we're not using that full data science job data set that we've been using previously because I find it's not common to really use these date and time functions on a large set of data because it's going to slow down your sheets. So that's why we're using this smaller data set for this. Once again, if we're needed to actually clean up date and time stuff, we're going to use something like Power Query, which we're going to be getting to in the advanced chapter. Anyway, we're going to be focusing on two main types of functions: first up, our date functions, which are going to be able to extract out things like month, day, and year, and then from there, we're going to transition into time functions, extracting things out like hour, minutes, and seconds. Finally, we're going to move into that final analysis, looking at what is the time that is most likely for applicants to apply to jobs. For this, we're going to be using the date and time functions workbook, and we'll be working in this data sheet for this, filling in certain values as we go through this. I'm going to go ahead and hide some of these unnecessary columns so we have more space to work with this. Anyway, jumping right in, if we want to calculate what the month is, we have something like the MONTH function, putting that in, that's D2. Similarly, we can get the day by using something like DAY, and once again, providing it D2. Then finally, something like YEAR, we can provide D2; we get 2023. Now if I wanted to only extract out of this date, out of this date time, if I were to use this DATE function, it returns the number that represents the date in Microsoft Excel, okay, date time code, got it. We're going to put in the year, so we need to provide the year first, month, and then from there, day. Boom. And analyzing this, we see it is February 14, 2023. One quick refresher on how Excel stores those datetime objects. So right now, it's in as a the number format of date; if I change this back to General, it's going to shift to this number, and if we recall, this stores the values in it; if we start at something like one, converting it to a short date, we can see that it starts at January 1st, 1900. Now if you're working with dates before 1900, let's say we put in something like negative one, I converted it here to a date; it's going to just provide all these different ampersands here. There's a few different workarounds for that; that's beyond the scope of this course. Main thing to understand is how it's actually stored within Excel. Anyway, I'm going to convert this back up into a date, and for each of these, I want to actually fill in the values all the way down. Bam. All right, close up this home ribbon. All right, next up is TODAY. Say we needed today's date; well, I can put in the TODAY function; this actually takes no arguments and will provide us
The date I'm filming this on is September the 3rd. Now, the last common function that I find myself using all the time is when I want to calculate the days since something happened. In this case, we want to find out how many days it has been since they applied to the job. So we can use the DATE DIFF function for this.
Now, the one thing to note with this is I'm typing it in; there's no—if I type in just "date," there's no DATE DIFF in there. There's no documentation that Excel natively actually includes for you to use this. So this is like a function you just have to know about. Anyway, it takes three parameters: basically, the start date that we want to start from, the reference date that we want to basically subtract from this—which is today. We want to actually go ahead and lock this; I'm going to lock this with F4—and we want to provide this in the format of days, which we provide this text character of "D". And this tells us it's been about 567 days since Valentine's Day in 2023. Anyway, updating all these cells for this, we now have this data.
Shifting gears into our time functions, as we can expect, a lot of these are going to be the same. HOUR we use the HOUR function; MINUTE has a function, as well as SECOND, but this doesn't really show seconds, but we can see up here it is; it is actually included in your data. Similar to the DATE function, for TIME we have to provide three parameters: hour, minute, and then also second. Drag and drop this all the way down; we can see that, yep, it's correlating correctly. One note for the HOUR that we previously calculated: this is in military time, or if you're in Europe, you also do it this way. Anyway, I really like this for an analysis purpose, especially when we get into analyzing it.
Conversely, we can also use for time and also date; you could use the TEXT function, which we previously saw when we were extracting out the month out of date times, by providing a value and then the format text, which we're going to say in this case is just "hour hour minute minute". If I wanted that AM/PM format, not that military time format, I can just add in here "AM PM", and it converts it appropriately. Dragging this all down and then filling it in, we get it.
Now moving into that final analysis, we want to analyze when are these job postings happening by hour of day. The first thing we need to actually do is get a column here of the hours in the day, so we can do some sort of like COUNTIF on it in order to calculate that. So for this, I'm going to use the SEQUENCE function, and I went 24 rows with it; the column's going to leave blank, and I want to start at one, and it's going to fill down from 1 all the way to 24. And now we need to run a COUNTIF, basically for each one of these conditions, run down this list, basically matching to see what is the hour for these things. So I have it hidden, but I'm going to go ahead and make a column again for HOUR, and I'll put in here "HOUR", and unlike last time, I'm actually just going to put the whole range in here, and it's going to provide me back it in a modern array. Now with this, I can actually now use this in the COUNTIF. We want to first provide it a range, which is our modern array, so it's going to do I2#, and then a criteria for the hour we want to search for. We want to search for that one, A2, from here; we want to fill it all in, and we have some reference errors because we didn't lock our cells; specifically, we didn't lock this cell right here, this I2, so I'm going to press F4 on that to actually lock that. Then dragging it all the way down, we have it.
Okay, our last portion of this is actually visualizing this. So we're going to go in, select all that data, go to INSERT, go to RECOMMENDED CHARTS, and I'm more of a fan of column charts with this type of data, so I'm going to go ahead and put this in, and I'm going to change this to "Job Postings per Hour," and BAM! Now from this, we're seeing that basically people are applying during normal working hours, and apparently they're waiting until the end of the day to actually submit their job applications, maybe to get in before a deadline or something. All right, this is the last lesson on functions and formulas. In the next chapter, we're going to be moving deeper into understanding how to actually make these different visualizations. I've only been showing you a sneak peek at it right now to get you familiar with how to easily create it, but we're going to go into a lot greater detail, up coming up next.
Now, we spent almost nine lessons on these functions, and it's because I feel functions are one of the most important things to understand about Excel, because it also transfers to other portions; specifically, we're going to be learning more about the DAX language in the advanced chapter, and we're going to apply a lot of our knowledge that we already know about these Excel functions to DAX functions. They're very similar. Anyway, you got some practice problems to go through and work in order to understand better how to use these datetime functions, and from there we'll get into that chart chapter. With that, I'll see you in the next one.
Welcome to this chapter on charts. And as much as I love using something like Python, a programming language for making visualizations, I feel that Excel has some capabilities built into it that allow it to basically exceed any programming language, and the customization that you can do to charts that we'll be finding out in this chapter. For this chapter, we have four lessons. This lesson right here is an intro to charts, so we're going to be focusing on understanding the basics of using charts and specifically looking at three types of charts: specifically, line charts, pie charts, and bar or column charts. So technically, that's four. In the second lesson, we're going to move into more advanced charts, such as scatter plots and also map charts, along with understanding more advanced customizations that we can do to these charts. In the third lesson, we're going to go harder on the paint in order to understand statistical charts, specifically histograms and then also box and whisker charts, which are imperative to understand statistical distributions of our data. We'll finally wrap this all up with a final lesson focusing on sparklines, which basically allow us to put charts inside of individual cells in Excel. Pretty neat.
All right, for this lesson, we're going to be using the Charts Intro workbook. First thing to understand is terminology. Microsoft refers to all these different visualizations, diagrams, plots, whatever you want to call it; they refer to it as a chart. Basically, they want to use a safe term that encompasses all the different types of visualizations we can build with this. So you may hear me from time to time call this a plot or visualization; I basically mean a chart. Anyway, why do we use charts? Well, looking at these six examples here, we can see some different characteristics about this data that we're looking at, but what if we looked at just the core data itself, which is this table right here? Looking at what is the number of job postings per month, if we look at this visually, we're not able to see necessarily what is the highest month and also what is the lowest month. I mean, you can figure out eventually, but it's not easy to spot, and that's why charts are so powerful. And so I have a variety of visualizations here in order to showcase that same table that we were just looking at in basically a variety of different forms. Here, even have a few below here, down below it, but we need to understand which chart to use because let's say we wanted to use this pie chart here; is that actually a good chart to use to visualize this, or instead should we be using something like this line chart to better show a trend over time while also showing a magnitude of difference? Anyway, as we go through this lesson, I'm going to be calling out when you should use certain charts as best practice, along with my recommended tips for how to customize it to show them best.
So for our first chart, as I hinted to, we're going to be making this job posting count into a line chart, and this is the chart I'd use typically for any time series-like data, as it's great at showing a trend over time and how it's connected. So how do we do this? Well, we're going to select all the data here, all the way from A1 down to B13, come up into INSERT, and we're going to dive into each one of these charts individually, but I would encourage you to actually just start with RECOMMENDED CHARTS. I really jump to it every time I use it. Anyway, first thing, they have two tabs here: RECOMMENDED CHARTS and ALL CHARTS. For RECOMMENDED CHARTS, it usually provides a lot of good tips that you could potentially use for different charts. Sometimes, however, I do find that I want a particular chart, and it's not here, and that's when I'm going to go to this ALL CHARTS tab, and frankly, it provides a lot more control while allowing you to actually visualize your different data. In our case, I know I want a line chart on this, but now I can go in and actually plot it with markers or even change it into a 3D line chart. Highly don't recommend this; we're going to be sticking to a line chart for this, and I'm going to go ahead and click OK. I'm not going to lie; this chart is getting us 90% of the way there. Now, if you notice, for this, when we clicked on the chart, we have certain values highlighted here; basically, this purple outline is showing that these are the X values right here, and then the blue coordinates right here are showing the actual values themselves, and then conveniently they put the "Job Posting Count," which is highlighted in orange, as the title. We'll be jumping into how to customize this area in the advanced section, but that's in the next lesson.
Now, for those new to charts, there's a bunch of different elements, and I can come up here, and I can click this plus icon right here, and it shows all the different elements on here. I can use the checkbox to control whether I want to include the axes or not. In this case, I do want to include it, and then I can even fine-tune it further to select which one I'm talking about: am I talking about the horizontal, or am I talking about the vertical? Just going through these in rapid fashion: axis titles allow us to provide titles for the X and Y axes; the chart title shown above, I can remove it or keep it on; if I want to include data labels, I can do this along with controlling what position of them I want to go with; I could also include something like a data table below, but personally, I find this is sometimes sensory overload; I don't really use that much. Next are error bars, for data, gridlines, whether I want to have horizontal, vertical, some minor ones, or some other minor ones; a legend, if there's more than one data, I probably want this; a trendline, which will be adding in this a little bit; and then up and down bars, which are going to show whether the data goes up or down based on each set, but not really necessarily applicable to this one. Now, I find this plus icon is where I go most of the time, but I could also go to this CHART DESIGN tab up here, and it has this box of ADD CHART ELEMENTS, and basically you can go through and adjust all the different ones along with showing a more visual indication of what's going on here, here showing that I was actual up/down bars to actually see what they actually look like. You can also use these QUICK LAYOUTS to quickly try out different themes that Excel has, so I do this myself from time to time using this. So this chart is almost done. All I do want to do first is change the title, and I usually like to either provide some sort of snippet of information from it or ask a question that I want the reader of this graph to understand or take away from this chart. So I can put in something like, "How Did Jobs Trend in 2023?" So it also tells what year what's going on here, and it asks them to look at, hey, what is the trend going on here, which it looks like we have a peak up in January and a peak up in August. Now, I try to minimize the amount of axis titles on here because, like in the months' case, that's pretty self-explanatory. However, the number in the Y-axis is not so self-explanatory, so in that case, I would want to include it. In this case, give it a representative name of "Counts of Jobs". The last thing I want to do with this is just add a trendline, and there's multiple different options for this. We can do linear, exponential, a linear forecast where it actually goes into the future, and then even a two-period moving average, which is pretty neat. I'm going to just stick with the basic one right now of linear, and BAM! That's our first chart. So let's move into the next one.
Now, if we go back to our original data set, in the DATA tab, we have a column here on "Job No Degree Mention," and basically this column right here includes whether there's a mention of a degree in a job posting. So in this case, where we have two different values, we're trying to determine what are the proportions of each. A way to compare this, we could either compare this in like a bar or column chart, but I feel a better one for this is a pie chart. So I've gone through and calculated a count of the jobs with a "No Degree Mention," along with those that have a mention of a degree. I calculated the total, and then from that, I calculated their individual percentages. Now, I'm not going to just select all the data here because I don't want to plot all of it. I'm going to select the first two values here of A2, A3, press CONTROL, and then also select C2 to C3. Then from here, now I'm—I'm going to go INSERT those RECOMMENDED CHARTS. Like, got a lot of bar and column charts come up, but the one we're going to be using for this is a pie chart, so I'm going to go ahead and insert that in. Now, personally, I'm not a fan of this layout here, so I'm going to come up into CHART DESIGNS into QUICK LAYOUTS, and I'm going to just experiment with different ones, looking at them, and frankly, I like the one—this one right here actually, where we've removed the legend and put the actual values themselves along with their titles inside the pie chart itself to make it super simple to see which one is which. Now, Excel sometimes gets crazy with the colors. I actually don't recommend using a lot of different colors because it could be very confusing for viewers on where to look. Personally, I want to highlight more of the "No Degree Mention," so I'm going to use this single color palette right here, or this monochromatic color palette right here that has these different shades of blue, and—and I feel the eye is going to go more to the darker blue. Now, with each of these labels here, I can actually select it; I double-clicked it over time. I can actually drag it and drop it and move it around where I want it to be. I would probably want it to be more over here; I want the "Degree Mentioned" to be stacked; basically, I want them opposite of each other. Now, you may have noticed I can't really read this text right here, and even this text is hard to read as well. So what I can do is I'll just click outside real quick, and clicking back in, I'm going to double-click, and this is going to bring up the FORMAT DATA LABELS. If double-clicking isn't working, you can just select it, go into the FORMAT tab up here, and select FORMAT SELECTION. Anyway, there's a lot to unpack in this pane, and we'll be unpacking it as we go along this entire chapter, but the main thing to understand is they have LABEL OPTIONS and TEXT OPTIONS. We want to adjust the TEXT OPTIONS, and this has things like text fill and outline, text effects, and then also the text box. For this, we're trying to fill the text fill and outline; specifically, this drop-down here of TEXT FILL, we want to change the color, so we want to change it to white. Now, if you notice, only one of these changed, and that's because I only had one of the boxes selected. So actually, actually click out of this, double-click back into this, and then make sure both of these are actually selected, go back into TEXT OPTIONS, go into TEXT FILL, and then change this color, and then it's going to change both of these colors. Now, I'm fine with this text now, but let's say I wanted to customize further the percentage here; maybe I want to include one more decimal place. Clicking on the box itself, I can now have this option for LABEL OPTIONS, and then under—well, LABEL OPTIONS again, I can scroll all the way down, or I can actually cover this up and then unhide this number. I can change the number formatting itself. In this case, I do want to still do a percentage, and then maybe I want to do one decimal place. Personally, I think there's a little—a little bit too much data, so we're just going to keep it with the zero. All right, that's the final customization. The last thing we want to do is just add a title, and I want a very compelling title. What do they want to look at for this? I want them to understand what jobs mention a degree. And now with this, we have a pretty great visual indication of that—about one-third of jobs have no degree mention in them, which personally I think that's a pretty high percentage, and hopefully gets higher.
So we have data similar to our first chart that basically explains how many counts of jobs for the different job titles. Now, this isn't chronological, so I don't necessarily recommend using something like a line chart for this; that's why we're going to be making column and bar charts for this. Also, let me explain the difference between the two. Anyway, I'm using the formulas that we previously have covered; you can dive into it if you want to; basically using UNIQUE and then also a COUNTIF formula in order to count each one of these in their DATA tab. Anyway, if I actually go to graph these by selecting all these things, go to INSERT, and RECOMMENDED CHARTS, here provides the recommended charts, and we're going to start with a column chart first. I start with this one first because we're already running into problems with how long these labels are. We can see that we have these three ellipses here, basically telling us that the rest of the name is hidden here, so not all the names are shown here. The other problem that we're getting into with this column chart—named after the fact that it looks like columns—is that it's not in an organized manner. I would expect to see it high to low to make it more easily to compare values to each other and also how they rank. So we'll go ahead and delete this bad boy. Anyway, this table is organized based on this UNIQUE function, which doesn't necessarily put things in the correct order, and I won't be able to actually go through and filter it or sort it appropriately. So below this, I made a different table that I basically use SORT to sort these values from above by their job count in descending order. Now, since it's in this order, I could actually select a few less of this. Remember how it was cut off last time? I could select only the top six, go into here, go into RECOMMENDED CHARTS, and once again, and put in our clustered column chart. Now this one I can play around with, and as you see, as I expand it out, I can actually see all the different names here, but once again, I'm not a fan of this column chart; I'm not going to be using it for this case. Instead, we're going to try out a bar chart instead. So selecting all this data to show the power of these bar charts and then coming in, I can put in that bar chart. Now, I do like this one better because all the titles are organized, and they're right off to the side, and so this is a much more easier read. The problem now is—I'm really nitpicky with my charts—the problem now is I don't like the order that this is in. What happens is is Excel starts plotting these, although
It's in descending order in our table, as shown over here. It's going to be plotting them starting at this zero axis up here and then plotting from there. So, technically, we don't even want it like this. Instead, what I can do is reverse the sort order here. I'm just controlling it by using uh, either one or negative one in that sort order portion.
Anyway, with this order now, now we can finally get into the final bar chart that we want to actually put in. And I'm just going to skip this recommended charts come up here into the column and then the bar charts; we want this one inserted in. And I'm also going to zoom out some. Now this is more in lined with what I want. Let's actually clean up this visualization to identify what we want. I'm actually more curious about what are the top jobs in data science, so that's what we'll name it. Additionally, feel the titles are pretty self-explanatory based on that title, but I would need something for the x-axis down here, so we'll add an axis title calling this count of job postings.
Now, with this question I'm asking of what are the top jobs in data science, I'm not really feeling like we need to include things like machine learning Engineers, software Engineers, cloudware Engineers, or business business analyst. How could I actually adjust this? Well, one way is I could control what areas are highlighted over here, and I could actually drag this and change this to whichever ones I want. Um, but I'm not necessarily going to recommend that. Instead, I'm going to select our data, make sure all the columns are selected themselves, right-click it, and then go to select data. This new window is going to pop up here. This tells us a lot of great things about our visualization. First is the chart data range; it tells us we're selected from A25 to B35, so we could change that here if we wanted to. The next thing is the two windows down here of the legend entries and the horizontal axis, so this controls our job count. I'm going to scroll this over here; we could just remove job count, but it's not going to do anything. This guy's mainly right here, the access labels we can control. So I know I want data analyst and all the way up down to senior data analyst. I can actually go through and select remove business analyst, machine learning engineer, software engineer, and Cloud engineer, and then click okay, and it will remove it from this visualization while still keeping this data here, so I can easily go back and add or remove job titles as necessary. And now we have our final visualization.
Earlier, I did go through and actually delete the chart and start over, but you do have this option in the chart design tab of change chart type and allows you to basically go through and try out different ones. If I wanted to go back to that column chart, I could, and it would show me an example of what it looks like. Now there is one last thing that I want to format on this. I do find it a little difficult to read exactly what are the amount of job postings that they have here, so I'm going to add data labels to this. We have a couple different options: we can be inside end, which can't read at all; inside base; outside end, which I'm more for; and then also a data call that's just too much. There, we're going to do outside end.
Now, with this, these numbers, I don't like the level of detail; I don't need down to the single or the on-digit place to tell what it is. Instead, I would rather it shows something like 9.6k or 9.6000. So we can actually format that. So double clicking on one of those labels, this format short area is going to pop up again, and for this I'm going to go under label options and then label options again, and finally number. And for this I'm going to use, use instead of uh any one of these, I'm going to use a custom type. Now I have a few of these already built into here, and so they may not pop up to you, but this is actually sneak peek; this is actually what we want. But if you don't have this popping up right now, what you can do is actually go in, in this case I'll just show a different value. What we're going to first say is how we want this formatted with how many decimal places. So I want all the values before the decimal place, then a decimal place, and then I only want, in this case let's go with two places after the decimal place, and then from there I want a K on the end, so basically to show this as a thousand. So I'm going to use a parenthesis, put a k, and then close parenthesis, and I'm going to click add, okay. So now this changes it to the double digits for explaining that this is the thousands. This automatically whenever I do that K parenthesis, it automatically does the math to basically divide that by a th and transfer this to K instead of the thousands. Anyway, I don't really, I'm going to go with the original one I had of only one decimal place, and bam, that's our final visualization. And we can see from this that we have a lot of insights into understanding that more Junior roles like data analyst, data scientist, data Engineers are more prevalent than the senior roles, and that luckily it seems like there's a lot more data analyst roles than data scientists and data Engineers. All right, you now some practice problems to go through and get more familiar with those four major type of visualizations that frankly I feel I'm using on a daily basis anytime I'm making visualizations. So don't think that they're just too plain or too simple; they're really powerful and explaining data.
In the next lesson, we're going to be jumping into not only more advanced charts but even more advanced customization. So with that, I'll see you in that one. We're going to crank this up a notch and get into some more advanced visualizations, specifically on this; we're going to be doing a deeper dive, dive into the pay of different jobs, not only based on the different job titles but also based on where a job is located, using things like a map chart. And so for all these charts also, we're going to be looking into how we can further get into deeper customization of these.
Scatter Plots are great at comparing two numerical values in our data set. We have these two columns here, one on the salary year average and the other on the salary hour average. Just as a background on why it's called average at the end of these, sometimes job postings have a range of salary, and so I took the average of the Min and Max, and hence I named this average. Anyway, we have yearly salary data, and we have hourly salary data. What it did next is get the unique value of the job titles, and then from there, using that median, basically modified median IF function, got the yearly median salaries and then the hourly median salaries. So because we have these two numerical values to compare, basically we want to see if there's a trend correlated between the two, because well there is, we're going to find out. I'm going to go ahead and select these all, then from there go into insert, and we can come into charts. I know I want a scatter plot, and if we go to insert it in, can't see cuz it's hidden behind here. Well, we'll just go ahead and show it. This isn't necessarily showing us what I want us to show with this; it's basically showing, hey, this is the yearly data up here in the blue, and then this is the hourly data. Since hourly data, it's super low, it didn't work out how I wanted to by selecting all the data like we've previously been doing. Instead, I'm going to go ahead and delete this. What we're going to do is we're only going to select basically this B and C column of data. Once again, we're going to try again inserting that scatter plot, and at this point it's actually working correctly as we want it. Unfortunately, we can't tell; there's no basically like data labels for this to understand what are the different job titles associated with it. Even with the graph, we can see that it's only highlighting this also the incorrect titles up here; it's not just hourly median salary. We're going to fix all this.
Anyway, the first thing that I want to clean up is actually the selection of data. Right now, we can see these numbers are overlapping down here; also it goes all the way down to this zero axis on both the X and Y. I want to change that. So I'm going to double click this x-axis, and format access pane pops up, and we can see that we have bounds here 0 to 180,000. I can see that there's no values under about 75,000, so I'm going to go ahead and put that in for the minimum and press enter, so it's going to jumate this way. Now I want to do the same thing for the Y AIS. I'll just double click it, and this one didn't necessarily go where I wanted it to go. I wanted to actually change the values here, so we can go under access options, under access options again, and under access options again, we can change this minimum maximum. I'm going to change it to, looks like there's nothing above 20 or below 25, so we're going to go with that. Now, even with this change in the formatting of the values here, the minimum, I can still see that there's overlap here, so I want to update this similar to last time, basically cut it off the thousands place and place and put a k at the end. So under access options, access options again, I'm going to close this drop down of access options also. Instead, we're going to go to number. For this, we want a custom type, and I do have some values in here, but we're just going to go, if you don't have them in here, we're going to add a new one. Specifically with this, I wanted to show one; I wanted to show a dollar sign at the front, and I don't want any decimal places whatsoever, so I'm just going to put a zero in there, and then from there, like last time, I want to format this in the thousand's place, so I'm going to put a comma and then double quotes to put around the K, which signifies I want to formulate this in the thousand's place. I'm going to go ahead and click add, and now this is much more readable, not so much sensory overload. For our y AIS, I don't care at all about this decimal place right here, so going back into numbers again, I can just format the decimal place places as zero, and I'll just leave this one as an accounting category.
Now, which one's yearly and which one hourly salary? Well, we need to include actual access titles for this, so I'll go ahead and enable that, and then for this we're going to do something a little bit different. I'm going to select this ya AIS title and instead of actually typing in values in, I want to use actually the column header right here, so I'm going to come up into the formula bar, type equal to, I'm going to select C1 and then press enter, and now this updates for that column head. I can do the same thing here for the x-axis title, selecting it, then from there going to the formula bar, put an equal and selecting cell B1 and pressing enter. For the title, we don't want that hourly median salary; we're really trying to find out what jobs have the highest pay, and we can basically tell it from this. All right, so let's actually finally get to adding data labels to this, and we can see what data labels are actually available, but scrolling over the different options here, we're going to just go with above for the time being. Then I'm going to close on out of this, and I'm going to select the data labels themselves, and format data labels should pop up. If it doesn't, you can also go about doing it by right-clicking this and going to format data labels. Anyway, for this, I don't want to actually show the X or the Y value for this. Anyway, uh, I made I made it disappear by actually closing out of that, so actually I going have to add those data labels again again. Anyway, going back into it, under label options, label options, then label options again, I'm going to leave that y value selected for right now, but what I want to do now is provide the job title itself right next to the data point, so I can do this option here, so label contains value from cells, and it's going to ask me to select the data label range, and so now this is when I'm going to select all of these different job titles here and press okay. So now we have these values from cells. I no longer want this y values, and I do want to include this leader lines because we're going to be actually dragging this around because, as you can see, some of these values are overlapping now. Also, I'm noticing that this is really busy right now with all this text and stuff, so I'm actually going to remove the grid lines for the time being, actually for the remainder of this, cuz I I don't feel like it really needs the grid lines in general. And now I have a little bit less sensory overload, so I can go through and actually clean up where a lot of these different job titles are located by just selecting it and then dragging it, and you notice uh we had that leader line selected, so I have arrows or basically lines going to each of these ones to signify which one is which. So now I've dragging these all over so that way they're basically more represent I want. Sometimes if I dragged off of this and drag maybe the whole chart itself and make a mistake, I press just control Z, and it reverts it back to where I'm going, and then I just continue on to selecting the box that I want and moving it. Anyway, this is pretty neat. Now I could actually go in if I wanted to and add a trend line to this, and basically it shows for an increase in that yearly salary, I expect the same with the hourly data. In this case, I don't find it as much useful, so I'm going to just keep leave that off, but in general it is pretty neat to see the trend that's going on with this that senior data Engineers, although they're underpaid compared to senior data scientist in yearly salary, you could get the hookup if in instead you look for an hourly gig instead in order to get a little bit higher pay. A similar dynamic happens between business analysts and data analyst, so if you're a data analyst and you're looking for a job maybe on upwork, maybe you should advertise as a business analyst instead.
All right, going back to our data set itself, we have another column in here I want to investigate, and that's specifically around the country is called job country, basically where the job is located at, and I like to visualize these type of things well on a map to actually see how it affects others. So I've made this table here under the map chart tab where we have our all the different countries in the data set, then from there we use a count if to determine how many counts for each of the countries, and then our modified median if in order to determine what the median salary is in each of these countries. I've also had to wrap this one in an if error because some of these if there's no values it throws an error, and I didn't want that popping up in the chart, so so I had it disappear or make it basically a blank value if it does have an error. Anyway, let's get into visualizing this. We're going to first just visualize what are the counts of these different jobs based on the country, so I'm going to select column A and B, go to insert, and then maps and go to this map chart. Now you may have a pop-up warning that comes up during this that says data needed to create your map chart will be sent to B, and I'm fine with sending this data to being; you should be fine too with it, so feel free to accept this, then you shouldn't get this pop up anymore. Anyway, this chart's pretty neat because it goes and shows we have a heavy concentration of jobs basically from the United States. For my job scraper, I'm heavily aggregating jobs from this country compared to other countries, sorry, other countries out there, but I am still n less collecting from other countries like US has 25,000; India is around 580. For this one, I'm going to change the title to where are most jobs in Luke's data set from. There's not to say the United States has more jobs than other countries; this is just how my data set is and how I extracted the data, so don't want you to come up with the wrong conclusions from this.
Now, the visualization that I really care about is comparing these countries to the median salary. So holding control, I select A and then C. I'm going to do recommended charge from this cuz I'm having problems using the maps one. Anyway, I see that it has the filled map here; I'm going to select okay, and I have all the data filled in. All right, with this visualization, we we can now dive in; we can see that we have a range of these median salaries from over 157,000 down to 30,000, with country like China having around 68,000, and then over in Africa we have Algeria at 45,000. So looks like we have a lower salary in the African continent, over in North America and also South America, pretty high salaries along with Australia as well. Anyway, pretty cool visualization we were able to generate out of this. I mean, I love data, and I just love this visual. A with this, I'm going to change the title to what are top paying countries.
Now, the last thing is a minor point. Sometimes if you're going ahead and actually moving maybe columns around, you'll notice that my visualization is also moving as well, and this can wreak havoc especially whenever you've made your dash or made your chart a certain size and then move columns around and it messes everything up. We can fix this, so I'm going to go ahead and control Z both of those column moves to get it back to where I had previously, and then from there I'm just going to double click on the chart itself, go under chart options, and once again this like resizing one here and going under properties. Right now it's selected under move and size with cells; we don't want to do that; basically we don't want to move or size with the cells, so I'm going to select that. Now, closing out of this, whenever I go to adjust the column size, it's not going to adjust the visualization at all. This is much more of what I want. Also, one last note on this, I do do have a filter currently applied to this data set, specifically I go into it; it's a custom filter, and I wanted to make sure that I had basically removed any NA values, so I put hey, I want values that are median Sal greater than zero and are less than 200,000. So if I go ahead and clear this filter, we can see that we have some other values up here, basically rushes up here at 300,000 for a median salary, and if we actually go in investigate Russia, we'll see that they only have around four jobs with salary data listed, so I feel like this salary is more of an outlier than anything, so that's why I'm applying this filter of 0 to 200,000. Applying this filter again, we get final visualization. Now you could also play around with this and filter it based on the number of counts to make sure you have values that are above a certain count; that's also an option and probably maybe even a better option as well. All right, chch turn now to dive into those practice problems to try out some different Advanced visualizations and along with some Advanced customization. With that, in the next lesson we're going to be diving deeper into understanding how to use statistical analysis, specifically box and whisker charts and also histograms and how to read them. With that, see you in the next one.
This lesson is going to be focused on actually visualizing a lot of the things that or a lot of the functions that we used in that statistical functions lesson, where we're looking visually at things like the median and core tiles. Specifically, we're going to do a refresher on histograms; we've seen it a few time reality, but we're going to dive into further understanding how salaries are distributed, specifically for a target audience of data analyst in the United States. You can feel, feel free to do whoever you want, and then from there, based on the limitations of it only be able to visualize one job.
Title: We're Going to Shift Vex to Looking at Box and Whisker Charts
These are great at also showing statistical distributions like a histogram, but we can take it a step further and we compare different values; specifically, in this case, we're going to compare them across the different job titles on how they're distributed.
Now, box and whisker charts aren't probably a chart that you're familiar with, or most people are familiar with, so we're going to go through a review and understand and break them down to understand those concepts we talked about previously about median and quartiles and where they fall into this.
For this, we're going to be using the Charts Statistics workbook. Specifically, we're going to be starting in this Data Tab, and for all this, we're going to be analyzing salary data. In this video, we're going to be focusing specifically though on that yearly salary data.
So let's actually go back into breaking down how to read a histogram. We go back into Insert, Recommended Charts, and then from there select Histogram and insert. In the histogram, I don't like where it is right now; I'm actually going to move this chart into a new sheet.
Now, quick refresher on histograms: each one of these bars represents a count of values within a range. So, in this case, there's 920 values between the range of oh my gosh, so hard to read—75,000 to 81,000. And as we're noting by this, we have a large number over here; if it gets even out to 960,000, this would be called a skewed right distribution.
Now, this is different from a column chart because this data down here on the x-axis is basically continuous data. When one bin stops, so this first bin of 15,000 to 21,000, the next bin picks up.
The first problem with this histogram is this is for all salary data, specifically all job titles across all countries. I want to actually fine-tune to look at my specific use case of Data Analyst in the United States. So you can come here into the Histogram 2 Tab, and I have the four columns of interest that I want to use from the Data Tab. And I already have the filters applied, but if you want to, you can come in here and actually select to clear these filters, and I'll just select it here from that Home tab.
Then, from there, I'm going to go through and select Data Analyst roles that are full-time only, that are in the United States, and then finally, I don't want any of these blank values here, so I'm going to uncheck this value here for blanks.
Now, we'll say filtering this data did take some time to actually do, so don't be alarmed if this taken more than 10 or 15 seconds. All right, so back in—let's actually make a histogram with this data. We'll go into Insert; from here, I'm going to insert in a histogram.
Now, once again, this distribution is, so the last one, skewed right, and we have a heavy amount of outliers right here, even out this one value around 370,000. I don't think this provides a lot of value; instead, I want to actually focus more into these, this actual distribution, and not actually on this portion out here that we have, just outliers anyway. I'm going to come in here into our filters up here, insert a number filter, and that it's less than 300,000. Click okay.
All right, this is looking a lot more readable, which we can actually see now. The x-axis—now each one of these bars right here, or what, what you would see in like a column chart, are called the bins, and they're all equally spaced, but we can control the width of each one of those bins that they encompass. Specifically, I can double-click on the chart to bring up that pane to the right, selecting the x-axis. I can then go into Axis Options, and then once again, Axis Options. We can go into something—right now we're noticing that the bins are automatically determined; we can actually change this bin width. I'm going to change this something to like 15,000. Notice that it is bigger; in this case, the bins are bigger than they were previously. You can feel free to test different options if you will. I feel if you go too small, in the case, let's say we went down to 1,000, it just gets too noisy, and also you can't necessarily see the distribution as well. So really, you just have to play around with it until you get to what you want to find as far as the access goes.
This is a little bit—this is sensory overload for me—way too many zeros in here, so I'm going to move this—selecting the x-axis, we can see that has Format Axis. Now I can go under Number, and once again, we can go in our custom type—none of the ones that I've previously done are here; sometimes it pops up, sometimes it doesn't. We're going to go ahead and just put in—we want the dollar sign, zero, and then formatted with the K value, basically removing all those uh thousands zeros, and I'm going to go ahead and click Add.
All right, this is a lot more readable to actually see what those different ranges are, and from there, I'm going to change the title to "How Much Do Data Analysts in the United States Make?" Probably also best practice here to add a title on the y-axis for "Count of Jobs," and bam. Now we have this final visualization shown on our histogram. We can see that a lot of the salaries are more around the range of 85,000 to 100,000, which 70,000–85,000 is coming up next. So this shows really visually great and at where I can expect to have a salary as a starting Data Analyst.
But now what if we want to analyze multiple different job titles? Which we're eventually going to get to is this box plot here, where we're plotting it for all the different job titles. We'll be able to actually compare different values across each other, but before we get to that, we need to first understand how to read a box plot. Also, sometimes I call it a box plot, but it's also known as a box and whiskers chart.
Anyway, I made this visualization here—you don't have to do it; there's a bunch of customization along with it. The main purpose of this is to demonstrate or help understand how to read a box and whiskers chart. So I took our data that we previously were analyzing for Data Analyst in the United States; it was a full-time role along with all the salary data, and then I used, like we previously did, calculating things like the Min, first quartile, median, average, third quartile, and Max. Just ignore this portion right here; it was used to make, build this visualization right here. Anyway, I tried as best as possible to line up this histogram where we have the x-axis going from 25,000 to 285,000 with the box and whiskers chart I made below it from 25,000 to 285,000.
So the box itself signifies what that nerds call the interquartile range—basically all the values between Q1 or quartile 1 and quartile 3. Had a typo there, got to fix that. Anyway, that's why it was so important that previously we calculated that first quartile and third quartile, and if you remember from that, there quartiles—so 50% of the data falls within this box. And if we look up, we were to draw imaginary lines into our histogram, we can see that about 50% of the data does fall within this. The next up inside of here is a line that is for the median; in this case, our median is 99,000, and then we have our average of 90,500, which, as we discussed previously, the average is going to be higher here because we have things all the way out here called outliers, basically dragging that average higher. And outliers are signified by these dots outside of the whiskers themselves. These whiskers are the lines, and the lines themselves extend to the minimum and the maximum, and these are just relative mins and maxes; they're not necessarily the true min and max.
Anyway, so that's a box and whisker chart, and frankly, by themselves, I don't think they're really great, but when you pair them with other categorical values, I find them super interesting. So let's actually build this visualization. So you can come over to this Box Plot 2 Tab, and I have our data inside of it; none of it is filtered; it has all the different job titles and all their associated salaries. For this, I'm going to select column M, and then also holding Control, I'm going to select column A. Then, from there, go in and Insert, and go to Recommended, and from there look at the box and whiskers chart, which looks like it's already pulling it up for us. So let's pop this bad boy in.
Now, one drawback of these box and whisker charts in Excel is, unlike that last box plot that I made—I custom-made this in order to make it appear in this horizontal fashion—you can actually do that; you can only have the option to have them vertical, up and down. Anyway, this is pretty close to what we want to get. The main problem I'm noticing right now is we have outliers up to 1.2 million, and it's really with the data around 100–150,000; it's really hard to actually look into those boxes. So I'm going to change this y-value scale; double-clicking on the y-axis, I'm going to change the maximum to 300,000. Additionally, since we're here, I'm going to change that number formatting to use that 0K value. Then also, I'm finding the color is a little hard to actually see these x's in here, so under Series Option, selecting Fill in Line Fill, I'm going to change this color to more of a lighter blue.
Okay, and that's definitely easier to read. I'm going to add a vertical access of Salary USD; I'm also going to bold it all to make it a little bit more readable, and then from there change that chart title to "What Are the Top Paying Jobs in Data Science?"
All right, getting into actually analyzing this and getting insights from it now. One drawback out of this is there's not an easy way to sort these values right here. Right now, I'd normally put them high to low; I'd probably put them high to low based on median salary, but they've been put into this graph based on the order that they first appear over here in column A, and that's when they pop up, so that's the order. So technically, I could go through and sort this column alphabetically, but that's going to take a little bit too much time. If you want to do that, feel free to try that out.
Anyway, it looks like roles like Machine Learning Engineers and also Software Engineers have a pretty large interquartile range, or that where that 50% of that data falls. So there's a basically a wide range of data or salaries you could find with that, whereas Data Nerds, Data Scientists, Data Analysts, and Data Engineers have a tighter band. Also, as expected, those Data Analysts and Business Analysts have some of the lowest median salaries, where something like the Data Engineers and the senior roles have even higher median salaries overall. This is pretty great at going in comparing values; I would probably work with this more to fine-tune it to only have a couple of job titles in it, and for that we can use something like slicers, which will be covering in an upcoming chapter—well, the next chapter—when we get into Advanced Techniques in Excel. So we'll be able to customize this further once you have that knowledge.
All right, you now have some practice problems to go through and get more familiar with those histograms and also box and whisker charts. In the next lesson, which is a quick one, we're going to be moving into sparklines, which is the final lesson in this chart overview. With that, I'll see you in the next one.
Moving into this last lesson on charts, focusing on sparklines. Sparklines are basically ways to insert mini charts into a cell that summarizes data that's next to it. If your data is coming in a horizontal form, similar to this table, you probably have the possibility of considering inserting a sparkline. We're going to be going through how to make them, but also customizing it.
All right, for this, we're going to be using the Spark Lines workbook. For this, we have, like usual, our Data Tab and then our Original tab that calculates data off it. And for this data set, we're just looking at what are the counts of the different job titles based on month. So this is basically horizontally oriented; this is great for a sparkline.
So how we're going to do this? Well, we'll go ahead and select the data only, so C4 to N10, then come up into the Insert tab. Then right here, we have this section on sparklines; we can insert a line, column, or a win/loss. We'll just start with column to start with, and it fills in for the data range C4 to 10, but it wants us to choose where you want the sparklines to be placed, so the location range, and click this arrow here, and then from there actually drag it next to it. All, close this arrow back and click okay.
Anyway, I wanted to demonstrate that bar chart because it's not really that great for here. Remember, anytime we're doing continuous data—in this case, we're doing that monthly data—I'm going to want to use something like a line chart instead. So I can easily change it by coming up here, selecting all of our different data, selecting that Spark Line tab, and then just changing it to—I can change something like win/loss, which—no, really data from this line chart—that's what we really want from this.
Now, getting into the customization of this—I really personally, I'm like blue, so we're going to stick with the blue color, but we could change the color if we want to. And the other thing we change is the marker color. Right now, we don't have any markers on it; we can actually change which markers are right here in the Show selection right here. So I can select the high points; right now it's going to highlight all of them red; uh, low point also red; negative points—there's no negative points. You also do the first point, which I don't really find much value in that, or last point, and then actual finally the markers itself; you just put every single one of them with a marker. I really like this high point and this low point, and we can customize this. The high points, I would really want to call out to be a green color. Right now, this green that's sort of hard to see, so I'm going to change it to something a little bit darker, and bam, we can see that one a little better. The red for the low point, I'm going to keep it as is. And the last thing is all this data has basically a grid around it; I'm just going to add that in real quick by selecting all the cells, come up into Home into the borders; I'm going to put in all borders around it. Then it looks like I have a double line right here for this lower one, so I'll insert this bottom double border, and then finally I'm going to put a thick border around this all. Bam, we have our final visualization there. Now I can go through and see things like, okay, with that Analyst and Other Analyst, we saw spikes in January, but things like that a Engineers, we didn't see a spike. However, all the job titles ran to a similar problem where apparently they ran out of budget, and the least amount of jobs were posted in November and December. So this is a pretty cool feature to show some quick snapshots about the data you're looking at.
Right, you now have some practice problems to go through and basically practice making some of these sparklines. We're going to next be jumping in the next chapter; it's our final chapter of the basic section, and it's going to be focusing on Advanced features inside spreadsheets, such as tables, formatting, and how to collaborate with others. It's our last section before we build our first project, so with that, I'll see you in the next chapter on Advanced Spreadsheets.
Then nerds, welcome to this last chapter in the basic section, focusing on Advanced features and spreadsheets. There's a last chapter we're going to be covering before we get into our first project, and this chapter is broken into three different lessons. This one right here is going to be on tables—how to use tables, how to use things like slicers, and how to manipulate them. The second lesson is on formatting—not just on making cells look pretty, but developing conditional formatting rules in order to highlight cells according to, well, a certain rule—pretty interesting feature within Excel. And the third lesson is on collaboration. For a project, we're going to be making a dashboard, and so we need to enact certain measures in order to protect it and prevent people from going in and messing it up, and so we're going to go over a lot of features in order to set it up properly.
Anyway, back to this lesson—what are we going to be doing for it? Well, first, we're going to start out by using a smaller subset of our data set—basically 15 rows—and creating your first table. We're going to be manipulating it using custom formulas that we really haven't seen before, along with using some other ones that we have seen before in order to calculate totals, subtotals, and aggregates. By the end of this lesson, we're going to be building a mini dashboard to analyze that histogram that we talked about in our previous lessons; specifically, we're going to add slicers to it in order to be able to filter down and look at a subset of data that we're most interested about. And that's all could be done without the help of tables.
For this lesson, we're going to be using the Tables workbook in Chapter 4. For this, you're going to start in the Tables Intro Original sheet, and then the final one's going to be what we're going to eventually get to. All these are going to be labeled similarly with the Original and Final, and we're all going to be working with the Original; it should look like the Final when you get done with this. So let's dive into creating our first table. First thing you have to do is make sure that we're selected somewhere in here; we don't necessarily need to select the full table, but just somewhere in here. From there, we'll go into the Insert Tab, and we'll insert a table. Also notice that we can use the shortcut Control + T for this, so I'm going to do that instead. And for this, it automatically pinpoints the rightmost cell and the bottom-most cell, and we need to make sure we have this check mark enabled of "My table has headers" because we have, well, call them headers, and bam, we just made our first table. This lesson's over, but seriously, let's actually get into exploring this Table Design tab that now appears anytime you're selected to the table. If I click off of it, it disappears. Anyway, we're going to first look at the table name, and I like to have a table name that's easy to reference, so I'm going to just name it something like "Jobs"; it's going to come in handy naming it something simple whenever we're making formulas later for this. Now we'll get to this section in a little bit on tool and external table data, but I want to move over to the Style Options. You can play around with some of these options here where you can highlight the First Column or you can highlight the last column; has a lot of different formatting options with it, but what I really like is this color formatting. If I'm not really liking the color that it's given to me, just come over here, select a new one.
So we'll get back to Table Design in a bit, but what's really the benefit of this table? Well, one thing is you can easily add data to a table, and it will, will autofill. Let me show you. Let's say I wanted to add a new column with a "Solid Year Average Copy." Whenever I enter this new column name and press Enter, it automatically fills this in. I can—the skills are sort of covering this up right now, sorry about that—and I can make this a little bit bigger, but you can see we have Salary or Average Copy now included within this table, and I can verify that it's included also in this table by—if I want to go to Resize Table—it will say that now it goes to K16. Now for this, I just want to copy the results of the Salary Year Average column over here in H. So what I'm going to do is press equals to, and I'm just going to select the cell over here of H2. Now this is what I was talking about whenever I said tables have their own unique formulas. What it's going and doing here is it's referencing the Salary or Average column, which is this portion right here, and then it's also using this at symbol to basically refer to this is the same point in the row of H2 that is a K2.
When I go ahead and press enter, watch what happens. We actually fill in all the different values of this. So, if I were to actually double click into this one down here, we still have that same syntax of: we're selecting the "Sal your average" column and we're using that at value value to get the one that corresponds in that same row.
Now let's dive deeper into these different formulas we can use for this table. So I'm going to come over here into column N, and for this, remember we named our table "jobs," so I'm just going to type in "jobs." I have two tables in here, one called "job," one "jobs." You only have one popping in here anyway; it automatically pops up. So I'm going to select "jobs," and now whenever I do this, I'm going to press enter. It's using our modern dynamic arrays basically to fill in all the data that we have over here inside of our table. So pretty unique in how we can reference this.
Now, what happens if we wanted to also include the column headers up at the top? Well, I can type in "jobs," and then from there I'm going to add a square bracket, and we have a few options popping up right now. It looks like it's just column titles, but if we scroll down, we have these values here with hashtags in it. Specifically, I want the ones with the column headers, so I'm going to put "#headers," I'm going to put a close bracket on this, and then press enter. And now we have the column headers across the top.
Now that's a little bit too much work having to do two different formulas for this. If instead I wanted to do "job" and then square bracket and see the options available, I can see I have "all," "data only," "headers," and "totals row." "Totals row," we're going to get to a little bit, so we'll do the "all" for now. And if I go ahead and press enter, bam, we now have our data with our column headers and also the data itself.
But what happens if you want to just access certain columns? Well, I thought you never asked that. Well, once again, I can type in something like "jobs," but the square bracket, and then we have a list of different columns available. Let's do the "salary year average" and do a close bracket. Once again, this is going to provide the data values only. If we wanted to include the specific header for this, I once again need to put in "jobs," and this time I'm going to need to specify not only the headers, so I need to put this in its own square brackets, but I'm also going to have to do a comma, put another square brackets, and put "salary year average" within its own brackets. So it's almost like a list of items. If you're familiar with Python, this would be like a list. Anyway, we have the headers in brackets, and we have "salary year average" in brackets. Pressing enter, we get "salary year average" up at the top.
Now, honestly, an easier way to do this all is to, well, use that "all" command, or "#all," but it has to be put within its own square brackets. Then, from there, a comma, and then we want to say, hey, the subset only that we're providing for this is "salary year average." Close that bracket, and then close the entire brackets for "jobs." Now, from there, when we run it, we get the "Sal year average" along with all the column values. At any time, if you forget that, it's not that big of a deal, as you can just go through and put an equal sign, and like we did previously, I could just highlight, well, not that, um, our "salary year average" column, and look, it automatically populates with that same formula above here. And when I press enter, boom, it pops up there. So don't think you have to memorize these formulas that I just went over.
But what do all these formulas actually provide any value value for? Well, let's look at a use case. Let's say I wanted to identify jobs that, whenever we looked at the skills, we could find out if they contained the skill of "Excel" or not. So I'm going to create this new column over here and call it "Excel," and for this we're going to be using the search function, which we need to provide what text we want to actually find. Conveniently, I put it in the column header, so I'll go ahead and just select it, and it automatically populates the formula for this. Then, from that, we need to go to the next parameter of within text; we're trying to look at that "job skills" column. It puts that "@" symbol at the front of "job skills" to basically signify look at that row. Then, from there, I'm going to go ahead and close the parentheses and press enter. So for that search function, it provides the N numerical location of "Excel" in here. "Excel" is 36 characters deep into this. So I'm just going to modify this because I don't really care about the number of that. I'm going to say I'm going to use the "is number" function, which checks if it's a number and then returns true or false. In this case, we have True Values, so we know that for these columns, if they contain "Excel" or not, they'll have true. So that's how I find myself using these different formulas and understanding how to actually manipulate them.
Anyway, let's get into our next step. Let's say we wanted to include some sort of totals row in order to maybe calculate median salary, how many job postings there were, etc. So we'll go into this table design tab, and I'm going going to select the total row. And now down here in row 17, we have "total" written down here along with a bunch of, well, blank values, except for all the way to the right. Looks like it puts us the number of 15, which is the total of these. Now going over to that "salary year average" column, I can basically select this totals row right here, and you notice a drop down appears right here. From here, we can select some basic statistics: average, count, min, max, variance. Go ahead and select average; that's the average of this column right here. So pretty neat. I'd go through and if I wanted to do other columns as well that. Now you can also go into here and select more functions, and then like we said, we want to calculate median on this salary. We could go ahead and select this function of median, but I'm actually going to recommend another approach. You see, if we double click inside of here, we actually see that this totals column is using a function, specifically the "subtotal" function function. So let's actually build this out from scratch without selecting it. Luckily, we have the "salary year average copy" column over here, so I'm going in, and I'm going to type in "subtotal," and it returns a "subtotal" in a list or database. First is the function number; what do we want it to actually do? And this has even more values available to it that you can actually select from and perform on this. So in this case, let's say I wanted to find out what the max value is; I would plug this in; it would be 104. And then for the reference for this, well, we're just going to select this "salary year average copy" column. It automatically transformed into this special syntax and then add a closing parenthesis and press enter. And so now we have the max salary, which looking at this, it's true. But if we go back into this and actually inspect what values are available in this function number, we can see that median is not available in here. So what are we going to do? Well, there's another function we're not going to use median, but that I recommend instead of using "subtotal," and for this one we're going to use the "aggregate" function, and this returns an aggregate in a list or database. It's similarly designed where it has a function number, but with this one we have a lot more options, including things like CTO and stuff like that. Anyway, it has median available as number 12. Now the second parameter on options allows us to select a host of options, uh, no pun intended, for allowing us how we want to actually perform this aggregate. Basically, do we want to maybe ignore hidden rows, or do we want to ignore error values? In my case, I don't really want to ignore anything, so I'm just going to do number four. And then finally, we need to insert the array or the column itself. In this case, we want "salary year average," closing the parentheses on this and pressing enter. We get our median value of 94,000.
Now, depending how fast your computer is, you're going to run into some limitations here. I have in the "table limits original" tab, which is the next one we're going to be working with in this uh portion of the lesson, it has around, well, 32,000, which is in the data set. Anyway, we're going to run into some limitations, as I'm going to show. I'm going to encourage you to just watch along uh me do this, and then from there basically decide if you think you have a strong enough computer or not to continue on to do this. Um, but if you have a pretty uh basically slow computer, I wouldn't necessarily follow along with this. Anyway, I'm going to convert this into a table by selecting any portion in here, pressing Ctrl+T. It selected all the different values, and that table has CS. So now we've converted this into a table, and one of the benefits we haven't really discussed yet is the ability to actually filter data because it automatically provides this filter up at the top. Now I'm going to go ahead and filter this down based on a data analyst job title, and when I go through and actually select this to just select it at analyst and press okay, it runs pretty quickly, but I have run into problems in the past, especially working with smaller computers where it takes a while to do this. I'm working with about 24 GB of RAM on this virtual machine, so if you're something at like 8 or even 4, I'm going to highly recommend that you may not perform this exercise.
So moving to this last exercise of this lesson, I've gone ahead and condensed down this data set. You can go into "histogram original" and our previous data set. I basically shorn it down to these four columns and limited to only positions that have a "salary year average" value listed. Basically, if there's blanks, I remove those rows, so it's about 208,000 rows. Anyway, this is what we're going to be manipulating for this. This shouldn't lock up your computer if you have a basically a computer with less RAM, and we're going to convert this into a table first, pressing Ctrl+T. I select all the values on here and press okay. So now we have a title. Now also in this sheet, you may have noticed hopefully that it's been on the screen, I have this histogram here, which is basically aggregating the data from this Delta column on "salary year average." Anyway, we're going to be manipulating this further. We want to basically make this into a dashboard so we can go through and maybe filter for different job title, different job schedule types, or different job countries, and it can be mildly inconvenient to come up here and actually select this arrow and then go through and select the values you want. That's why slicers are great. So with our table selected, I'm going to go into table design, and then from there under Tools, I'm going to go to insert slicer. We're going to be entering in both a job tile short, job schedule type, and a job country slicer, so all three are here. Now I'm going to go ahead and position them, make them look a lot neater. All right, got them cleaned up, and then from there I can go ahead and actually select the slicer, sir. And if you notice, this slicer tab pops up, conveniently labeled. This slicer has a caption on it or a title as well, and I can just rename it basically to a better visually appealing title. In this case, I want it to call "job title," and then it updates here for "job title." I'm going to do the same for the other two, updating it to "schedule type," and then also "Country." Now by default, this slicer and all the slicers have all the value selected. So if I wanted to to go in to actually select a value, I could do something like, well, we want to look at "data analyst." I just select "data analyst"; it's going to clear all those other ones and then only select that analyst. As you notice, it took a second for it to actually load; that's why with this 20,000 rows of data, even that's a little high for tables. I recommend it around 10,000 if you're using tables. Anyway, we have it filtered down to "data analyst." I could also do it down to "full-time" along with filtering it for U, basically uh, I want to do "United States." If you notice these values are grayed out, that means there's no country basically available with the current selections that I have of "data analyst" in "full-time." So that's what that means there, but I can go into that for "United States," selecting it, and bam, we now have our final basically visualization. But what happens if I want to maybe look at multiple different values? What if I maybe want to look at both "data analyst" and "business analyst"? Well, in that case, you want to select this box up here, and it allows multi-select, and so I enable it, and now I can go through and select something like "business analyst," and this provides both those values along with I wanted to look at "full-time" and also "part-time." I could enable the multi-select on this "schedule type" and select "part-time," and bam, now we have multiple values selected for this along with the "United States," and this makes the dashboards that you're building a lot more interactive and a little bit fun to play around and to visualize the different data.
All right, we have some practice problems for you now to go through and dive into not only creating tables, manipulating them, but also adding and playing with slicers as well. With that, we'll see you in the next lesson. We're going to be jumping into formatting, specifically conditional formatting, so see you there.
In this lesson, we're going to be focusing on formatting, and not just self-formatting where we're going through and adding borders and colors, but also conditional formatting where a cell's basically formatting, highlighting, will update dynamically based on a value. In the first example, we're going to focus on cell formatting; specifically, we're going to go back to that table that we've worked with previously that does a count of data science jobs over the month. Anyway, we're going to go through and actually format it using all the different functions we can in order to make it look pretty, like I made it. From there, we're going to move into our first conditional formatting example where we're going to look at basically highlighting based on a job title, those that are basically high and those that are low, highlighting them appropriately green or red. And then in our final example, we're going to move on besides using color scales to also using things like data bars and also icon sets to make it look a lot more dynamic. We're also going to go over best practices on what not to do, because sometimes you can go overboard in how much you're actually coloring a table, and you can make it a little distracting and, and ultimately, not meet your goal.
For this lesson, we'll be working with our formatting notebook in chapter 4, as usual. All the data is located in the little data tab, and we'll be starting with the _original of each of these sheets, and then it, we'll get to, in this case, "format original." We'll have what it looks like; "format final" for the cell formatting, we're going to be using this "format original" sheet, and we're going to be focused on this Home tab here. So I'm actually going to leave it expanded, and for this we're going to make this to where, well, what this table looks like by going through and actually formatting using all the different features in here. So the first thing we need to do is highlight it all and actually remove the formatting. So with it all selected, I can go to editing and then clear, and I can either clear all, which is what I don't want to do; I want to do clear format, and bam, now we have an ugly table that doesn't really make a lot of sense. Now previously we were mess with tables, so I could highlight from B3 to O10 and make this into a table by coming up here to format as table, basically selecting the color that I want, saying that it has headers, and allowing it to update. There's definitely an option, um, but I'm not necessarily a fan of this, so I'm going to clear this by pressing Ctrl+Z.
Now, an underused feature of formatting is this cell Styles tab right here. So I'm going to go ahead and select the months up here, basically the titles, and for cell Styles, they actually have a lot of pretty unique formatting you can see happening in the background. So I'm going to try out, in this case, I'm going to try out "heading 2," which is pretty neat because it makes it bold, slight bigger, and it puts a little line underneath it. I could do something also where I highlight all the rows over here and then make this into maybe "heading 3," and then all these values in here are calculations, so technically I could just highlight this all, and for the cell Styles, I could come up to the top here and select, hey, this is a calculation, and this not a bad looking table, uh, but not necessarily all I want to do. So I'm going to just remove this all. Instead, I'm going to start with my months; I'm going to make them bold and also add a light gray background. I'm going to do the same thing over here for the values in my rows, and then from here we're going to get the actual column grid lines put in. I'm going to only select C3 all the way down to O10; I'm going to show you why, and I'm going to add an all borders. So this is NE; it adds it adds all borders to it. What I'm going to also add this, which will add a little bit of flare to it, is a thick outside border. So now we got a thick outside border around all of this, and I'm going to do the same with this one of an all borders and then a thick outside border. Now it did remove that thick outside border that I had on this line between B and C, so I'm actually going to go ahead and put that back in by just clicking it. Next thing I want to do is format these with a comma, so I'm going to come up here and, well, add a comma, and then unfortunately it adds this space in here and makes this table bigger than what you can see. Now I'm going to first remove the decimal places, and then in order order to fix this, I'm going to highlight all the different columns through here to January and just double click on one of them to make them slightly smaller. Anyway, it's still not fitting completely on here, and I want this to fit within the view here, so I'm just going to select this all, and I'm actually going to make these values slightly smaller. And I'm not liking the positioning of these; it looks like it's lower now that I made this smaller, so I'm going to actually center this, this do a middle align, basically move it up slightly. All right, my OCD is now looking good. All right, now this is looking good. Now the last thing we want to do is add a title to this, basically describe what is this table that we're looking at, and I want to insert this in up on the top row, but I basically want it centered over this table. So what I can do is highlight from B1 to B11, and from there select up here for merge, and also I want to center because that's I want my text center during this, and from there I put in, hey, this is the "Data Science Job Count Tracker," and for the cell style, I'll make this "heading 1."
Now let's get into conditionally formatting this table, and specifically I want to say if I'm looking at "data analyst," I want to be able to look across here and see which ones are the highs and the lows. Right now I have this grid lines, and I can see that based on the green and red or the highs and lows, but I want to actually be able to see this in this table right here. And so underneath the Home tab, we have this conditional formatting available. We're going to focus on these three right here first, and that is data bars, and you can see if I put it in, it's basically looking like a, you know, like a bar chart. Color scales allows us to do, well, different
Color formatting with it and then an icon set basically allows us to put in a nice-looking icon. And we're going to stick simple for now. We're going to do color scales. Right now, I have C4 through N4 selected. I'm going to go ahead and select this green to red, which is not bad if we're looking at this right. This is doing exactly what I want. I want August, which is the highest, to be highlighted green to attract my eyes to it. And then I want the red to be November and December because I want to attract attention to it. But we want to highlight the entire table here. So if I were to actually select the entire table, if you will, from C4 all the way down to N10, go into conditional formatting color scales and do the same thing, you're going to notice it basically does these bands, but it does this entire table all formatted together. And this is not what we necessarily want. Of course, the total row is going to be the highest. I want to look through that row and actually see where I should be actually looking.
So anytime we need a clear mess with any rules, we come into conditional formatting and go to clear rules. You have clear from selected cells or entire sheet. We're just going to do the entire sheet. Then we're going to go back to where we were before, of selecting just the data analyst values, going into conditional formatting color scales, and I'm going to go to this green white red. I actually want to try to limit as many colors as I do; two is enough. So I'm going to go green white red. And I really like this one better. Now I don't need to necessarily go through once again of selecting senior data analyst doing this again. What I would do instead is I'm going to select data analyst here and then come into this home menu up here, and you notice this paintbrush. This is a format painter. In the instructions, it basically says select the content with the format you like, click format painter, and then select something else to automatically apply the formatting. So from here, I can just paint my formatting on. Unfortunately, this doesn't have a shortcut, so I have to go do go back up every single time it removes their marching ants, reselect the format painter, and go through and select it. But now we have this formatted how I want it, where I can look at a certain row; in this case, I look at data analyst, see what some of the highest are, and senior data engineers. I can see how they contrast to the other job titles.
Additionally, which we're going to jump into a little bit more later, is we can go into manage rules and we can see the current conditional formatting applied. Right now, I have show formatting rules for current selection. I'm selected the top cell right up here, so there's no conditional formatting. If I were to change this to just this worksheet, I can then, if I expand this down, I can see how this applies; this type of formatting of the red white green applies to each of the different cells. And if I needed to actually control what cells are actually selected, I could do that. I could have also gone through instead of done that copy formatting and pasting; I could have done a duplicate rule and modifying the code as well, but I decided to do my way instead. Anyway, this is where you need to go if anytime you need to manage conditional formatting. We click okay.
Let's crank this up a notch and get into using some more advanced functionality with conditional formatting. Here we have a new table you haven't seen before. Basically, it has all the different job titles, the counts of those jobs aggregated from our data sheet, the median salary, what is their work from home percentage or likelihood based on the jobs, and then finally, I have this job rank right here, which basically uses these cells that are hidden right here that, if we actually expand it out, goes through and normalizes the values. So in this case, the job count normalized it between zero and one, so this job count is 90 is the highest, so it gets a value of one, where it's the lowest gets a value of zero. Anyway, I did this for all the different values, and then from there, provide a certain weighting factor of like 0.453 and 0.15 in order to wait it appropriately. This is all my bias and how I wanted to actually do it, so feel free to adjust it to what you want. Anyway, we have this final job rank in order to assess based on these three values. And this is commonly done, especially in like KPIs and stuff like that. So we're going to be making like icons for this column. So let's get into formatting our first column. We're going to do job count first, and for this one, I want to have a data bar. So I'm going to come down into conditional formatting into data bars, and we'll add these data bars right here. I like the bars in this case because we're dealing with a count, and we can really see, especially data analysts, scientists, engineers, they really make up the majority of the data here, so it really draws your attention to it.
Next up is a median salary. We're going to do similar to last time, maintain a color scale. We're just going to do this first one right here, where green is the highest salary and red is the lowest. And then one more, we're going to do that work from home. We're also going to do it in a color scale, but for this one, let's actually do a different color. Go into more rules, and in this case, we have this new formatting rule window right here. I have two colors. Just say I want to do one color; I'm going to do white from the lowest value, and then we'll do like purple for the highest value. Anyway, this is all basically to show a point; this is becoming entirely, entirely too much visually distracting. If you were to give this to somebody else or a stakeholder, where are they supposed to look and actually organize their thoughts on where they should potentially pursue a job? Right now, I'm thoroughly confused at looking at this, so let's clean this up a bit. And for this, I want to make it to where I like maintaining a solid color across, so that way you know, like, hey, if this color is darker or there's more of this color, I should be looking there. So in this case, we'll make this job count; we're going to just clean it up slightly. For the data bars, we're going to make this like gradient appearance, because then I feel we can see the numbers better, and it's not too visually distracting. For the median salary, I really; my goal of this is to find jobs that are, let's say, greater than 100,000. So let's actually just make highlighting that highlights those jobs that are greater than this value. In this case, I'm going to come to conditional formatting and enter a new rule. This new formatting rule popup comes up again once again, and we have a select a rule type. This allows us to do things like format all cells based on the value, format only top or bottom rank values, format only values that are above or below average. I personally like this one: use a formula to determine which cells to format. And in this case, I want to say I'm going to collect this formula thing right here. I want to look at; you can just select the first item in the item selected, so I'm selecting D3. It's going to go through and actually do all of these; don't worry, we'll see. And for that, we want to highlight those that are greater than 100,000 and press enter. And then right now, it doesn't have any format set, so I'm going to change this to format, and we can control a whole host of things such as the fill, border, font, and the number formatting itself, but we're going to stick with that blue theme. I'm going to just come down in here, and I'm going to just select this blue color right here and click okay, and then okay again. Now you notice my formatting is not appearing. That's because we have multiple formatting applied to a cell, which you can do. So in order to fix this, we need to come into manage rules, and as we see, we have both of these applied to it. So I actually need to select this one, and I need to delete this rule and click apply, and then okay. Now we're running into our second issue, and I slightly misled you earlier when I said that D3 works. If we go back into manage our rules and we see our formula right here, I'm going to double click it. We don't need to actually provide an absolute reference to D3 because it's actually going to evaluate all those cells based on D3 instead. We want it to be D3 without the dollar sign, so it's not an absolute reference. And therefore, whenever I click okay and okay again, bam, now it knows appropriately to check the actual cell that it's looking at within the range on whether to highlight it or not.
Moving on to the work from home, we're going to keep this similar in that it's not going to be purple though; we're going to change this to blue instead. So going into manage rules, we have the actual color right here selected. I'm going to just go in and change this to this color that we used previously and click okay, and then okay as well, so that way it applies it all right. The last thing is this: the job rank itself. And for this, we're going to be using icon set. Specifically, I like this one over here on ratings, but this becomes a little bit overwhelming when where we have this rating and also the number next to it. So we can actually remove this number in the column. We go back into manage rules; we can double click on that icon set rule, and we can even further customize when these stars are appearing, but I'm going to just go ahead and get to this portion where it says show icon only. This allows us to only show the value. So going into applying this, bam, it's now showing the icon. I want that icon centered both vertically and also horizontally, so bam. Now whenever I look at this, I can see, especially since it's all one color, my eyes really gravitate to well, data scientists and data engineers based on this full star rating and more of the blue being in this region, and that's what I would hope people would go to or gravitate to as well when they're looking at it. One quick note: in this conditional format, we didn't cover this highlight cell rules where you highlight greater than or less than, or you do a top/bottom rule where you can highlight the top 10% or top 10; you can also adjust that number. Anyway, I find that myself more using custom rules instead by coming in here into new rule and then actually fine-tuning what I want to do. So with the practice problems, I'd really dive into actually relying on using these types of options instead. And so, as I subtly hinted to you, have some practice problems now to go through and really practice how to do formatting and also more specifically conditional formatting. In the next lesson, we're going to be moving into collaboration and covering how to actually protect your workbooks and your worksheets so that way whenever you share these with coworkers or friends, they don't go through and actually mess them up. All right, with that, I'll see you in the next one.
Welcome to this last lesson in spreadsheets advanced. Before we jump into our project, and this lesson itself is on collaboration, which sounds sort of cheesy, but in order to demonstrate what we're actually going to be learning in this lesson, we need to actually jump fast forward a little bit and jump into our project. So I'm going to open up the salary dashboard, which is located under Project One dashboard. So here's the dashboard that we're going to build in it. They have three boxes that you can go through and select. This is going to be using data validation, which we're going to be learning about in this lesson, but it allows you to basically standardize the inputs that we want somebody to actually select in order to get the results, and it prevents them from putting in values that maybe don't exist and then breaking our dashboard. So for each of these job titles, country, and types, we have an associated visualization for each, showing the salary by job title, the salary by region, and then also salary by job type. Finally, at the bottom, I have some; I call them KPI cards, basically outlining certain characteristics or certain indications of the median salary, what is the top job platform, and then what is a count of jobs. But I can come in here and select something like maybe I wanted to look at business analyst, and it's going to filter down based on this, telling me what their median salary is, that LinkedIn is probably the best place to go to for this, what are the different types of roles available, and what's available in the job database. So the other feature we're going to be going through besides this data validation process that we can do right here is actually protecting your sheets, which you can find this here underneath review under protect. But anyway, if you try to move these cells around, you're not able to at all. So we're going to be able to design this dashboard in a way that other coworkers won't be able to destroy it. Additionally, if you notice down here at the bottom, there's only one sheet in here; there's actually other sheets. If I go to unhide here, there's other sheets. I'll just unhide one of them; we'll just unhide data. There's other sheets inside of here, but if they're not applicable to my coworkers or stakeholders, I don't need to have them, so I can hide them. So that's the another feature we're going over in this. All right, nothing be yaen. Let's actually get into this lesson. For this, we're going to be using the collaboration workbook in chapter 4. Now we're going to be building out these three sheets as we go along, and as a sneak peek, in this first example, we're going to be building out this little portion right here. This is going to be basically preparing us for our project, so a lot of this work is going to be put to good use. Anyway, we're going to be building the simple one right here. I'm going to zoom in where we have based on the job title, we can go through and select it. So senior engineer, it's going to pop up with our median salary. So that's what we're going to be building with this, and specifically, we're going to be using this feature of data validation. So I'm going to create a new sheet to start with because I don't want to start with the answer right there. I'm just going to call it calculator. I'm going to put in job title here and then median salary below. I'm also going to bold these by pressing B, and then these are where next to it in column C is where we're actually going to use the actual control of this. Now we need to get a list of job titles to put in this. So I'm going to create a new sheet and call it validation, and basically what I'm going to do with this is create a sheet of all of the different job titles available. Specifically, I'm going to say this is going to be from the column job title short, and we're going to be using in order to get the unique values of it; well, the unique function, we need to provide it an array. So I'm going to come back over here down to column A2, use control shift, select all the way down, close the parenthesis, press enter. Okay, so now we have all of our different values. I'm going to expand this out. I'm also going to zoom in a little bit. Now whenever I do this drop-down menu, I want it in some sort of order. Specifically, I wanted it in probably what is the highest count value; I wanted it appearing at the top and those that are less likely down at the bottom. So what I'm going to do is actually just copy this value right here because this is what we actually want to use. What we want to do is a COUNTIFS. We want to count based on a condition. For the criteria range, we're going to be providing that job title short column from our table, and then for the criteria, we're going to be selecting right next to it, A2. There, we'll just autofill it all the way down. And then finally, we want to now sort it by this. So I'll use job title short sorted. From there, we'll use the SORT function to then sort this by the second column position in descending order. So bam, this is more like I want. I want those data analysts, data scientists, engineers at the top and the senior roles and so on, cloud engineers, car Bel. So we now have this list available that we want to use for data validation. We speak of; I'm going to go back to the calculator tab that I made, and for this, we're going to go to the data tab, specifically under data tools, they have this selection available where data validation actually is. And now this is going to allow us to, well, customize it. Right now, the data validation for this cell is any value; I can place any value into it. I could limit it to a whole number; I could limit it to decimals, a list, a date, a time, a bunch of things. We're going to limit it to basically a list of values, and we need to basically, so, provide a source for this. So for the source, we're going to go in and select the validation tab that we just made, and I'm going to select all the different jobs right here and then press enter. From here, I'm going to accept this and press okay. Now, as you can see, we have this little drop-down right next to it, and I have different selections actually available of data engineer. If I were to go into here because I have this set to data validation, if I was going to put in something like data nerd, which isn't available, and press enter, it says this value doesn't match the data validation restriction defined for this cell; therefore, I have to go in and retry. So only values within there are going to be able to work in this. So now let's actually get into calculating that median salary. And for this, we're going to create a new sheet, similar to this median salary sheet. We're going to call this one salary; wrong spot; need to actually enter it down here and call this one salary. Throw this all the way over first. I need the names of job title short and all that kind of good stuff. So what I'll do is I'll come over to our validation tab, and I've selected equal to already. I'm going to select these cells right here, press enter. So now they're all appearing here. Now I'm going to calculate the median salary for all these jobs. I know our calculator or dashboard has only one value that is calculating at a time, but in our dashboard we're going to build, we're actually going to build a graph with all these median salaries, so we just need to calculate them. Now all the median salaries and then basically calculate using data validation and also an XLOOKUP what the median salary is going to be here. So for this, we're going to be using the MEDIAN function, and specifically, we're going to be using that IF inside of it, because MEDIANIF isn't available. We first want to check: does the job title here of data analyst meet our condition of the job title short? So I'm going to type in the table itself of jobs and then the column of job title short, close bracket, and set an equal sign, equal to A2. Then I'm going to close the parentheses on this, and actually, we need to wrap all this in parentheses because we have to do multiple different conditions; we're going to do some array multiplication. The other thing we have to check is that the values are not blank or not equal to zero. So once again, I'll put in jobs again, and we're going to be using that salary year average column, and we want to make sure that it doesn't equal to zero. And so that's the condition we're checking for. And so now what do we want to return if true? Well, we want to return the salary, so we'll do jobs and then salary year average. I'll then close the brackets on that, then we need to close one parenthesis. I can see a red parenthesis still and then a final black parenthesis. NOS, I'm good. Press enter. Looks like I got it right on the first try. Let's actually drag this down. Boom, this is pretty nice. So now we have all the median salaries for these different job titles. I'm also going
To take this a step further, of actually sorting this by the med CER, because I know I'm going to be actually visualizing this in the Project's lesson, so we'll go ahead and sort this as well, sorting it on the second index in descending order. So now we need to provide the value; in this case, Data Engineers, there is selected. We need to provide, based on this value, the median salary, and I want to just calculate it over here, just in case I need to go back to it. So for this, I want basically 125,000 to here, right here in G2. So I'm going to provide an X lookup, and the first thing is this lookup value. Right, we're going to look up the data engineer in this. Now I'm not going to use a cell reference of going over here of selecting this cell of data engineer, which is calculator C2. I'm actually going to escape out of this; we're going to stop this right here. I want to go back to this. I actually instead, because I'm going to be referencing these cells specifically, well this what's right here a lot, I'm going to just rename this from C2 to title. So right now I can see that it is named title. So going back over to that salary tab again, now we can perform our X lookup. And for the lookup value, we're trying to look up the title. For the lookup array, we're looking up through this job titles right here, and then for a return array, the actual salary values. So now we're getting that data engineer value of 125,000. Similarly, I also want to name this cell as well. I'm going to name this one median salary, pressing enter, boom, locks it in. So now when I come back over to my calculator tab, I can just put in here equal to median salary. I'm also going to go through and format this to make this look better. So just playing around with this, I can see that I can put in something like senior data analyst and then a job, the associated Med and seller is going to come up with it. But let's say now I want to give this to a coworker. Right, how can I prevent them from going in and potentially, you know, entering in this cell and then breaking it? Well, we can come up here to review, and in this case, we're going to select this of protect sheet. Now the first thing you can do, you can set a password to unprotect sheet. I'm not going to put a password, but say you wanted to put one, you could. And then we have these options for for what you can actually protect, whether that's select lock cells or select unlock cells to protect. We're just going to leave both of these checked for the time being. Click okay, and now while one we can see that underneath protect here it now says, instead of protect sheet, it says unprotect sheet. Whenever I go through this and say I want to change it, any value whatsoever, I can't change it. So it's good because the numbers can't change or the median tile can't change, but now I can't change B job title, which is a little bit of a pain. So unfortunately, Excel doesn't necessarily make this the easiest. I'm going to start over again and just click unprotect sheet. And what we want to do is we're going to select all the cells in here. So with all the cells selected, I'm going to press control and unselect C2, then right-clicking it, I'm going to go into format cells. Now under this protection tab right here, we're going to notice we have options for locked and hidden. We want to actually be able to lock all the cells except for C2. We don't want to hide any, so we're not going to adjust that right now, but now we're going to have the ability to adjust whether it's locked or not. This doesn't actually change anything right now. So if I go into here, yes, I locked those certain cells, but if I were to type into here, it's still going to allow it to be changed. So now what I can do is go into protect sheet, and previously we had both of these selected, of Select lock cells and select unlock cells, and in this case, because we locked all the cells except for C2, we only want to allow people to select the unlocked cell of C2. So I'm going to uncheck this, click okay, and now I can't click anywhere else except for where I've set up that data validation in this cell, and I can still change it, and it will manipulate the value. Now we could also go through and protect the workbook itself. I don't necessarily manipulate with this as much. Instead, what I would want to do in in this case is actually hide all these other sheets with the exception of this calculator, and so I can do this by right-clicking a tab and selecting hide. So I'm going to go through and actually hide all of them. So now we have everything as shown by this tab down here of calculator; we have every tab hidden except for that. And if I wanted it to reappear or get a sheet to reappear, I would just right-click it, click unhide, and then it's going to allow me to select which option I can unhide, and and if I do want to make it to where a user can't go in and necessarily unhide sheets, well I can go in here and select protect workbook. Once again, I can enter a password if I wanted to. I'm going to just set this up, but now when I come down here to right-click it, there's no option to hide or unhide a sheet, so the entire workbook is now protected. So I'm not going to lie; that was definitely an advanced intro into Data validation and also protecting your workbooks, but I promise it's going to just come into great use for whenever we're building this project, which will we get to next. Now we do have some practice problems for you. Go through and just test out all these different features, and with that, we'll be jumping in the next lesson and actually building this data science salary dashboard. With that, I'll see you in that one. All right, let's now dive in and build our first project with Excel, which is this data science salary dashboard. This project is going to combine everything that we've used and learned up to this point, from formulas and functions to charts and then even to data validation. We're going to start first by looking at the dashboard itself. You can just go to the Project One dashboard folder and Open salary dashboard workbook. Now in this, right now you're only going to see one sheet, and as you try to click around, you're not going to be able to do anything. So as a refresher, if you want to actually dive in and see what's going on behind the scenes, you'll need to First, if you want to actually touch any of these points, actually go into the review Tab and click unprotect sheet, then you'll be able to investigate how I name certain cells and whatnot. Additionally, if you want to investigate any of the workbooks that I worked on, you'll need to go into unhide and select the appropriate workbook that you want to, well, unhide. So for this, we're going to be building it out section by section; specifically, we're going to start up at the top, building these data validation drop-down menus. Then from from there, we'll go into building the different graphs associated with it, and then finally we'll end up with these KPI cards. Now powering each one of these major topics, I've built individual sheets. So for things like jobs, I have all the jobs along with any key information to then build the visualizations in it. So here is the basically the table that I made in order to show the graphic right here. Similarly, for Country, I have all the different countries and then they're Associated Med and salaries, and I use that to not only make the drop-down but also make the graph. Same thing for type, and then finally for platform. Anyway, that's just a quick overview to make sure that you're under familiar with how we're going to be working through this, but let's actually dive into it. For this, I recommend picking up where we left off in the last lesson on collaboration. Did a lot of work for that, so we're going to use this workbook. First thing I'm going to do, once this is open, I'm going to go in and actually save it as this final dashboard, and I recommend that during this you're saving this pretty frequently, so we don't lose progress. First thing I'm going to do is start moving this around. I basically know where I want to get these different titles of these drop downs and then where I want to put the drop downs. We're not going to be using meeting salary for a little bit, so I'm just going to take that control x it and place it down at the bottom, then take the job title, put it in C3, and then move the data validation to right below that. We'll fix all the format add in when we get later on it. Okay, so we have the job title. Now the next thing we need to jump into is country, and we'll be putting that right under this portion right here. For this, I'm going to create a new sheet and call this country. With all these sheets, I want to have them pretty much similar to what the title is above it. So in this case here, where we had median salary, it's actually the titles um you have named it in the previous one salary, so let's go ahead and just name this title. Anyway, going back to that country tab, that's where similar to the title tab, if you see, we first grab the names of the job titles from there and then calculate the median salaries for each. We're going to be doing something similar in the country tab, with first putting in the country names and then from there putting in that median salary, but I want to keep a similar format as in this title case. Remember, we actually pulled this from the data validation tab, which we're pulling here, so I want to keep this consistent. Anytime we're creating anything for those drop downs, we're going to make it here in this data validation tab. So I'm going to create a column here called job country, and then in this I want to get the unique values from our data set, specifically that jobs table; it's still named that jobs table, and of that column job country. Go ahead and close the brackets and then close parentheses, and now we have all of these different countries. Not sure why, but this is bolded. I'm going to go ahead and remove that. Anyway, I want this in a sorted format. I'm not going to necessarily sort it like count like we did here with the job tiles; I'm just going to sort it in alphabetical order. So I'm going to use the sort function, and I'm just going to identify that we wanted to use G2#, and Bam, now we have all of this. Also name this appropriately of job country sorted. So now we have our list; we can go back into here and actually put in the country for the data validation portion. We do that by going to the data tab, selecting data validation, and the values we want to provide a list to this, and for the source we go back to that data validation tab, close this out, and we basically want to select all these values here, so I'll just do control shift down, pressing enter. We now have everything; all the criteria for this, I'm going to go and click okay, and I get this error message, and there's a problem with this formula. For some reason, I guess when I move back, it added this extra sheet in here. I'm not too sure this extra data; I can't even select in here. Anyway, just make sure it's only one sheet there; it's going to work fine. Country is now in here; I can s something like Argentina. Next value that we're going to be looking at is the job type, so part-time, full-time, whatnot. With this, although we're not going to use it yet, I'm going to create a new sheet and call it type and also move that to the end. But now we want to get the unique values of job schedule type. So I'm put in the column here of job schedule type, and then from there we want to get the, once again, unique values for this. We're using the jobs table, specifically that job schedule type column, and Bam. Now you will notice from this one, this one, it's a little bit; this needs some data clean up with it. There's a lot of values in here like it sometimes it has combined values like full-time part-time and internship and and whatnot. We really; I'm actually going to expand this column out; we really just want the single values from this, so something like fulltime, contractor, part-time, internship, and then also temp work. So the first thing I'm noticing about the thing ones we want to remove is that they contain the word and, so we'll first identify those that contain and. We do this using the search function, which is a text function to find text; specifically, we're looking for that keyword of and, with intext, we want to just look through the whole array, so we'll put in J2#, and I got a little error message; I need to make sure I use double quotes for the text itself, and running this now I have basically number values for where the and is located at, and it looks like yeah, it looks like we're good on everything with the exception of the zero, which we'll get in a little bit. Okay, so we need to convert this into basically Boolean values, because we're going to end end up using this to to pull out that we want using a filter function. So we're going to wrap this in the is number, and we're going to get false or true and whatnot. Anyway, all right, so now we have false or true. The last thing we need to do is use, well, not the last thing, second last thing, we're going to use the filter function, and in this we provided the array, so in this case it's going to be J2#, and then for what we want to include is this other array that we just did. So I'm going to go ahead and close this and see what we get returned back, and we're returning now only the values that have and in it. We actually wanted to do the opposite of that, right? We want the values that don't have an and. So in order to do that, we're going to fix this entire statement right here for the include portion; we're going to wrap it in a giant not to turn everything around, add an extra parenthesis on the end, bam. Now we have full-time, contractor, part-time, we got the zero in there, internship and temp work. We just need to remove this zero out of it, so we just need to modify, once again, this right here, this portion of this include; we're going to do some array multiplication, basically, once again looking through and making sure no values equal to zero. So I'm going to do a multiplication, do an opening closing parenthesis, and basically we're just checking whether J2# is not equal to zero. Let's go ahead and enter this, boom. Now we have it down to the values that we want. For this, I'm going to name this appropriately job schedule type sorted. Also, for some reason, this is in this column; we're going to move it over. Looks like we're building one spacing. Anyway, now we need to go back to our basic calculator tab, and we need to enter data validation in this portion to make sure can select the right type. So going select data validation, once again allow values of list, and then for the actual Source itself, we'll go to that data validation tab, select all these values in here, press enter and enter okay. So now we have the type in here, so all of our data validation portions are now built. Next thing up is moving into building the three different charts here. We're actually going to start with the country chart because it's the easiest, and a sneak peek of what data is actually needed for this, I can go to the country tab inside my final salary dashboard, and all we really need to do is for each country calculate the median salary and then throw it into a map graph. So back to our Excel worksheet, first thing we need to do is get those list of countries, and remember we already have that, so I'm put equal sign; it's inside of our data validation here with these sorted values; I want all these values here here, so I'm going to do H2#, press enter. We have all them all. So let's actually start developing the formula for building this out using only; we're just going to calculate first the median salary for that country, and then also remember in the past we've have to filter out any values that basically equal zero. So for that if condition, for the Logical test, we're going to do, we're going to have to do array multiplication, and for our first array we're going to be checking for the job country, right? So we do that jobs table, and specifically that job country column, and we want to make sure that it's equal to basically A2 in this case, the country right next to it. Additionally, we want to check that there's no zero values, and so we're going to be checking the salary year average column and making sure that it's not equal to zero. So now moving on to the value if true, we basically want to use the salary year average column value, false not applicable here. Go ahead and close this. Looks like we have a typo; it went ahead and added that extra parenthesis, and we have a median salary now. And go ahead and copy that all the way down. Now this is great, but remember in our, if I go here back to to the basic calculator tab, we also want to not only filter for a specific country, but also we're going to need to filter for a job title and also for a job type, so we need to include not necessarily the country because we're doing it for each country, but we need to include the job title and the type. Now in order to add that, this formula is going to get a lot longer, and it's now getting hard to read, so I want to actually, I want to one; I want to operate in this formula bar. If you press control shift U, it expands it out, and then from there you can actually change it to the desired length that you want. So what I'm going to do now is actually break this into new lines. I can press on a Mac; you're going to press Alt Enter. On the Mac, I'm pressing option return. Anyway, I've went ahead broken this into different lines; I've also inserted some spaces in there to basically put in some indentation so I can read it better. Don't have to necessarily do that, but now I feel like this is much readable for my eyes. Go ahead and execute this, and Bam, we have all the results, and if I do a drag and drop all the way down, all the other ones are updated as well. So the first thing we need to add to this is to check for the job title itself. So I'm put a multiplication there, go to the next line, pressing Alt Enter, and for this I want to check jobs, specifically I want to check that job title short column and whether it's equal to basically title. Remember we created title, so I'm going to go ahead and press enter, and it looks like we have a typo because I forgot to insert a parenthesis at the end. Press enter. Looks like I misspelled the actual table at itself; my bad. Press enter again. Now I'm getting this name error right here, and that's because of this title that we're using. If we go back to that basic calculator and select that cell C4 right here, it's named titlecore exe, and I can inspect the different names assigned to cells by going to formulas, Define names, and then the name manager. Now I started directly with this workbook before we actually created all these variables here, so what we'll do is this: I'm going to go ahead and actually just delete this titlecore ex; that was just an example; that's why it says ex. Then from there, I'm going to just rename it; I'm going to select the cell itself of C4, and I'm going to change it back to title. Okay, now it's Title. Here, back to the country tab, uh, we have this updated for the title; it's actually appearing now; no name error, and I'll go ahead and drag it all the way down. There's going to be a lot less values for this cuz we're further filtering this, so I'm seeing some num errors; that's as expected. All right, the last condition we need to now take into account is this type right here, and we haven't named this cell already, so I'm selecting K4, and I've come up here, and I'm going to
Select type. And now I've renamed that as type, so we can finish this formula off. We wanted to—I'm going to do a multiplication sign, start a new line by pressing Alt+Enter, then do opening and closing parentheses. For this, we want to check if the job schedule type column is equal to type. Okay, I'm going to go ahead and press Enter for this. Looks like we have a value; I expect a few more, even filtered from here. Okay, not a lot. Now, one note on this: this formula is perfectly fine for checking the job schedule type. I'm going to make it slightly better and actually slightly more correct if I go over to that data validation tab. I'm going to press uh, Control+Shift+U to actually close that formula bar. If you remember from our job schedule types, yeah, we narrowed it down to this list, but actually, there were—the true list is this. So what we actually need to do is check if a value is in here. So, in our case, we want to check whether the type is in here. So if we select part-time, we will also match on this job type here, where it says full-time part-time, or this one here where it says full-time, part-time, temp work. And we can do that using the search function. So we can find something like part-time within the text of right here, and it's going to give us back a number. And then if it's not there—if I were to actually drag it down to something like the third column—it's not there, it's going to get a value error. So I'm going to come back into this and expand out the formula bar, and I'm going to change this formula right here to basically get that condition. Remember, we want to use the search function; we want to find the text of the type, which is that variable that we have for the job type, and we'll be searching the job schedule type column. Now, remember, this is going to return back a number of the position if it's there, so we're going to need to wrap this all in an ISNUMBER function and then put closing parentheses. So I'm going to autofill this all the way down again, and it doesn't look like any values, at least in view, actually changed underneath this formula bar for right now, so I'm going to go ahead and hide it.
And then for this, when we go to plot it, we actually need to remove these #NUM! values from here. So in order to do this, I'm going to—I'll create this new one called job country filter, and we're going to be using the FILTER function. And for this, we need to include the array, so everything from here downwards—pressing Control+Shift+Down to select that—and then what do we want to actually include? Well, we want to check to include anything in that B column, so ISNUMBER—we're going to check those values are equal to a number—so I entered in that B column then as well. All right, let's go ahead and run this, and it looks like it has all of our values. I don't like the order; I'd rather it sorted. This is just me preference; I'd rather the numerical values be sorted. So I'm going to wrap this all in a SORT function, and this is the array we're applying to it. We want to sort it on the second index, and for this, we wanted to put it in—we'll say descending order. And well, Puerto Rico has some of the highest jobs; may have to move there. And okay, we're going to get into applying this now. I want to make sure that we have the maximum amount of values present. There's a lot of countries missing that I know we have available, so I'm going to just select the most basic job possible to make sure that we have all the jobs that we can appear. So, so we'll just select Data Analyst, United States, full-time. Okay, now we can go about selecting column D and E and then inserting in our map. Now, I don't want this here, so I'm actually going to grab this map and then come over here and put it in. I'm only going to do some minor cleanup right now. I'm going to remove the chart title and also legend, but we now have this chart map available for countries that shows the median salary. One quick note: you are going to have this sort of warning right here. If I click on it, and it says, hey, we plotted 74% of the location from the data with high confidence. Basically, some of the countries in there couldn't align properly. In my opinion, it picked out a lot of the major countries, so I'm really fine with that. I'm fine if I didn't identify all of them; 74% is good enough.
Back to the final dashboard. So we made this country map right here. Now we need to make these other two. One thing to call out with this, which I don't think I've called out before: if we notice whenever we select a job—so in this case, I'll select Data Scientist—it makes that bar—all are a darker color blue; the way your eyes go towards it, and then you can compare it to the other ones. So how did I do this? Well, if I go to my Jobs tab, my final Jobs tab, what I'm doing here is I have all the median salaries, which we calculated already in ours, but I added this over here. Basically, I have one column without—we have Data Scientist selected right now—so I have one column without the value appearing in, and then one value with it appearing in. And then what we'll do from there is just some basically manipulation of the graph to make it to where, in this case, Data Scientist appears. So going back to our worksheet of our fancy-dancy dashboard, we have so far—going to go to that Title sheet. Remember, we already did all this portion of the last section. First thing we do is—well, we need to do some cleanup. We need to get rid of this #NAME? error. Also, we are going to create those extra columns right here for basically what job title is selected, but we need—need to more importantly—if I expand out the formula bar—we need to update this median salary, similar to what we do with job type, to not only take into account the job title but also the country and the job schedule type. So I'm all for not repeating our work. I'm going to go back over to the Country tab, select the median salary, and I'm going to basically just copy all that portion that's in there. Anyway, I'm going to escape out of that, come back into the Job Job Title tab, select B2, and I'll go ahead and just press Alt+Enter, insert all that in, and then now I just want to clean this up. We do want this country, which we're going to have to fix, but we don't need these middle two right here that we already basically have, specifically with the job country though. So remember, this thing's calculating the median salary based on the job title selected in this column here and column A. So this A2 is going to work here. Previously, we were doing the same thing with country; we don't need to do country anymore. We need to actually put in a variable of country, which we haven't created yet. So I'm just going to enter country in; it's going to give me an error, this #NAME? error. I'm going to come back over to the Basic Calculator tab, select this, and then rename G4 to Country, press Enter, come back to the Title tab. We're no longer getting that #NAME? error; looks like it's executing just right. I'm going to go ahead and drag it all the way down, and we do have an error in my formula. I have this comma right here; this is supposed to actually be an array—this whole thing is supposed to be an array. So now let's try it again, press Enter. Okay, $990,000 for Data Analyst in the United States. I know that's true, and now we're filling it in for all the rest. Okay, so we have what we need. I'm going to close out the formula bar, and remember, we want to basically—in one column, if it has the word Data Analyst, we want to not include it, and then another one, we want to only include that one. So we're going to use an IF for this. So if this value—which we're going to go ahead and lock the column—is not equal to the title, then we're going to basically display those results, which I'm going to lock the column for this. Otherwise, I just wanted to display an #N/A and not a value. Okay, go ahead and enter this, and it is Data Analyst, so it's not going to appear there, but it will appear all the rest of these. And so I locked those columns, so I can just drag this over. And now with this other one, I want to do the opposite: basically, if it's equal to title, I want it to appear, and then I'll drag and drop it all the way down. So these are the values I want to plot. So I'm going to select D2 to D11, then holding Control, also select these values right here, go in and insert recommended charts, and the first one up is actually the one that I want, so we'll go ahead and insert that. So I'll take this chart and also move that right here into the Basic Calculator tab. With this one, once again, I don't want a chart title, and I don't want a legend. The other thing are the values—the horizontal values down here. I'm going to go ahead and double-click on that, scroll down here all the way to Number, and we're going to do that custom formatting that we've done previously. If it's not appearing, uh, feel free to type the code in, but we're going to use this to basically format it as—with the dollar sign in the front and then also the K for the thousands place. All right, the last thing is, you know, I don't like to use a lot of different colors in this, so making sure the graph is selected, go to Chart Design and then into Chart Colors. Right now it's set under Colorful, which I think is an awful default value. I'm going to come down here and select—not this Monochromatic Palette 4 5, sorry, the—but the Monochromatic Palette 12. And that's because now Data Analyst will be the darkest blue; the other ones will be light, so that way my eyes go to that one instead. So now what we just did with the job title, we need to repeat it for job type. A lot of copy and paste in, so we're going to move a lot faster with this one because we've done most of this before. For this, we're going to be entering in the Type sheet, and I'm going to go ahead and pull all those things in from the Data Validation tab. Now we need to get the median salaries for that. I'm just going to come back over to the Title sheet, come into here, and actually just copy this entire formula. Then expanding this out with Control+Shift+U, pasting this in here. Now we need to just change this up slightly. So for the job title, we need to actually use the job title, whereas conversely, for the job type, we no longer want to use type; we want to use what's available in A2. Pressing Enter, we get our value for full-time: $990,000 for Data Analyst; that's correct, and then drag it on down. I'm going to go ahead and close this formula bar. And for this, I'm going to use—similar to what we did in that Country sheet, in where we not only filter the data to make sure we include ISNUMBER but also we sorted it. And that's because sometimes these values—sometimes we may not have values—and we go back to this Type tab; sometimes there may not be a certain job schedule type. So I'm going to go ahead and paste this in. Now it is working. I know there will always be five values, so I'm going to actually change this to B6 here and also B6 here and press Enter. Now I also realized I made a mistake earlier whenever I went to the Title sheet; this is only doing the SORT function, and we may have a condition where in certain countries they don't have all these different job titles available, so we need to do its similar here as well. So I'm going to paste that formula into here and then adjust it because I know there's always 10 job titles, so it's going to go down to 11 in this case and 11 here. We go ahead and run that; there's going to be no change. The one issue though is, in this case, if I go back to that Basic Calculator, it doesn't do it in the order that I want. So going back to that Title sheet, I'm going to change that sorting value from a -1 to a 1, so that way it goes in basically ascending order. And I need to do the same thing here, here as well in the Type sheet, where it's also in ascending order, because we're going to be making the same graph. All right, similar to last time, I wanted to—if the value is selected, I want it to be highlighted, so we need to make those same columns again. So if this is not equal to the type, I want the value to appear and it be #N/A because right now Full-time is selected. Dragging it over and then adjusting it for equal instead and then dragging it down. I do want it to appear if it's Full-time. Now I'm going to select D2:D6 and then these values in F and G. Once again, we're going to go to Insert Recommended Charts. I don't like these clustered columns; I prefer a clustered bar chart. So I'm going to take this and then put it in here, make similar format and changes as well of removing the title and then also the legend, updating the x-axis by going into Numbers and changing the format to a custom format to using the K value instead, and then finally the actual color itself by going to that Monochromatic color palette 12. So bam! Now we have a lot of this made, so I can go through now and select, say, Data Scientist; it will update for selecting Data Scientist, and then you see all these other values update as well. I can also select the different type—part-time in this case—and then the values still remain the same; it just changes the bar that it's selected to.
All right, the last major thing before we get into formatting: we're going to make these three KPI cards. One is for the median salary, the next is for the top job platform, and then finally on the job count itself for how many counts of jobs for all of these. Now, one quick thing: Excel doesn't necessarily have KPI cards like if you use something like Power BI or Looker; they provide cards to this. We're going to do some sort of backdoor approach, if you will, to make this into a KPI card. Basically, I'm going to insert in a text box, and we're going to put a cell equal to it. You'll see what we're going to do with it, but the main point is these values—this value itself is not—as you can see, it's a rectangle; it's not in a cell per se, but it is calculated within the workbook. Anyway, what we're going to be doing—I don't need this down here—this median salary—what we did from the last lesson, I'm going to go ahead and delete this, but the first we want to calculate is that median salary, and we basically have it already, and I'm going to calculate it right here in this column of I2. And for this, we're just going to use a simple XLOOKUP. The value we want to look up is based on the job title selected, so Title, and the lookup array is this array right here, and then the final return array is right next to it. There's a missing value right now because Cloud Engineers is not available in the currently selected. So make sure you're selecting the full values, and we're going to go ahead and close it. But we have now the median salary. So I'm going to actually rename this I2 cell to Median Salary and then going back into our Basic Calculator tab. Remember, I'm not going to insert it into a cell in here, but instead, we go into Insert and then Illustrations, and I'm just going to insert a simple old text box. I'll drag it right there. Now the thing is, I don't want to type inside of here; what I'm actually going to do is I'm going to select the box itself, so you no longer have that blinking cursor in there, come up into the formula bar up here, type in =Median Salary, and bam! Now if you notice, it copied the formatting that we previously have right here as a cluster number looking at right there; it copied the same formatting that we're using here in I2. So what I'm going to do is just go in here and change this formatting to a currency with zero decimal places. And then once we have this value actually updated, go back to Basic Calculator; we can see boom! Looks a lot nicer. We'll adjust the formatting as far as the size and stuff in a little bit after we calculate all the other ones. The next one from our final dashboard is the top job platform. So we've only calculated things associated with the job title, the job country, and the job type, so we need to make a new sheet, and we'll rename it Platform, and technically the column name is Job Via. And for this, we need to get the unique values of the Job Via column. Now, for this one, we're trying to get the top job platform, so we're not necessarily doing that based on what is the top median salary on this; I just want where are the most jobs actually located. So we're going to be doing a count using Control+Shift+U to expand the—we've been using this median with this IF array in it; we've already built this out already, which this formula does, so you could—so we're going to use this. I'm going to go ahead and copy it by pressing Ctrl+C, coming over to Platform, and then pasting it in with Ctrl+V. Okay, and instead of MEDIAN, we're going to use COUNT, and the only other thing we need to update on this is—we stole it from the Job Country page—is we need to update the Job Country to be well, Country, and we need to check one more condition. So we need to add to this array—I'm going to press Alt+Enter to create a new line—and we want to check that Job Via is equal to, in this case, A2. And we go ahead and press Enter. Looks like 10 were available for Via, Script, ZipRecruiter, and then it calculates all the way down. Now remember, our data set also has hourly data in there as well, so technically, if you wanted to—which I'm going to—I'm going to remove—move this condition right here that we're checking that it's not equal to zero; basically, it's also going to include if there's a job that has an hourly salary included. So I'm going to go ahead and backspace out of that, press Enter, and then from there, drag and drop it down, and I can see we added a few more values because of this. I'm going to close this formula bar, Control+Shift+U. All right, so now I need to sort these values basically from high to low, selecting all the values using Control+Shift+Down, the sort index; we want to use the second index, and we want to put this one in descending order because we want the highest one up at the top. And for this, it looks like SnagAJob is the highest. Anyway, uh, this is what we want—this first one actually appearing in our KPI card, but if you notice all of these have "Via" in front of it, so what I'm going to use is a TEXT function of SUBSTITUTE, which replaces existing text with new text. And for our text in D2, the old text that I want to replace is "Via" with a space, and the new text is just a blank value. So SnagAJob is now up at the top; this is what I want to be known as. We're going to rename this variable to Platform. Then we do the same thing on our dashboard of inserting a text value, and for this, I'm going to select it and say that it's equal to Platform. All right, so SnagAJob. And for this one—this one is well, somewhat simple—but in our Data Validation tab, we were—in the very beginning in the last lesson—we were calculating the count, and we were calculating a generic count of all of them, so we need to once again modify this because we want the count based on our three conditions here. So what I'm going to do is just basically steal it from what we did previously. Go
Into that B2 cell in the platform sheet, go ahead and copy this all. And then, then in here, I'm going to expand this formula out. I'm going to go ahead and replace that in B2 with this. Now, a few modifications we can make to this: we're no longer checking the job via column; we're not trying to check that for the count that was specific to where we stole that from, so I'm going to delete that and also this uh multiplication point. And then, this is checking all of the things selected of country, title, and type. We're wanting to check the count of a certain title, so instead of having title, we'll put in A2. Pressing enter, we have a lower value because we've the current filters are lower, and then we'll fill it all the way down. Closing the formula bar out, we now want to get the count for whatever is selected. So I'm going to go to an empty column over here, right here, and we're going to be doing an X lookup again. The lookup value is what is the title that we're using; the lookup array is we'll use this one right here; and then, for as far as the return array, right next to it. Pressing enter, boom, get a value of 537.
Now, just to be safe, in case there aren't any results, like say it was zero or something or not applicable, it's going to be basically not applicable. I do want to include if not found, I'm going to enter in "no results," and I'm going to do the same thing underneath the title sheet for where we calculated the median salary, put for "no results." So I'm going to go ahead—we want to get that count in there, so we insert that illustration again for us. We're going to insert a text box, and that textbox is going to be equal to count, which I don't think we actually named yet, so I actually need to go back, escape out of this, go back to the data validation tab, rename this "count," and then from there, with the text box selected, I'm going to put that equal to count.
Now, for each one of these text boxes, I need to go through and, actually, as you can see, the we have a text box for the value, but I actually want to use a shape, basically background, to tell us what we're actually performing or calculation that this KPI is showing. So I'm going to come in here and to insert illustrations for shapes, we're going to keep it—actually, we'll say a rectangle this time—and then we'll go ahead and draw it. Now, for the shape format itself, I'm going to go to this one right here, basically a blue around with white on the front. And with these shapes, you can still put in text in here, so I can put in something like "median salary," and I can open up the Home tab and I can actually customize this further, so I can make this bold, I can put in the center—I actually want Center top—and I'm going to make this slightly bigger by 20 point. Also, I'm noticing this box is a green outline; I don't really like that; I'd rather a blue outline, so we have that now. Okay, so how do we get that number? If you notice the number is no long—it's hidden behind here. We can do a couple different ways, but I'm just going to right-click this object, and then under shape format, you can go to send backwards; specifically, I want to send all the way to the back.
Now, getting into the actual text box itself, if you notice there's a little bit of a a box around it; I don't really like that. I'm also going to exp-expand it all the way to the edges. I'm going to format this one as well to be centered, bold, and then we're going to make the font much bigger on this, and I'm going to—once I like—I talked about remove that shape outline. Right now it has a a light one; I'm going to say no outline. Okay, so now it looks like a KPI card. Copying this, I'm going to then make two more, and for each of these, I'm going to send them back to the back, name appropriately to Top Job Platform and Job Count. For this, I'm going to just copy this text box here that has the median salary in it, and I just want to copy the formatting to the other ones as well, so we can conveniently use this paintbrush, this format painter. And I'll select this one—it disappeared; I have to reselect it—and I'll also select this one. If you notice the names are cutting off, so it's really important that you extend it all the way over. Same thing with the job count as well.
Now we're getting into the format portion of actually just doing some final touches on here. I don't like grid lines, so under view tab, I'm going to select remove grid lines. For each of these charts, I don't really like those outlines; I want it just to sort of blend in to make it look like it's there. So for the shape outline, I'm going to change each of them to no outline. Up in our data validation point, I want to make the spacing right. I'm also going to make these titles slightly bigger for the dropdowns themselves. I want them to basically pop out, so I'm going to change this formatting. I'm going to go to the cell Styles, and I really like this one of input because it sort of calls your eyes to what you need to go to. I'm going to make this G column slightly bigger and then shift the type over some. The other thing I want to do is add a title up here at the top for what this dashboard actually does, so I'm going to select cells B1 through L1. I'm going to do merge and center, and I'm going to change this to "Data Science Salary Calculator," along with going to the cell style, we'll do Heading 1 for right now. I want that to still be slightly bigger. Okay, now we're going to to start moving stuff around, but I want to get in—it's like its final form that I'm going to give to colleagues and co-workers—and I'm going to give it with the Home tab closed and also with—if I view this—can remove headings, so it moved the column headers, the A and the B, and then the row numbers as well, so it looks like everything's updated correctly.
One minor thing: this job count, I want to make sure after I select it, full-time, I saw that the formatting of the thousands with the comma is not there. So going back into that data validation tab, I'm going to select this, go to Home, make it a comma, and remove all the decimal places. Okay, looking good. All right, now we need to get this set up to give to colleagues. I don't want them to have all these other tabs or all these other sheets, so I'm going to go through and actually just hide the ones that aren't applicable for them. Additionally, the sheet of basic calculator doesn't really make sense anymore, cuz that was for that first lesson. I'm going to actually name this to "Salary Calculator." Now, colleagues could still potentially go in, and they could mess up these formulas, and so we need to now protect our worksheet, and we only want them to be able to manipulate these three cells. So we're going to be going through protecting the sheet, but we need to actually recall that we have to pick what cells that we want to lock, right? We need to select all the cells, and I preemptively told you to hide the headings. You need to go back into View and show the headings again, cuz we need to be able to select this triangle in the upper left-hand order to select all the different cells, and then from there, holding control, unselect these three cells, and then from there, we're going to right-click in there, go to format cells under protection, and we want to make in that case that they are locked, or basically we are going to be able to lock them. Conversely, we need to escape out of this and now select the three cells that we want to unlock, right-click, go to format cells, and for these, we want to make sure that they are not checked for this, so basically unlocked whenever we go ahead and protect the sheet. So now whenever I go into Review, go to protect sheet, I want to be able to select unlock cells. Once again, if you want to enter a password, you can. I'm going to click okay. So now I can't click anywhere else except for where we have our data validation, so I can go through and select things like data scientist and Turkey. Now I'm just going to add that last final touch of removing the headings. Bam, we have our dashboard.
Now, I promise, last last thing before we go, I'm noticing—and you're probably noticing as well if you're going through and manipulating these values—in this case, let's go from data analyst from previously selected data scientists—this me talking in real time—I want to show it takes how long it takes to load, and it takes forever to load. Why is it doing this? This is not good for stakeholders; they're going to get annoyed if it takes this long. I'm going to go ahead and unhide some of our sheet repeats, specifically that platform one. Now, these formulas that we're using—um, the array formulas to calculate these values—it's F—so in this platforms one, we have like—oh my gosh—in this case, we have close to 200—oh no, it's like slowing down even going through this—we're executing this hundreds of times in here, whereas if I compare it to something like the title sheet, we're only running this, you know, n 10 times, which I feel isn't that big, but if we're running this formula hundreds of times, it's going to slow down this sheet. So I have a quick fix for this, and it involves—we're not going to—especially for this sheet here, platform sheets—we're not going to use this um array multiplication order to calculate this; instead, we're going to use a COUNTIFS. The first thing we're going to do is check that the Java is equal to the criteria one of A2, so basically job platform is what it is says it is. From there, we'll check the job title short column to make sure it makes up with title; we'll check the job country is equal to Country; and then finally, we're going to check that the job schedule type is equal to type, and then we're going to go ahead and execute this, and then we're going to autofill it all the way down. Notice that 1490—it's actually going to go down slightly to 1426—and that's because we've now changed this condition inside of this COUNTIFS, specifically—if I go back to that title sheet—you remember whenever we match for this, we did a really in-depth search, so if any job schedule type contain those keywords, we match to it. Now we're only matching it if it exactly matches, but since this job platform is just providing—it's not providing a numerical value; it's providing what is the Top Value—I don't think the Top Value is going to change that much, so I don't think we're being inaccurate about this if we change this formula anyway. Going back to the actual dashboard itself, now whenever I change this from data analyst to data scientist, it is much faster. So now I'm going to go ahead and hide those sheets, and we are done. So that was a heck of a lot of work. So in the next lesson, we're going to be getting into how you can actually go through and share this dashboard, specifically for those that have a Microsoft description; you can use something like Microsoft online because it has all these features that we have within here and host it there for others to use. Additionally, we're going to get into my recommended method of sharing any your projects, and that's via LinkedIn. Now, just a heads up, we will be getting into Git and GitHub after Project 2 at the very end of this course, and during that portion, we'll talk about how to share not only Project 2 but also this project here, but that's more complicated, and I really want to focus on Excel. So with that, we're going to be shifting in the next lesson to quickly share it and then moving into the advanced chapter. All right, with that, I'll see you in the next [Music].
One first up, congratulations on completing your first project in Excel and building this salary dashboard. It's been nothing short of your hard work, and you shouldn't let that hard work go unnoticed. So in this lesson, we're going to be going over different methods you could go about actually sharing this project to your social network and to others to help out in the job search or future employment. Now, if you were just learning these skills for fun, you had no intent getting a new job or increasing your pay in your current job, then you can feel free to skip this and go to the next chapter on pivot tables. So there's a few different ways you can go about sharing your work that you did. We're not going to go dive into deep any of these; we're going to look at these more at a high level before jumping into one of the options. First up is a portfolio website. Here I have lukbru.com, and if I wanted to, I could come inside of here and edit it and include my project here along with what I did for others to see. Another option, even if you don't have a big following on YouTube, is you could actually go in and record and describe what you did within your dashboard and host it somewhere like YouTube.
Now, for both those options, you may be like, "Luke, how do I actually actually share my Excel file that actually went through?" Well, that's where we run into a little bit of issues, as as yes, we created this Excel file right here, but how do you actually go about sharing it with others to see your work? Well, one option for this is actually hosting your file online via something like OneDrive, which if you're paying for a subscription of Microsoft service, you have access to OneDrive, and you can host your dashboard online. All I need to do is navigate to onedrive.live.com, go to this add new and files upload, from there select my file that I actually want to upload online, and then we can go to it, and our file is actually uploaded here, which we can actually go through and select something like data scientists, and it will actually calculate based on the changes we make to it. Now, one note: the country chart inside of Excel online doesn't work, but I have a fix for it, and mainly it's to just remove it. You go into the review tab under protection and go to manage protection, and then you turn off sheet protection. Then from there, you can delete it. Next, all you need to do is just take those charts and actually extend them over so way they take up that extra space, and then once you're complete with that, turn back on the sheet protection, and now you can go about actually sharing this. So here I'm coming into share, and you can add an email if you want, or if you just want to share it in general with a link, you can come down here and fine-tune the control of a link to provide. In this case, I'm selecting that I'm going to share with anyone; they can edit it. You could make it view, but then they can't change the dropdowns, so I recommend that you still leave it on edit. You could set an expiration and even password, and then from there, click apply, and now you have a link to your dashboard that works even if you don't have a Microsoft account. So here I am in incognito mode within my browser, so I'm not signed in at all, and I can actually go in and access this dashboard and go through and select something, and it updates in real time, and because I got that sheet protection on, they can't go through and change anything except for these dropdowns. Don't believe me? You can check out my project via the link below. But what happens if we want to not only maybe share our file but also write up what we did, the work we did with this and all the different skills that we used? Well, that's the case of using something like GitHub. GitHub provides a location to store Excel files, like shown here, along with giving you the ability to go through and perform a write-up detailing all the different work that you did. Now, if you wanted to see this, you could just navigate over to my project where you download all these files from on GitHub, navigate into that Project One board, and in here has our Excel file and also this readme, which then appears actually underneath here and details all the different work that we did for this. Now, getting this project onto GitHub, if you're not familiar with GitHub, is fairly complex. We're actually going to be saving this for after Project 2, and in that case, navigating back to the project itself, we'll not only be uploading Project 1, we'll also be uploading Project 2 as well. So after we finish the last chapter, Chapter 8 on Power Pivot, we'll be getting into all of this, and you'll be learning more about Git, GitHub, and how to manage a project. Now, from what I found working in data science, it's that the best way to share your work and your project and potentially collaborate with others is use something like LinkedIn, a social media platform for networking, in order to share your project. Specifically, here I am on my profile right here, and if we scroll on down, they have a section in your profile to basically show all your different projects that you've worked on and contributed to, and adding a project is super simple. I got to do is click this plus icon, include a description—in my case, I was trying to help out job seekers in inves salaries for their desired jobs—put in a few skills—up to five—of Microsoft Excel, data analysis, or Excel dashboards. Now, for media, they do have options to add a link or media. In the case of the media, it doesn't support Excel files, and then if you try to insert your OneDrive link, I ran into errors. So I find the best way to actually just share the link is to post it inside of the description. From there, specify when you start and stopped on this project, anybody that contributed to it, this, or anything that is associated with, and then from there, click save. The other option that I recommend is actually just going in and making a post here. I just write up a short little description of what you did with your project, and then if you want, include something like an image or even something like a GIF, which shows an overview of the project, and then probably the most important thing is actually sharing that link to your OneDrive online. You can also post in the comments and not include in the description; it's really up to you. Anyway, go through there and then post. So bam, that's how you share your project. As a reminder, we will be going into greater detail into how to share both this project and also the second project on GitHub using Git and also use things like markdown in order to write about your project, but that'll be included after we go through all of the different Excel content. Just wanted to have a quick way of you going through and actually sharing what you've done so far, cuz I know you're probably excited and proud of it. All right, in the next videos, we're going to be shifting gear into the advanced chapters, getting starting off first with pivot tables. With that, I'll see you in there.
All right, welcome to the advanced chapter, and because we're getting into the advanced section, you know it's time for a new flannel. And with this Advanced chapter, we're going to be focusing on a few core topics that I think is going to make your life a lot easier. Specifically, we're focused on things like pivot tables, Power Query, and also Power Pivot. All of these are great at automating my Excel workflows to make it a lot easier to do repetitive analytics that my boss may come to me back and back again for. Instead of with something like a formula where I have to go through and make and copy and paste that formula all over again and rerun that whole analysis, these Advanced chapters are going to make your life a lot easier. Anyway, in this chapter, we're going to be focused on pivot tables. This lesson specifically will be getting an intro into pivot tables, how to make them, how to manipulate them, how to even read them. In the next lesson, we'll be going into advanced pivot tables, looking at things like grouping and even aggregating such as getting percentages of grand totals and whatnot, and then the final lesson in this chapter is on pivot charts, which allows us to basically take what we have in our pivot
Tables and convert it into a usable chart; hence the name pivot chart. All right, so let's actually get into it and understand why these pivot tables are so important. So, in the basics chapter, we made this table right here, which uses hardcoded values for the different job titles along with the different months. And then, from there, uses formulas—specifically some product along with some array calculations—in order to calculate how many job counts per month. This is cool and all, but what happens if we wanted to add another job title? So, say we have some like business analyst or we have software developer; we'd have to actually manipulate and upgrade all these different formulas that we have here.
Well, here's that same table, but in a pivot table. And by its name, that's what they're great at—they're great at pivoting and thus aggregating data based on certain values and whatnot. So, what if we want to add more job titles? Well, I can just come in here, similar to how we manipulate a table, select this filter dropdown, and then go from there and select things like, "Oh, I want to include something like a business analyst," and then the data automatically updates for this—no readjusting formulas—makes it super simple. I can even take this table a step further, and if I wanted to, I can actually filter by the job country. In this case, I'm filtering by the United States, and we now have these values; makes it super simple. Anyway, we're getting ahead of ourselves; we actually need to get into creating our first pivot table.
All right, so for the advanced chapters, it's going to be a little bit different for what files you're going to use for this. The final results of this lesson will be in the lesson title of "Pivot Table Intro," but what I want you to do whenever you're going through or following me along in this lesson is actually revert back to the previous file of the last lesson; in this case, or the first lesson, so we don't have one. So, I have this one called "0 of just pivot tables"—that's the one you want to start with. So, in this case, "Pivot Tables" itself just has the data tab of the data we want to work with, and this sheet of the table that we've been familiar with in the Basics chapter, which by the end of this we're going to make a pivot table out of—and when I say "out of," I mean actually of the core data itself. Anyway, for the actual "Pivot Table Intro," this will have also those similar tabs, but then also the lesson itself will have all the different work that we've actually done to complete what we need to do. So, feel free to just have both of these up during a lesson, so that way you can consult back and forth in case you get lost.
All right, so let's get into our first pivot table. We're going to be using the data that we previously been using of all the salary data for those job titles. Anyway, if I go into the Insert tab up here in the top left-hand corner, I have PivotTables, but I also have Recommended PivotTables. If I don't have an analysis in mind, I could come into Recommended PivotTables; a pane's going to appear on the right-hand side, and notice here that it actually selected the data range. I know that's the data range, and it goes through and provides some recommended different pivot tables that you could put into here, whether you put it into a new sheet or an existing sheet. But I know what analysis I want to do—specifically, I want to do a count of the different job titles, so data engineer; I want to find the accounts of this senior data analyst, and so on. Right now, it's not providing any of that. I don't typically find that anytime with Recommended PivotTables that it provides me what I want, so I don't find myself using that often. Instead, I go directly into PivotTables right here, and then we have three options, but we're really going to focus for this lesson and this chapter is from table or range. I'm selected inside of A4 right now, but it automatically knows that this is the data range all the way down to the bottom. The other thing it says is choose where you want the pivot table to place; you can either do a new worksheet or you can do inside the existing worksheet, but you have to specify a location. We don't want that. I typically like it in a new worksheet to keep my analysis in one standard location. The last thing it asked is whether you want to analyze multiple tables—specifically, add this to the data model. We're going to be going into Data Models very heavily in the Power Pivot chapter, or chapter eight, or last chapter. This is a super powerful feature when you have multiple tables you need to combine it. We're not doing it in this lesson or in this chapter, so we're going to leave it unchecked.
So now I'm in this new sheet that I'm going to rename to "Job Count," and I'm also going to move it over here to the end. Anyway, this PivotTable—this "PivotTable 2" that is calling it—is there's nothing in it right now, and you notice there's a few things that popped up. First is the PivotTable Analyze tab, which is available with this, and also the Design tab. We'll be going into these in some upcoming examples that we're going to get into. We're, however, going to be focusing on, for this example—example—on the job count. I'm going to close this out on this PivotTable Fields pane right here. Now, the layout of this you may see is somewhat different; is we have the columns over here on the left. So, if you remember the Job Title Short column, Job Title column, Job Location, and then these fields on the right-hand side are things for like filters, rows, columns, or values. So, I can take the Job Title Short column, put into something like the rows, and get basically all the values in the rows. Now, your layout may be a little bit different. If you come up and select the tools icon right here, you may be under this Field section and area section stacked, which has the fields down here on the bottom. I personally don't really like this because look how short my column titles are, so I like having them like this instead. Anyway, I think we understand this columns area right here, but I don't think we understand these filters, rows, columns, and values. So, let's explore this by calculating the counts of these different job titles. Now, anytime I add something to the rows or any of these columns, I can either remove it by grabbing it and pulling it off—notice they have the x mark on it—or similarly, I can also just come in here and click the uncheck mark box. That's more applicable if, especially for having it in multiple different panes and want to move it completely; makes it simple. Besides rows, we also have columns, and so instead of the job titles being in rows, they're in the different columns. I don't really like this too much; I typically find myself using rows. So, we're trying to calculate what is the count of these Job Title Shorts, so I'm just going to take that Job Title Short again and put it into the values, and it automatically aggregates this by counts of that. But what happens if I don't want to do that count aggregation? Well, one way is to come back into that values right here, and I'm going to just click it—not right-click it, just normal click it—and then go into Value Field Settings, and this pop-up is going to come up. First up is the name of the column itself. I actually don't like this for of a name; I'm just going to rename this to "Job Count." Under here, under the Summarized Values By tab, you can select a lot of different aggregation methods. We're going to stay with Count. You can also change how you show value as—basically, if we wanted to do a percentage of some total or not. We're going to be jumping that in the advanced lesson, so stand by for that. The last thing to note with this is the number format, so I can come in here and actually select—in our case, we have thousand values, so I like to use a thousands separator along with zero decimal places—and then clicking OK to apply this all; it updates the formatting and the name. So, we've gone over rows, columns, and values. What happens if we want to then filter, let's say for only United States jobs? Well, I could drag something like the Job Country column into filters, and right now it's selecting all. You have—you see this pane come up right here—and from there here I can actually go through and select something like the United States, click OK, and now the values, as you can see, they reduced and are only United States values. Other types of filterings I can do: I can filter the row itself. So, if I wanted to, I could select the different job titles that I want to appear in this and click Apply. I could also do something where, let's say I wanted only job titles, so we're going to do a label filter, and jobs that contain the word "data," so I could just type in here "data," and whenever I filter it, I get all the different jobs that contain "data." Similarly, I could also filter by this Job Count here, and that's by the values filter. So, I'm going to remove this label filter to start with, and we can go back in here in the values filter, and we could do something like, "Hey, we want to get jobs that are only greater than," let's see here, Cloud Engineers 33; I don't want to see that anymore; I get to greater than 100, and it filters down. But we're not going to use any filters right now, so I'm going to first clear this filter for the table and then also remove this filter from filtering for the United States.
Let's get into taking this analysis a step further, and we're going to want to now analyze the average salary of these different job titles. While we're going through this, we're also going to be exploring the PivotTable Analyze tab. A quick tour of this tab: First up over here on the left is PivotTables. If I wanted to, I could go through and rename this; I'd probably name this typically something similar to what is my sheet name itself—in this case, I named it "Job Count." Additionally, inside of here, we have options which allows us to do a lot of detailed control of how we're building our pivot tables; it's a very advanced feature. I don't find myself going into it quite often unless I need to fine-tune the functionality of it. Active field, so that tells us basically what's the active field. Grouping is something we're going to go into in the next lesson; we actually go and perform groups of different job titles. Slicers and timelines; we're going to be going into the last lesson on pivot charts in order to basically use these slicers and timelines to filter data. Section is used to control our data, so I can click something like Refresh or Refresh All; it's going to refresh the data that we have. So, in this case, remember Business Analyst is around 1,100. If I go back to our data itself and I find this entry on Business Analyst and then—and let's say that that's not correct and I delete that out of there—whenever I come back to this table itself, it still says 1,100. What I have to do is—well, we've updated the data, so I have to—well, refresh it. Now that I refreshed it, it's down to 1,000. I actually don't want to remove that entry, so I'm going to just press Ctrl+Z and bring that right back and then also click Refresh to make sure it's up to date. If I want to change the data source or maybe the range, I could go into something like this of Change Data Source. Actions allow us to clear, select, and even move a PivotTable. For calculations, they have things like Calculated Fields and Items, but we're going to get into measures, and I feel they're way more powerful, so we're not going to cover this much. The last thing to cover with this is over here on the right-hand side is the Show. Sometimes, whenever you're navigating, you'll click into your PivotTable, and that PivotTable Fields pane won't pop up. You can also pan it on and off by clicking this field list, and if you didn't want something like row labels at the top, you could just remove the field headers as well.
So, getting into that actual analysis, we want to analyze the Salary Year Average; what is the average value? Now, I can't see all the different values selected in here, so I'm going to actually—I'm going to go ahead and close this pane up here to have a bigger view. Anyway, what it did was it did a sum of the Salary Year Average; we don't really want that; we want to go to Average, and I'll change this column name to "Average Yearly Salary." Now, if you've been following along since the Basic chapter, you probably know that I prefer me performing a median for this salary data over an average, but if you actually go through this, there's no median value for this. That doesn't mean you can't do median in PivotTables; you actually can; you can actually do even more advanced stuff, which we're going to get to in chapter eight and Power Pivot, but for now, we're just going to stick to only performing average for this. I'm going to click OK. So, the formatting on this is all jacked up, and we could go into that field settings and adjust that, or I can actually go in—as long as I have all the values selected here—I can select, "Hey, I want to convert this to a currency," and that I don't want any decimal places, and it's going to format all the values, and I feel this is a little bit easier because now—actually, if you go back and in exploring the Value Field Settings inside of Number Format—it actually applied this custom formatting for me, so it knows to apply that since I applied it to all the values that were visible. Now, since this is so easy, I could also do something like get the average of the Hourly Salary. Once again, it's doing the sum of that, and I don't want that; I want the average itself, and I can change that column name by just going in here and typing in "Average Hourly Salary." Inspecting the Value Field Setting, it also updates inside of here, and I'm going to go ahead and adjust the formatting as well—changes to a currency with two decimal places.
So, let's get into actually cleaning how this table looks up, and we can go and do this by going into the Design tab. Now, I'm going to start over here on the right in PivotTable Styles, and we can actually change what it may look like. In this case, I sort of like this one right here, the simplistic look. I can also change things like column headers, which I like the formatting on it, or whether I want banded rows or banded columns. In my case, I kind of like the banded rows; we'll go with that. Last portion is around the layout. If you notice down here, we have this grand total over here; this is a grand total based on—well, the column values; it's adding up all the values in the column, so this is on for the column. So, if I wanted to turn it off for rows and columns, I could come up here and actually do that. I kind of like this, so we're going to leave it on. I could also turn on—on for the rows and columns, but in this case, because we're doing different aggregation methods—so a count here and an average here—it's not necessarily going to do anything over here for the row grand total, whereas for something like the columns grand total, that knows that, "Hey, for a job count, I probably need the total count; for the average, I probably need an average," and that's what it does for both of these. There's some additional ones up here on adjusting the report layout, adjusting for blank rows, and then also subtitles. We'll be exploring that as we go along as we build out more complex pivot tables.
So, let's now get into that final analysis, and we're going to be creating basically this pivot table that we did previously with formulas and functions. So, what we'll need to do or think of—right—we're going to need the Job Title Short in the rows, and we're going to need the month—the Job Posted Months—in the columns, and then we'll need to aggregate this by count for the values. Now, I can navigate back to the Data tab and once again go to Insert PivotTable. If you notice here, it says from table or range, so that's the really good thing. If we actually convert this to a table, we'll now be able to—once we do this—press OK and rename this to something like "Jobs." Now, we can really be anywhere in this workbook. In this case, I created a new sheet; I go, "Hey, Insert from Table/Range," specifically; I want to do a table of "Jobs," and we want to do this existing worksheet in A1, and all the values from that Jobs table are now here. So, we know we need the Job Title Short along the rows, but then we need the Job Posted Month across the top, which right now we have a date. We could put the date into the columns, but we get this error message: "Hey, you cannot place a field that has more than," well, 16,000 different values for it, so we're not going to do that. Also, before we forget, I'm going to rename the sheet to "Monthly Count." Anyway, we need a monthly value here, so what we're going to have to do is—good thing about the table itself is—now that we've created this as a table, I know next to this Job Posted Date column, I want to insert in a column called "Job Posted Month," and for this we'll just use that TEXT function that we already know, using the value of Job Posted Date, and then for the format we know we want three lowercase "m" to get the month itself. It's going to fill all the way down. Okay, so now we have Job Posted Month. Going back to our PivotTable itself, remember we're not going to see Job Posted Month in here until we actually go back into PivotTable Analyze and click Refresh. Now, Job Posted Month is inside of here, and conveniently it's also in the correct order. Now, this thing is completely blank right now; we need to actually add what values we want. So, I'm going to drag Job Title Short into values, and it's going to do a count. Notice here we do have column value—Val—which go up and down, and then the row values itself, so we can see what the count of Business Analyst is—around 101. I'm not really a fan of these things that say row and column labels; I'm—so I'm going to toggle off field headers to make this look a little bit better, and I'm also going to change the name of this to "Monthly Job Count," so bam! This is looking good, and we compare it to our basically non-pivot table just to make sure that our values are correct. We can see we have 982 for Data Analyst; come over—over to Data Analyst; we have 982. All the last thing we want to do is actually filter this down and better sort our values—specifically, I'm curious about roles in the United States, so I'm going to drag that Job Country over here and select United States from here to apply to it. Additionally, I care about the most important jobs at the top and the least important at the bottom, mainly by this grand total right here, and so what I can do is I can sort it by the grand total. But if you notice, I remove that—that filter button right here—whenever I actually remove the field headers, so I can also go in—instead—right-click this grand—the value inside of grand total, and I can say Sort from—in our case—Largest to Smallest. So, I feel like that makes it a lot more convenient. Also sort. Additionally, I'm noticing the formatting isn't correct for this; I'm going to put in that comma separator and then remove the two decimal places. Similarly, not only did we sort by the grand total, let's say I only wanted maybe the top six of these right here; I could right-click any of these job titles right here and then go into Filter; in this case, I'm going to go Top 10—instead—I'm going to select Top 6, press OK. Now that we have this all sorted, I can once again go into that Design tab, change the grand totals; we're going to turn it on for columns only, and bam! Now we have basically the same pivot table that we had before with our values.
Using formulas, but instead, now with pivot tables, and this is a lot more customizable. All right, all right. It's your turn now to get your hands dirty with some practice problems and exploring how to make some different pivot tables.
In the next lesson, we're going to go deeper with pivot tables, looking at things like grouping, hierarchy, and how we can show different values. With that, I'll see you in the next one.
So let's get into some advanced pivot table features. And for this lesson, and actually for everything in the advanced chapter, we're going to be sticking with that salary data set of over 30,000 rows in order to actually analyze for this. So I'm not going to be calling it out really any further into other lessons or chapters. The first thing we're going to focus on is hierarchy, which allows us to look at things like we want to aggregate not only the job title itself but also by the country. So what job titles are within a country and then look at specific values there for, say, like the salary.
Next, we're going to move into grouping, focusing first on automatic grouping, basically using that job posted date column to automatically aggregate by year, month, and whatnot. And from there, we'll then shift into some manual grouping. We'll be able to create groups of different job titles and basically break out whether we want to look at maybe senior roles such as senior data analyst, senior data engineers, and compare them to just normal data nerd roles such as data analyst or data engineers. With this, we're also going to dive deep into understanding a deeper method to analyze maybe percentages of totals or percentages of grand totals when analyzing these types of groups.
For this, you can continue working with that workbook you were working with on the last lesson if you did everything you did there, or you can just open the pivot table intro for this lesson. Once again, as a reminder, the solution is going to be in pivot table Advanced. We don't want to open that just yet because it could mess up what we're doing here. If I could, it is so we have four different sheets that we created with this. I only really care about the data tab right now, so I'm actually going to select all these other ones by holding control and then right-clicking it to hide them.
So let's actually look what a hierarchy actually creates. I'm going to go in and insert a pivot table from table/range. Remember, we're using that table of jobs; you should have named the table that in order for this to work. And we're going to insert it in a new worksheet. I'm going to move this over, and I'm also going to create—uh—call this sheet "Hierarchy." So for this, we want to look at the salaries for job titles in a certain country. So we're going to start by dragging that job country over to Rows. And right now, there's no hierarchy, but if I drag job title short into the rows as well, when we close this tab up here, we can see that now we have two values in here, and how we have values underneath here. We've now created a hierarchy. So Albania is basically the parent or the top of this, and then we have data analyst, data scientist, senior data scientist. Notice there's only three values here, and that's because Albania's sort of a smaller country; they only have three types of jobs there, at least in the data set. Now we want to look at salary for this, so I'm going to drag the salary year average into the Values. It's going to do a sum. Once again, going into Value Field Settings, I'm going to change this to average, rename the title to Salary Year Average, and then changing the number format to currency with zero decimal places, pressing okay for all this. Bam! Now I'm also curious by this how many jobs we actually have with a salary value; this is just sort of an add-on. So I'm going to drag that salary year average over, going into the Value Field Settings. I'm going to do a count of this, and we'll call this Job Count. Click okay. So now we get more of a relative idea of how many jobs are so in Albania; we have, well, only five job postings.
So now I want to get into actually seeing what countries have the highest pay. Now, as a refresher, you can come in here and select the dropdown, and we could either select how we're going to filter the row labels or filter the value labels, but remember we want to sort them, and right now this is only the—or—option to sort A to Z or Z to A for those row labels. Instead, I can just click—make sure I'm clicking the Sal Year Average because that's where I care about—I can right-click it, and from there, go to Sort. In this case, sort largest to smallest. And what it did is it sorted the values; well, within each of these, it still kept this—kept the countries in alphabetical order. Instead, what I can do is select this cell for the countries because I want to sort the country's highest to lowest, and then I can sort largest to smallest as well. Now this is pretty neat because now we can see things like Belarus, Russia, Bahamas—I gotta go down there—have some of the highest salaries by country, and then what those are based on the different job titles there. Now, sometimes I find reading this somewhat difficult in this manner that it's laid out here. I'm going to show you how you can actually change this. So going back into the Design tab, remember we had this Reports Layout that we sort of breezed over last. Right now, it's in this Show in Compact form. We can actually change this to something like Show in Outline form, and it will basically shift this over and have this hierarchy basically in two separate columns. It also makes it nice that you can actually a little bit easier to sort with. Another method is Show in Tabular form. So now it basically crunches it up, and I actually like this one even better, and it's still breaking out the job country and job title short into two different columns, but now it's actually aggregated to less lines, so I can actually see more data on here. Now this is definitely a form that I'd like if I want to hand over to my boss, and even if I wanted to convert this even further to—what is this—Repeat All Item Labels. So now I could, if I wanted to, actually copy and paste this into its own table and analyze further. At least now I have, like, Bahamas with the software data engineer—not software data engineer, I mean software engineer or senior data engineer. Anyway, you may have noticed there are some blank values in here, and that's because it has an associated hourly salary but not yearly. What I need to do is actually apply a value filter because it's a value. So I come in here and click the dropdown, go to Value Filters, and then maybe put something like greater than—we'll put zero—and now those values will disappear.
The next analysis we're going to do is a count by the job month, but we're not going to use this job posted month column that we created in the last lesson. Instead, we're going to use automatic grouping for this. So we'll go ahead, insert in a pivot table; we'll insert into a new sheet, and we'll call this Group Automatic. I'll go ahead and move that to the very end. Okay. So what I'm going to do is I'm going to take the job posted date—remember, it's a bunch of dates—and I'm going to throw it into the Rows, and this is going to get into some aggregation. It's going to take a little bit to load; my computer's not even loaded yet, but it's about 15 seconds later, and it is now available. If you notice, now we have this hierarchy of this grouping, and I can now dive into—in this case, January—and then one Jan here, and then diving in further, we can dive into specific times of job postings. Going to go ahead and close this up. If we actually investigate over here inside of here, we can see that after I dragged that job posted date over, it basically created a month, days, and then the date itself, which is actually a date time, but anyway, three different values, its own hierarchy with this automatic grouping. And so now I can go in and do something like drag the job title short into here to get the job count. I'm going to change this to Job Count, also go in and actually adjust the formatting, but now whenever I actually go into each one of these hierarchies and look in, we can see how many job postings were having on a daily basis and how many were happening at a certain date time. So now let's say I wanted to dive deeper to understanding maybe why July had such a high number compared to all the other months. One, I could double-click it, or I can just right-click it and go to Show Details. This is going to show—well, the details—and if we actually go over to the job posted date column, it's going to have all the values for July inside of here. So this is a pretty unique way to get into diving deep and showing the details of what is the data being used to perform these aggregations and also double-check your work.
Now we're going to get into manual grouping. Specifically, we're going to create this where we actually go through and aggregate based on the job titles itself, assigning it into—well, a group—so put data analyst, scientists, and data engineers into Data Nerds, senior R&S into Senior Data Nerds, and then these guys into Other Data Nerds. So we're going to create a pivot table for this. Go in and select okay, using the jobs table, and we're going to be grouping the job title short. So I'll drag that into the Rows for the time being, and we'll just start by grouping just the data nerds. So I'm going to just select one of these and then hold down control and then also select data engineer and then also data scientist. Then I'm going to right-click it and select Group. The other way I could also do this is go into PivotTable Analyze and select Group Selection. The next one I want to group are senior roles. So I'm going to just select all the different senior roles; conveniently, they're all right next to each other. Then I'm going to right-click it and go to Group. So now they're their own group. The only thing left is getting the rest of these. I'm actually going to have to control—select these and then these as well—and then from there we'll group that. For Group 1, I'm just going to select it, come up to the formula bar, and type in Data Nerds. Name Group 2 to Senior Data Nerds, and then Group 3 to Other Data Nerds. Also going to zoom in a little bit to get a little bit closer. Now that we have all these grouped, let's actually dive into performing a basically deeper analysis on this to look at how or what percentages these make up of all the job titles and also of their respective groups. Specifically, we're going to be looking at—on going job—the job title short over here; we're looking at the count and how the counts of those jobs are going to be of the percentages. Anyway, I'm going to change this to Job Count, along with going through and updating the formatting to use a comma and no decimal places. So for this, I still wanted to use that basically count of the job title short. So with these counts, we're going to do the percentages. So I'm still going to use that job title short column; going to drag it into the Values. We did a count, but now let's actually go in inside of the Value Field Setting. Remember, we got to that Show Value As, and inside of here we can have different values: Percent of Grand Total, Percent of Column Total, Percent of Row Total. We're just going to go Percent of Grand Total to start. Press okay, and bam! This is now showing us the percent of the grand total. Now I'm not liking how this is ordered right now. I'm actually going to—I'm going to sort this, selecting one of the values inside of the Job Count, from—sort it from largest to smallest—and then also I want to do the actual grand total itself, sort largest to smallest. Anyway, we updated this to Percent of Grand Total; we need to update the title to specify Percent of Grand Total. And so we can see that Data Nerds, for their parent, are taking up about 76%; almost 34% of the jobs are that, and individually we can see that data analysts are nearly 30% of that, whereas we get down to the other data nerds; they're only taking up a very small percentage. Now what happens if we want to see—so what is data analyst of the actual parent, or what is the cloud engineer of the parent—Other Data Nerds? Well, I can drag that job title short into the Values again; it's aggregating by count, but I can go in, and this time I'm actually just going to right-click it, and we can have this Show Value As. I'm going to use that instead. We can do that Percent of Grand Total, but instead, I'm going to come down here to Percent of Parent Total, and in our case, it's asking us what is the parent. Now you didn't—we haven't gone over this—but it actually recreated that—that grouping as Job Title Short 2. So I'm going to click okay. We don't want to do the job title short; that's not the parent—Job Title Short 2—and that's—you can see it actually down here—Job Title Short 2; it created inside the Rows. But anyway, getting back to the parent, now it's showing the percent that it takes to make up the parent, and then obviously the parent is at 100%. So I'm going to rename this one Percent of Parent. Now we just looked at Percent of Grand Total and Percent of Parent, but the Show Value As has a lot of different other ones you can also do in here. If I wanted to, I can even do something like Rank Largest to Smallest. Once again, it's asking us, do we want to rank part of the parent or part of the job title short? I want to rank part of job title short, and it will show its individual rankings underneath each from highest to lowest. I'm going to go ahead and undo this; I don't want to necessarily keep that.
One more note before we go: For those that purchased the course, practice problems, and also note, I also go into Calculated Items and Fields and have its own little worksheet for you to follow along and try out calculated items and fields. I didn't necessarily include it in this lesson because I felt that it wasn't a very powerful feature; I instead use, like, using measures instead, which we're going to cover in the Power Pivot chapter. But if you're interested about it, I have content on it in our notes, and those calculated field and items is underneath that PivotTable Analyze tab in here, on Calculated Fields and Items. We're not going to be covering it outside of those notes that you can follow along and do your own self-study with it. All right, you have some practice problems now. Go through and get more familiar with these advanced features and pivot tables because in the next lesson we're going to be diving into actually making charts out of these pivot tables using pivot charts. With that, see you in the next one.
Moving now into pivot charts. So we did a lot of work already in analyzing things with pivot tables; we're going to take it now to the next level: pivot charts. Specifically, we're going to be looking at first, what is the average salary by a job title? Next, we'll be looking at which job has the highest percent of demand, and then finally, lastly, we'll be looking at how—basically how—are jobs trending over time? We're going to be building all these charts using pivot tables. Additionally, we're going to include the features of slicers and also timelines based on what chart we're using in order to be able to filter down and more easily make our graphs more interactive. As usual, in the advanced chapters, I want you to start with the Excel workbook from the last lesson, so Pivot Tables Advanced, and if you want to see the examples or the final answer, you could go to Pivot Charts. For this, we're not going to be using the hierarchy or that Show Detail tab, so I'm going to go ahead and hide those.
Let's create this first chart to analyze what is the top-paying job in data science. For this, I'm going to just create a new pivot table for this using that jobs table, and we're going to be aggregating by job title short in the Rows and then the salary year average in the Values. And for this, we want to summarize values; we don't want to do the sum; we're going to do the average. I did this by right-clicking it, but we do have to have these all formatted correctly in currency with no decimal places, and I'll update the title as well to Average Yearly Salary. So in order to insert in this pivot chart, we're going to go to the Insert tab, and we're going to come here to PivotChart. There's only one option right now because we're selected on a pivot table, and that's a pivot chart itself. So I'll go ahead and insert it. With this, there's no recommended charts, but I know I want a column chart, so we're going to go with that. If charts aren't that different from regular charts, I can come up here, select this plus sign, I can remove things like the legend; I don't really need that, and then I can change things like the title by just double-clicking this to something like What is the Top Paying Job in Data Science. Now you may notice these pivot charts are a little bit different as they have these field buttons on here that basically allow you to—with the chart itself—go in and filter it. This is really convenient if—let's say this chart was in a different page. Anyway, I want to have these salaries sorted from highest to lowest, so I can come into here, and you know, we can go sort A to Z or Z to A, and you can change it around. We want to actually sort from highest to lowest, so I can come in here under More Sort Options, and I can change this from the job title short column to that average yearly salary column, and we want it to be descending, and we'll click okay. So bam! Now we have our salary oriented from high to low with our values. If you don't like these field buttons right here, you can come in and right-click it and go Hide All Field Buttons if you want, but if you want to get them back, you have to come back underneath the PivotChart Analyze tab and select Field Buttons and unselect this Hide All.
The next thing to analyze is which job has the highest percentage of demand. We're going to use that percentage of grand total column before, and we're going to be adding a little twist with this one as we're going to be also building in some slicers so we can slice the data for what we want. So back inside the worksheet you should be working in, so we want the Percent of Grand Total only. So I'm going to move out Count, Percent of Parent, and also that Rank Count to only have what we want. Next, move into getting a pivot chart built for this, and once again, I'm going to be using that column chart. I'll go ahead and insert that. I'm going to rename this to Which Job Has the Highest Percentage. Once again, I don't really care about that Legend. Now I want my—basically target audience—whoever I give this to—to have control to be able to select which group they can filter for, whether that's Data Nerds, Senior Data Nerds, or Other Data Nerds. So in order to control that, I'm going to first zoom out; we're going to insert some slicers for this. So if we come into the PivotChart Analyze tab, we can have—with this chart selected—I'm going to go into Insert Slicer, and we're going to do it for—remember that that group from last time is actually Job Title Short 2—and also we're going to filter this one also by country. I'm going to click okay; they're going to pop up here on top of this. I don't really like this; I'm going to drag it over, and I'm going to fix the formatting real quick. So now with these slicers, I can make it a lot easier for somebody using this to come in and say, hey, I only want to look at Data Nerds, or I want to look at Other Data Nerds and see what their appropriate percentage is. When you click on a slicer, you will notice that this slicer tab comes up; there's some different formatting options.
The one thing that I define myself do changing is the appropriate label or the slicer caption. In this case, I would rename this one to something like “job group,” and then for the job country, I would just rename this to “Country.” You can see they update appropriately here. For it as a refresher, right, if you wanted to select multiple different op options, I would select this multi-select right here. And then with that enabled, I can then select “data nerds” and also “senior data nerds.”
The last visualization we’re going to be building with this is a line chart looking at how jobs are trending out of time using that previous pivot table we made on the job count. For this one, we’re going to be using a timeline filter to be able to select down to maybe a certain quarter or month. So back in the workbook that we’re working in, I’m in this group automatic sheet. We want to create a pivot chart, so I go to “Insert” and into “Pivot Chart,” and for this one we want a line, so I’m going to go ahead and insert that. I’m going to give this appropriate title of “How are jobs trending over time?” Additionally, I’m going to remove that Legend and I want to add a trend line to it.
Now you notice by this one, the actual field values for this you have multiple different ones here. Remember, it did that automatic grouping in the last lesson, so you have not only months to filter by days and also that job posted date, so a lot more values here. Now to add a timeline for this, I’m going to go up to “Pivot Chart Analyze” and I’m going to go into “Insert Timeline.” There’s only one value that’s going to be available for this “job posted date,” and right now if I expand this all the way out, we have all the different months that are available. I’m going to go ahead and close this up right below it. So if I wanted to filter by a specific month, I could be like, “Hey, I want from February to November.” In this case, October, I actually need to select February. I’m holding down my key for this and then dragging to November. Anyway, I can also change this with this filter, not only months but also quarters and even something like years. I prefer, I typically analyze things in quarters, so we’re going to do it that manner, and I’m also going to shift it up here to the right-hand side. Similar to slicers, if I have the timeline selected, I can come up here and actually change the name. In this case, I’m going to change it to “date.” I could also change things like formatting or even things like color.
Now, one thing to note with this, with what I have selected here, it’s only going to filter what I have the chart set up to or what I actually created the timeline while the pivot chart was selected. So let’s say I came into here and I wanted to look at, in our case, just “data nerds” and then also go into looking at the counts themselves. This isn’t necessarily going to update for that; those slicers aren’t connected to other charts, but you can change it to do that. So in this case, I could select something like the pivot table itself, going into “Pivot Table Analyze,” and then here under filters where you can create things like slicers and timelines, which we did in the pivot chart. Anyway, they have this thing called “Filter Connections,” and I’m going to expand this out so we can actually see it. And right now we’re saying that, well, for “Pivot Table 3,” as we can see up here—probably need to give these even better names—only the date is actually connected to this. If I wanted to connect the other ones, such as country or job group, I’d have to select them and press “okay.” Now, I don’t know if you noticed that, but it actually adjusted these values; actually decreased because I have less values selected here, whereas if I actually select more, all of these going on this is going to increase the values. Anyway, that’s sort of hard to see. Let’s actually show this by with uh “Sheet 1,” which actually should be something like “Top Paying Jobs,” and in this case I can go into “Pivot Chart Analyze” into “Filter Connections,” and this is going to show us based on “Pivot Table 7,” which is this one right here. I should have renamed these; there’s no different slicers or timelines attached to it, so I can actually select all of these and apply it to this one. And now when I go to our grouping right here, right, so we had all of them selected, if I want to just look at “data nerds” here, so I can see the percentages of data analyst, data engineer, and data scientist, I can see what their salaries are for it, and then also I can see their counts for those as well. So this is definitely a useful feature if you’re looking to link charts or specifically pivot tables that are not necessarily connected. All right, now it’s your turn to get more familiar with using pivot charts. We have some practice problems that you go through and actually understand more about how to use them. With that, in the next lesson we’re going to be jumping into, well, the next chapter on Advanced Data analysis and using some pretty unique and pretty complicated features in order to analyze data. So with that, I’ll see you in that one.
Welcome to this chapter on Advanced Data analysis. This entire chapter is really focused on using add-ins, which are basically programs that people have built to incorporate into Excel to do very unique and specific tasks. Because of that, going from lesson to lesson, we’re not going to necessarily be building on each other as we go through these lessons. Every lesson is going to be sort of its own unique sort of Learning Journey about a specific feature or features. To start with this lesson, we’re looking at just enabling the add-ins and looking at some basic ones, such as what-if analysis. And we’re going to get more into it in a second, but we’re going to be focused on looking at if we had three different job offers, which one should we actually take? In the next lesson, we’re going to be continuing on with what-if analysis, focusing on data tables, and this shows us how values are going to be changing based on one or multiple variables. And then finally, the third lesson is on an add-in called “Analysis ToolPak,” that provides us access to a lot of different statistical analysis that we can just easily select what type of analysis want to perform, and it does all the analysis for for us and provides it in a sheet. Pretty neat.
Anyway, getting into this lesson, we’re going to start by first enabling these add-ins so that way you have it, and then from there we’re going to move into our first somewhat simple example, forecasting what’s going to happen into the future. Specifically, we’re going to look in at our past job postings and try to predict what’s going to happen in the future. From there, we’re going to be moving into what-if analysis, and for this we’re going to have a scenario where we have three job offers and we’re trying to find what is the most optimal one. We’re going to use things like “Scenario Manager” to go through and automatically calculate what it should be for those three different job offers. And then let’s say we need to actually negotiate one of those job offers, and we want to match another; we can use “Solver” or “Goal Seek,” and both of these have both unique different features of them that we’re going to dive into to allow us to adjust what we could potentially negotiate for better job offers.
One quick reminder on which versions of Excel will support this chapter on Advanced Data analysis: all of them will, with the exception of Microsoft online; it doesn’t have the ability to add in these specific add-ins, but you’re on Mac or the windows version, you’re going to be completely fine. So for this, we’re going to be working inside of the “Analysis Add-ins” workbook. I know it said previously you need to work with the previous workbook from the previous lesson, but this chapter in general doesn’t build on anything; it has everything you need within the workbook, so you’re going to be fine with this. Anyway, we just need two sheets from this: “Forecast Original” and “What-If Analysis.” All the others are just the results that we’re going to be getting, and feel free to go through and select the sheets that we’re not using, so these four in this case, and hide them so that way we only have the two sheets of “Forecast Original” and “What-If Analysis.”
So before we enable add-ins, I think you need to know what are exactly Excel add-ins. Here I am in Perplexity AI, and I asked the question, and it goes into to specify what it is by saying that basically interacts with Excel objects and data, and it will add custom ribbon buttons or menu items and thus providing custom functions. Now this is a little technical, but there are three different type of add-ins: they have web, Excel, and COM add-ins. Today we’re going to be importing in Excel add-ins, which are actually created using something like VBA. Anyway, the most popular Excel add-ins are things like Solver, Power Pivot, Power Query—you don’t necessarily have to add in unless it’s not included—and then also things like “Analysis ToolPak,” which we’re going to get to in that third lesson. All right, enough on the history lesson; let’s actually get into enabling your add-ins. If you go to the “Data” tab right now, you’ll probably see that you have this “Forecast” section, so you do have what-if analysis available, but you don’t have anything else over here right now; it’s um, well, usually blank, but we’re going to add to it. So I’m going to go into “File,” and then from there it’s hidden, but under “More,” I’m going to go to “Options.” On the menu on the left-hand side, I’m going to go into “Add-ins,” and this menu right here tells you what your active application add-ins are. Right now I have no active applications, and then your inactive application add-ins, so I do have access to all these different ones right here. So I want to enable them; specifically, I want this “Analysis ToolPak,” and then well, the one we’re going to use in this lesson, “Solver.” So um, on “Manage,” I have “Excel Add-ins”; that’s the one that I want to actually use for this. I’m going to click “Go,” and now we need to enable which ones we’re going to use: “Analysis ToolPak” for the third lesson and “Solver” for this one. From there, I’m going to click “okay,” and now over here on the right-hand side we have “Analysis” popup, “Data Analysis,” which is the “Analysis ToolPak,” and then “Solver” is the Solver added.
So let’s actually get into forecasting, specifically looking at what we expect job postings it to be next year. And right here in the “Forecast Original” sheet, I have “date” and then also the “job count,” and this goes all the way for or this is all the data for 2023. Anyway, this example is going to show the custom features that we really can do with some of these add-ins and also built-in features. So I can select the date and job count column, and then for this we’re going to go into the forecast and specifically to “Forecast Sheet.” In this, it plots in blue what are our values that we currently have for basically 2023, and then from there it plots into the future using this orange. I can toggle this between this a line chart and also a column chart, but I’m not really finding the column chart that useful; it’s time series data, so I’m going to go back to that line chart. The other major thing I control is the forecast end date, so if I wanted to only do maybe two months, I could change this instead to end in March. Additionally, have hidden underneath this drop down of options the ability to go in and actually change other things like confidence interval and seasonality and things like that. Right now it’s automatic; set it up to basically detect automatically, and seasonality is, as you notice in this data, it goes up and down, up and down, up and down; it has a seasonality to it; basically every single week there’s more postings during the week and on the weekend there’s less, as expected. So this seasonality is carried out into the predicted data, as you can see here because it’s still in the orange, actually goes up and down. Anyway, going to close this; this is great. I’m going to click “Create.” In this new sheet, it automatically has this popup here that says, “This table contains a copy of your data with additional forecast of values at the end. You can manually edit the forecasting formulas in the sheet or return to the original data to create a different forecast worksheet.” Okay, great, got it. I’m going to zoom out a little bit, and what this table did is it still kept that date and job count column, but it also built out three other columns to actually look at—scroll all the way down—what the forecasted would be, a lower confidence band, and then an upper confidence band. And then looking at the actual chart that it provides, we can see this where this darker orange color is what the forecast band is; this is the upper band, and then this is the lower band. Anyway, that’s pretty cool that I could generate this all by just clicking a single button of “Forecast Sheet.”
All right, now we’re going to move into what-if analysis, and we click this what-if analysis; we have three different things here: we have “Scenario Manager,” “Goal Seek,” and “Data Table.” For this one, we’re going to start with “Scenario Manager,” but let’s first go over what the data is here in the sheet that we’re trying to basically trying to calculate. First, let’s focus on these columns B and C. This is a, if you will, dashboard or calculator that I built, so I can put into here a base salary, a bonus rate, and then an annual raise amount, and it will calculate it. So let’s say our base salary is 12,000; I can put that into here, assuming the same 10% and 1.5%; it’s going to automatically update for this over here on the right-hand side in E through H. Over here we have three different job offers that we received, and they consist of the base salary, the bonus rate, and the annual raise. Underneath here, this fourth or fifth row, if you will, this is constraints that we’re going to use later on; I would just ignore this right now. So what’s going on down here in the result cell? Well, what we’re doing is we’re calculating what the expected salary is for year zero all the way to year four, and then from there we’re actually getting a total. So in this case, this is summing up all these values right here. So why am I doing four years? Why am I doing a total left for these four years? Well, the Bureau of Labor Statistics basically estimates that most people have the average tenure at a company of four years. So the idea with this calculator that I’ve made is that we’re able to calculate based on a job offer, we’re what would we expect if we were to stay at the basically average amount or median amount of time that a normal person stays at a job, like just looking at what’s the first year, because sometimes things like bonuses and annual raise may actually push us into higher salaries even though the base salary is lower than another salary. So it basically helps calculate this out and even the playing field for these three jobs that we’re trying to calculate. Anyway, you can go through if you want to and and understand what formulas are going on behind the scenes here, but basically I’m just taking into account these three parameters right here, and then every year basically starting with that previous years and then adjusting it for the annual raise and then giving it its appropriate bonus. So as expected, because there’s an annual raise on each one of these, the salaries are going up. So with that, what is going on here? Do I need to actually go through and actually put in every single one of those jobs? So I’ll put in job one and get the 566,000, and then now do the second job and third job? No, I can use “Scenario Manager” for this. So going into what-if analysis, I select “Scenario Manager,” and we’re going to add three different scenarios. So I’m going to come up here and select “Add.” This scenario name we’re going to call it “Job 1.” Next we’re going to move into what we’re going to use for the changing cells, and I’ve labeled these basically or made these into an input format; we’re going to select these three right here, so C3 through C5. We’ll leave the comment as is; protection as prevent changes and go to “okay.” Now it’s going to ask us what values we want to use for each; in this case I use 100,000, 10%, and 1.5; it’s already filled in, pre-filled in. From there, I’m going to click “okay.” Now we need to add “Job 2.” For this, I’m going to leave changing cells the same; this one I’m going to change to 880,000, 15%, and then change this bottom one to 1.2%. Then finally we need to add that “Job 3.” One of the last steps we need to do is now go into “Summary” right here, and for this we need to figure out what we want to actually have it provide for us. In our case, we want the result cell of C9 through C14 to be provided. From there we click “okay,” and bam, we’re going to get this “Scenario Summary” sheet that goes through in details based on Job 1, Job 2, and Job 3 for the value that we input into it, and from there it’s going to tell us what year zero is, year 1, 2, 3, all the way down to the total salary. Now one thing to note is you see these names of Base, bonus, raise, year zero uh, and then total salary. If I go back to what-if analysis, I’ve actually gone through already for you and actually named this. So in this case, I’m selecting zero; it’s named “Year Zero,” and “Total Salary.” If I were to use things that were maybe not named, it would just provide the cell. So if we’re using the values here, it would just going to be provide F6, and in that case we would have saw F6 here also. Back in the “Scenario Summary,” you may not have ever saw this before, but Excel allows this sort of grouping, if you will, to basically manipulate the sheets and what values are hidden or potentially shown here. Anyway, pretty unique feature that you may or may not have seen before.
All right, moving on to “Goal Seek.” Let’s say we have the scenario now where we got the job offer for Job 1 in this case, but we want to try to match that of Job 3. Specifically, if I go back to that “Scenario Summary” sheet, we can see that Job 1 is at around 566,000, but Job 3 is at 640,000. We’ll say we have some insider information that Human Resources told us, “Hey, we can’t adjust the base or the bonus, but we can adjust the raise, what raise you get every year,” and so you could potentially ask for a higher raise. What raise would you need to basically put into here to get equal to that Job 3? So the first thing I’m going to do is go in and make sure that we have inside of our formula input in the Job 1 actual statistics of it, so 100,000, 10%, and 1.5% for the annual raise. Now I could go through there, so I type 1.7%, and then 1.8%, and just keep on going up until I actually find what it is, or instead we can just actually use this “Goal Seek.” And for this, we’re going to be setting a cell, specifically cell C14, to that 640,000 that we want to get to, and we need to provide what cell we’re going to actually change. In this case, we’re going to change cell C5, which is the annual raise. No, for this we can only change one option; we’re going to be able to change multiple the next scenario, but not in this one of “Goal Seek.” So from there, I’ll go ahead and click “okay,” and bam, automatically goes through. I don’t know if you saw that; it stepped through it, and it went up to 7.6%, and that’s what we’ll need in order to get to that 640,000, and it even provides an old nice dialogue box saying that hey, it did find a solution. Sometimes you may put a goal in that’s not achievable, and in this case it would it would tell you. So 7.76% is a pretty high raise. Let’s say we get further information from HR saying, “Hey, we can actually change not only the annual raise but also your bonus.” We still have the same scenario; you can’t change the base salary; needs to stay at 100,000 for that first year. So we have multiple parameters now that are changing; this is when we’re going to shift.
From using this Goal Seeker, now over to Solver. One thing before we start: we need to actually reset these values in here. I'm going to change this back to 1.5%; both of these step up in value, so you want to reset it before you go.
So, opening up Solver, I'm going to set the objective as before, that C14 of that total salary, and we want to get it to a salary of 640,000. And we want to do this by, like we said, we can change two things in this case: the bonus and the annual raise. We can also add constraints, which we'll do in a second after we just run through this one. But I want to actually just go through and solve it first.
The last thing we need to look at is select a solving method. We're going to just leave it here; I really like this GRG Nonlinear. We'll leave it that for the time being, and we'll go ahead and click Solve. Now, for this, it says Solver found a solution; all constraints and optionality conditions are satisfied. As we can see, it increased the bonus and then also the annual raise, and we got to that 640,000.
Inside of this popup box, we can have it output certain reports. So I'm going to just hold Control and select multiple different reports along with clicking this for Outline Reports. That's it; it's going to actually print to different sheets. And from there, click Okay. Anyway, the most important of these three different reports that it gave to us feels the Answer Report. Basically, it tells us, "Hey, what were the original values put in for the bonus and raise, and then what are the final values in order to get to that final value of 640,000?" They also have these two other reports: one on Sensitivity Analysis and the other one evaluating the limits, which we're going to get to, um, but these I don't find as important.
So now, with this, with Solver, we found that we can input more than one different input. Now we can also specify constraints. If I come back up to Solver and it says, "Hey, in this dialogue box, subject to the constraint," right now the annual raise is sort of low, still at 2.1%, but that bonus skyrocketed. It was previously at 10%, and it went all the way up to 23%. So we could actually put some constraints in by clicking Add, and we'll say, "Hey, the bonus, we're not going to let that exceed 15%." We'll click Add for that, and then for the next one, we don't want the annual raise to exceed, we'll say, 4%, and we'll click Okay. Remember, I did name these cells, so that's why it pops up automatically as B and raise; makes it super easy whenever you name cells. All right, let's go ahead and click Solve.
So look at this: Solver could not find a feasible solution with these constraints. Basically, maxed out the bonus and maxed out that annual raise, and we didn't get to that 640,000. So what I can do is I can return to the Solver Parameters dialogue, click Okay, and in this case, I'll change the bonus to, we'll say, 20% now, and then for the raise, we'll change this to 5%, click Okay, and then try to solve again. And we found a solution: we have 17% and 4.4%. And for this, I'm going to Output the Answers; I'll click Outline Reports to export it, click Okay, close this out, and then go to the report. We can see what our final values are along with how we got to our 640,000 final value. All right, you got some practice problem, problem two, now go through and try these different features out of Scenario Manager and Goal Seeker and also Solver. And I think once you play around with them more, you can find out which one is more applicable to which scenario. With that, I'll see you in the next section where we're going to be going into deeper into what-if analysis, specifically on data tables, one my favorite features of what-if analysis. With that, see you there.
Let's now get wrapped up on what-if analysis by focusing on data tables. We're going to be focusing on building one-input and also two-input data tables. For the first one, on one input, we're going to be continuing on with that exercise from last lesson, looking into that job offer one and seeing how we could change the annual raise in order to thus affect different salaries at our 4-year point and mainly the total salary at this point. And from there, we're going to shift into building two-input data tables where we're not only analyzing that annual raise increase but also a change in the bonus rate to see what the different salaries are for that final total amount of those four years.
So for this lesson and also for this chapter, we're going to be starting with the actual workbook of the name of the chapter, in this case, Data Tables, and we're going to be working this original sheet. But I want to jump into that one input to basically show you what we're going to be building. We're going to be inputting into here the annual raise percentage; we're going to put it in increments of .5%, and then along the top in the row, we're going to be inputting the values from over here, um, and these values right here across the top. And then the data table itself is going to fill this in with the expected result. So in this case, year three at 2% increase in raise, it's going to get around 116,000. We also do for coloring at the end; the data tables don't do that; that's done with conditional formatting.
So here we are back in the original sheet. First thing we need to do is get the annual raise put up here. Remember, we want to go in; we'll say .5% increment, so I'll do 0, 0.5%, and then for the rest of these, I'll just drag them down. I end up messing up the formatting, so I'm just clear the borders and then put a border back around the outside. Now, for the salaries, I want that to be for what year zero, then year 1. I'll drag this on over for these, and then we'll put a total. So for this, I want to enter in that year zero; we're going to be doing this for all the different values right there. I'm going to go ahead and put it in. If you notice, it has this line through it, and I actually click it, and then from here, whenever I look into it, it provides the error of stale value. You may or may not see this, but I'm going to show you how to fix this. If you are experiencing this, you can go into File and then into More under Options. And what happens is, under Formulas, my workbook calculations went from basically automatically calculating to manually. Where under Manually, if I look at this little icon right here, they can be manually calculated by pressing F9 or going to Formulas, Calculate Now. Anyway, there's nothing wrong with having automatic calculations; that's actually what I want all the time. Somehow my thing switched into this manual; if yours does switch it back to Automatic, click Okay, bam, we're good to go, and we'll continue on.
Now it's important for up here at the top that we have them equal to the formulas here because this is what's going to be ultimately getting changed and manipulated, so I wouldn't want to go through and actually manually fill this in with a 110,000; it needs to be connected to the formula that actually is getting calculated. So building our data table now, I'm going to select this entire range right here, E3 all the way down to K12, go to the Data tab, and select Data Table. Now this provides us two inputs: a Row Input Cell and a Column Input Cell. We're only doing a one-input data table, so we only need to fill in one of these. Specifically, we're looking for the input either into the row or the input into the column. In this case, we're going to be subbing in this, this column, this E column right here; we're going to be subbing it into the formula here, and it wants to know what is the input cell for, in this case, the column. So I'm going to go ahead and select it; it's C5. I'm going to go ahead and click Okay, and it's going to automatically fill it in. Now what's unique about this is I could also go in here if I wanted to and maybe change this to something like 10%, and it will update this entire data table with that new value. I'm actually going to change that back to 3%, but pretty unique. Anyway, if I wanted to, I can come in also, and I'll so go in and to conditionally format it. I'm only going to select year 0 through 4, and I'm going to do a white to green, and then for the total, I'm going to do its own because it's almost in its own bracket here, right? It's a sum of all those different values, so I'm also going to do the same thing of the white to green, and then you know me, I don't really, really like green, so I'm going to go ahead and select this, and I'm going to end up changing this by going into Manage Rules and Conditional Formatting, selecting on this one, adjusting the color to blue, and also selecting this one and changing this one to blue as well. Click Apply and then Okay, and bam! So with that example complete, let's move into a two-input data table, and let's look at the final example for this. For this, we're going to have, as we had before, the annual raise in the column, but this time we're going to have the bonus up on that top row. And for this, we're going to be calculating, as we click here, it's going to be calculating C14, which is the total salary; we're not going to be calculating that 0, 1 through 4 anymore, and it's going to go through and calculate it for all of these different scenarios, if you will.
All right, to do this, I'm going to go back to that original sheet. I'm going to actually duplicate this by saying Copy it, Create a Copy, and click Okay. Okay, so now we have Original 2, so I'm going to name Original 2 One Input and then rename or two to Two Input. Now for this one, I'm going to end up just clearing the contents from here. I'll go to Editing, Clear, and I'll just select Clear Contents. And now thinking about it, I want to also clear any of the formatting that's in here because we're going to be doing something different with it. I can go into Clear Rules; I can go Clear Rules from Entire Sheet. All right, so we're going to have the raise and the rows, and now we need the actual bonus in the columns. For this, we'll go from 0% to 5%, and I need to actually change this formatting to actually be a percentage and then drag this all the way through along with fixing this formatting. So now a two-input data table is a little bit different in that we need in the upper left-hand corner what we actually want to change, whereas the one-input we did across, in our case, we did across the rows. In this case, we just want to have in the upper left-hand corner. There it is; I sort of grayed it out; you can make it a little bit darker if, if you want to, but I would just want to make it known that, hey, we're not necessarily using it. So similarly, we're actually going to get into creating it. We're going to select the entire data table, go to the Data tab, What-If Analysis, Data Table. For the Row Input Cell, so this row up here, what are we wanting to substitute these values into? Well, we want to sub it into the bonus, and then similarly for the Column Input, same as last time, that's the annual raise, so we're going to want to sub that into C5. Going to go ahead and click Okay. So I'm going to dress this up a little bit; I'm going to bold the header right here. Also, I'm going to merge and center this all so we can put inside of here Bonus, and then finally, I'm going to conditionally format it like we did last time using that white to green and then changing that green to a blue to get it more of what I want. So bam! Now we have a two-input table, and we can see what it's going to be across all these things. Also, with this, if you remember from our last lesson, right, we were looking at finding what is the value we'd want to be to get around 640,000. Now we have a few different values we can actually look at for this, and we can tell from this, well, we're going to need to be above a bonus rate of 15% to even be considered to get up to 640,000. So sometimes I like this visually better than going in and doing something like Goal Seeker or even things like Solver because now I have multiple different variables I can look at and analyze and try to adjust on my own. All right, so you now have some practice problems to go through and get familiar with data tables. I found when I first started with data tables, got really confused on the row input and also the column input cells, but really understanding how those are being applied into the original formula helps you figure that out. All right, with that, I'll see in the next one where we're going to be going into the Analysis ToolPak and diving into a lot of different statistical analysis you can do with Excel. So with that, see you there.
All right, this is the last lesson in this chapter on Advanced Ad Analysis, and specifically, we're going to be focusing on that Analysis ToolPak add-in. Now this add-in is packed full of features, and I can make a whole tutorial just on this add-in alone, but we're only going to be focusing on four core things of it that it does that I use from time to time on our job posting salary data set of over 30,000 rows. First, we're going to look at how we can get descriptive statistics of something like a salary column, so we don't have to go through and use formulas to get all the different statistics for it. Second, we're going to investigate how to make histograms, but these are a little bit with a twist in that I feel like they're more customizable than the previous histograms we can make. Third, we'll get into ranking and assigning a percentile for our salary data so we can understand where it actually ranks for percentiles. And then finally, we're going to be moving into looking at a moving average. If you remember our job posting data set had all the seasonality in it; basically went up and down a lot depending on where it was posted during the week. Well, we can remove those fluctuations by a moving average.
For this, we're going to be working in the Analysis ToolPak workbook, and all the answers in there, so you can feel free to go ahead and actually select all the different sheets in here and go ahead and hide them so we only have the Data tab in there, and we'll be working with this. So as a refresher, this is the Data Analysis ToolPak; you should have gone through in that first lesson and actually enabled it by going into Options into the Add-ins itself, and it should now be under the active Add-ins. If you didn't do that, remember you all you have to do is just go into, go into here and select it. All right, so let's open this bad boy up, and if I click that Analysis, it's going to pop up here, and this dialogue box allows us to select, like I said, from a variety of different tests that we can actually perform. There's a lot of different statistical tests in here, such as Regression and Sampling, and then even things like Correlation, Covariance, and whatnot. So let's start with the one that I find myself using the most, and that's Descriptive Statistics. When I want to perform EDA, or exploratory analysis, this is the first thing I want to do. Now the thing about this is we need to provide a column that has numerical values in it, so we could do the Date column, but what we're going to do is we're going to provide the Salary Year Average column. Go ahead and press Enter. For this, we do have labels in the first row, so I need to click this here. For Output Options, we want to go to a new worksheet, so that's what we'll leave for this. And with this, we do want the Summary Statistics. You can go in and also specify things like Confidence Level and the kth Largest and kth Smallest, but we're going to leave those default for the time being and click Okay. Now it's popped up in this new sheet called Sheet1, and diving into it, I'm actually going to expand this out and then format all these numbers real quick, so that's much more readable. So now we have all the key statistics from it; we don't have to go through and calculate a formula for mean, median, mode, standard deviation, the minimum, maximum, sum, whatnot.
All right, next up is Histogram. And previously, remember we could just select something like the M column, go into Insert here, and actually insert a histogram. Now the one problem I have with this is the formatting of the rows or the X values down here; it basically provides this range; this is a lot of data right there, and there's, it's really hard to format this. So let's look at an alternate option for this using the Data Analysis ToolPak. Specifically, we're to come in here to Histogram. For the Input Range, once again, I'm going to go ahead and just select that column M, press Enter. It does have labels. For Bin Range, I'm going to leave M; I'm not going to specify a width of the histogram or the bin; I'm going to leave it just default. For the output, I'm going to leave it as the new worksheet. I don't want either of these, the Pareto or the Cumulative Percentage; instead, I just want the Chart Output of this. Press Okay. And here we have the histogram; it's honestly not too special; it's a little hard to read based on the size of these bins. As you can see, basically the difference between these is around, it looks like they're doing basically a thousand increments, so the increments are way too small; we need to adjust the bin. Anyway, the one good thing is along this X axis, it's only one value now, so a lot easier to read. So now let's go in and adjust that bin size. So if I go back to Data Analysis into Histogram and click Okay, for the Bin Range, it wants me to actually put in a range or a selection, so we need to actually pre-fill out what range or bins we want for this. So I'm going to copy this header up here because we're going to keep the bin in frequency. Start a new sheet, paste it in here, and I want to go in; we'll say 50,000 increments, so 0, 50,000, and I want it to go to basically 400,000. So now going into Data Analysis again, Histogram, opening it back up, still has the Input Range selected correctly. Now for the Bin Range, I'll select A2 to A10. Select the Output Range to; I basically want it to be inside of this notebook, so I'm going to select up here on D1; we'll just start there, and we want a Chart Output on this page. Okay, I'll click Okay, and I'm getting this error message that the Input Range must contain at least one data point. Right now, this M is not referring back to the correct sheet; it needs to look at. So actually, I'm going to select right here; you can see it selected that other sheet. I actually want to select the M column of the Data tab. Now we'll press Okay. So now I love this because wanted output this; I didn't apparently need to do this frequency thing; I got confused. Anyway, we can actually go in and format this to remove the legend and then update the X axis title for Salary, and then we'll update this one for Frequency. Anyway, I really like this because now look at this; control, we were able to minimize it, not to go past 40,000 and have all these outliers and everything else that has past 40,000 is put into this basically more value. You, anyway, this is my preferred method for making histograms, especially whenever I need to control that X axis.
Next up is Rank and Percentile. And with this one, we're going to be doing a rank and percentile of that Salary Year Average column once again. Now this one, depending on the size of your computer, may take up, it may even crash your computer, so if you're concerned that this is not going to be able to performed on your computer, don't run it; just look at my example and understand what get out of it. Anyway, I selected Rank and Percentile, and then for the Input Range, once again, I'll select that column M, and then we'll output it to a new worksheet. I can do...
Something like, even name it. In this case, calling it "Rank and percentile" of the sheet that it's going to go to. So clicking okay, it says, "Rank and percentile input range contains non-numeric data." Basically, I forgot to click "labels in the first row." Clicking again, it's thinking. How long is it going to take? All right, so Excel just froze on me. Maybe that wasn't a great idea. Let's try that again, using "Rank and percentile" this time. Instead of selecting the whole column, I think because it had some blank values, especially down to a million rows, it sort of crashed it. Instead, what I'm going to do is I'm going to just select A1 and then select down all the way to the bottom. I don't know why it changed it over to column F, but the main point of me doing this is that way we select column M. And also, I need to remove this A1 at the beginning. Okay, and also need to update this to be starting the second cell. And we're going to try this again. I gave it the name of "rank percentile." I didn't have the "labels in first row" selected because we're going from the second cell. How long is it going to take this time? All right, so that was a lot quicker this time, and we have our, now in this "Rank and percentile" sheet, our actual data. It did take about a minute to do so. Once again, if you have a computer that's not necessarily that fast, don't try this at home.
All right, so some key statistics about this: it provides a point, which is the row number, it's itself, and then from there, what is the value? That's the column, the rank, and then the percentile. What's cool about this, because of the provided point, we could do something like the index function. And you provided an array, and then the row number. In this case, that's the row number. So if I wanted to find out what the job title is, I could select column B, and then from there, for the row number, go back to "rank and percentile" and select this value right here, then close parenthesis, press enter. Looks like it's a clinical NLP data scientist. And I can actually autofill this all the way down. Anyway, let's make sure this is actually correct. Okay, yeah, just double-checking the row number at 25589 is clinical NLP data scientist, so we have it correct. Anyway, I could go through now, and I did this for the job title itself, but you could imagine you could pull out things like the job country, job title short, all sorts of other key information and get this in a list of what it's rank is along with its percentile.
Our last feature to look at is moving average, and this is what we're going to be calculating here. The blue line already is data we already have of what are the job postings over time, but that orange line is the moving average. We can use this analysis tool pack in order to calculate this, and as you can see, it removes a lot of these fluctu- these weekly fluctuations, if you will, from it and makes it a lot more, are basically readable to see where actual the peaks and the troughs are. Now, in order to do this, I can't necessarily just put in that job posted date into it; I have to actually get a count of the dates and also what are the counts of the job postings per date. So we need to create a pivot table. So we go in, insert pivot table from table; we're going to do it from this table, which is named "jobs," and we're going to insert it into a new sheet. Similar before, we're going to put that job posted date into the rows, and I'm actually going to take out—you can see it aggregated by month—I'm going to take out the month from there so it does by days. And now I'm going to throw into the values here; it's going to do a count, so I'm just going to change this to "job count," and we can actually visualize this by itself by going to insert pivot charts, inserting in a pivot chart. We want a line, and that's what we saw before with our blue line before that showed how it basically went across, uh, went through time. All right, so goes ahead, and I'm going to delete this chart because we're going to be making it. And once again, we're going to that data tab into Data Analysis, and we're going to be forming moving average. For the input range, I'm going to select B4 and then select all the way to the bottom. Now, this grand total went into it, so actually I'm going to back up one and change this to 368. We didn't select any labels in the front row, so I'm going to leave that on blank. In the interval, I'm going to just set it something like seven for the time being. For the output range, I want it to go right next to my chart, so I'm going to copy this above and paste it below and change these B's into C's, so it's C values right next to it. And we want a chart output along with standard errors. I'm going to go ahead and click okay. Now, this chart is not correct; um, we made a little bit of a mistake, but I did want to show you real quick this moving average. We can see that it starts 7 days later right here, and so that's what's happening in this C column here; that's the actual moving average, and then the actual error itself is right next to it. It's pretty consistent around 30 to 40. Anyway, we need to fix this; we need to take this entire value, if you will, and move it out of a pivot chart. So I'm going to select this all the way down to the bottom and copy it. Then inside of a new sheet, I'm going to come in and paste it. I'm going to just paste—looks like a pasting with the pivot table formatting—I'm going to paste the values only and change this to "job date."
So let's try this again, using Data Analysis, going to moving average. For the input range, we're going to select B2 and then all the way down to the bottom. Remember, this has a grand total, so I actually need to change that to minus one. For the interval, I'm going to adjust it a little bit; I'm going to actually change this now to a 21-day moving average. And then for the output range, this actually needs to be adjusted to match what the input range is, but for B or C, sorry. Anyway, go—go ahead, leave everything else checked, click okay. And bam, now we have blue and also orange, if you will, for the actual and the forecast. Now, one thing I'm noticing with this chart is, well, the markers are pretty heinous; they're making—they're clogging up this chart. So what I can do is select something like this orange line right here; I can right-click it, go to format data series, and then here underneath this fill and line, go into markers, and then for the marker options, just basically do none. We just want to have a line instead. Additionally, we can just go ahead and click that blue—the blue line—and for the markers there, we can do none as well. Okay, sensory overload gone. Now looks a lot more readable, with the exception of down here, for some reason it didn't pick up the dates on mine, and we can adjust that by right-clicking that and going to select data. Underneath the horizontal axis labels, I'm going to go ahead and edit this. I'm going like from A2 all the way down minus one; we don't need to do grand total. Click okay. That changed the names; let's see if that updated the chart. And bam, it did. Now I'm going to do some minor cleanup; I'm going to remove that legend from there, and that looks a lot better. So now we have a graph of our moving average of the job postings, and as we sort of suspected, in August we had a peak, along with January seemed sort of high, then went down a little bit, but then up again in August. So we see a lot more trends and then tapering out towards the end of the year.
All right, now it's your turn to go through and practice with those practice problems and exploring some of these features in the analysis tool pack add-in. With that, we're going to be wrapping up this chapter, and in the next one we're going to be jumping into Power Query, which I'm super excited about, in order—how to clean up our data and load it in in the format that we want easily. All right, with that, see you there.
Welcome to this chapter on Power Query. And no pun intended, but this is one of the most powerful tools within Excel. It allows us to perform ETL processes, or extract, transform, and load, which is just some fancy data engineering talk for connecting to a data source and loading it in after you clean it up. Anyway, in this chapter we have five lessons. Specifically, in this one we're going to have an intro to Power Query, what it's all about, how to actually connect to a data source. In the next one, we'll be moving into the Power Query Editor, and we'll be covering that for three lessons in order to go in how to actually clean up your data and get it prepared to a format that you want. In the last lesson, we'll be diving into the M language, which is powering Power Query. Don't worry, we're not going to do any in-depth coding or anything like that; just want you to have some familiarity with—you so we have more experience with using Power Query.
So what's this lesson about? Well, in order to understand that, we have to understand is what is Power Query. And here on Microsoft's learning platform, they have this fancy-dancy diagram that basically shows this—what Power Query does. It allows us to connect to different data sources; it could be something like a database, a text file, or even something on the cloud. From there, Power Query will then pipe it in to a bunch of different products they have, and we're going to be using it for Microsoft Excel, but it's also famously also in Power BI. Now, if you have a Windows version of Excel, Power Query is going to work just fine. On the Mac versions, it is available; however, it's very limited, so a lot of the stuff we're going to do within this lesson you're not going to be able to do. And also, Microsoft online is just completely not available. So as a reminder, Power Query is an ETL tool, or extract, transform, load, and we can connect to as a data source such as this. Here's a Wikipedia page on the list of S&P 500 companies, and it has all the different 500 companies that are part of the S&P 500. Anyway, let's say I want this table. I could go through and try—I mean, as you can see, I'm trying to select it right now, and it's like selecting the whole page; it's a whole mess if I'm trying to get this. But we can actually use Power Query to extract all this components out. All I have to do is go in and provide the web address of this, which I know it's located right here. I'll then select which of the tables I want out of the web page, which is this one right here, and then I just load it in, and here it is in our workbook. Now, don't worry, I sort of ran through that example real quick; we're going to go more in depth and detail in the last example in this lesson, but I just wanted to show the power of this and how we can actually get data even from online into our workbook so easily.
So why do we need to use Power Query? Well, we're going to find that out as we go along, but I'm going to give you the tidbits right now: One, it automates the ETL process, so I don't have to do that annoying task of going to a sheet and copying it over every time I get new data. I can just get Power Query to do it for me. Additionally, with that, sometimes I may have mistakes I'm copy-and-pasting sheets over; therefore, I have reproducibility. And then finally, with this, I'm now allowed to bring data in that potentially exceeds that 1 million row limit of Excel, which we'll show how we can deal with that in a bit.
So let's actually get into performing our first example of loading in a simple data set, specifically from another Excel sheet. Like I talked about the beginning of the advanced chapters, you're not going to be able to actually work inside of the workbooks that I have given. So in this case, "Power Query intro" has the final results, but I don't want you working in that. I'll tell you what workbooks you need to be working with as we go through this, which you're probably getting the security warning of external data connections have been disabled, and we'll get to troubleshooting that at the end. So instead, we're going to be starting with a new blank workbook. I'm going to go to navigate over here to the data tab; this is where Power Query is located, specifically under this "get and transform data." It doesn't really say Power Query, but that's where Power Query is hidden. Now, anytime I'm importing any data, I typically go to this "get data," and then from there I navigate down deeper, depending on it's file, database, from Fabric and Power Platforms, or from even other sources. They do have for all these for it here; they also have smaller icons right next to it that you can navigate over and basically highlight, okay, this is from web, and then this is from a table or range and whatnot. We're going to be going over multiple examples in this video, so don't worry if you're not following along with which data sources you can actually import; I think you'll have a good idea by the end of this.
So what are we going to import first? Well, if you navigate into our course folder, under resources, under data assets, and then data jobs monthly, we have Excel files for every single month. We're going to start by just importing one Excel file to start, and then in the next exercise we'll go into how to import all these at once. Anyway, we're going to start simple first with just this Excel file. So for this, I'm going to go to "get data," and it's a file, specifically it's from an Excel workbook inside the course folder. I'm going to then navigate to the data set, going to resources, data sets, monthly, and then select that January data set and click import. With Power Query, you're going to find that it has this Navigator window pop up, and from there it will show you what is actually importing in. In this case, "January data jobs," the Excel sheet, and then if it had one or multiple sheets, it will appear there underneath it. Whenever I select "sheet one," it then shows me to the right-hand side a snapshot or a preview of all the different data in there. It doesn't show all the columns, but a snapshot of it. At the bottom, there's a few options to load or load to, and then also transform. We're going to keep it simple for the time being, and we're just going to load, so I'll go ahead and click it. So we just imported in this data set from another worksheet; it's already in its own table, and because it also was "sheet one," it's naming the sheet "sheet one (2)" to signify as the second one. So congratulations, we just completed our first ETL process of actually extracting, transforming, and loading an Excel workbook into another workbook.
So we loaded this table in, but how do we actually go about using it? Well, in this portion we're going to be demonstrating how we can manipulate it with a pivot table and how to basically control all our different queries. If you notice, we had over on the right-hand pane this "Queries & Connections." Now, if it's not popping up, you can go up here to the data tab, and then you see "Queries & Connections"; you can navigate it on and off by clicking this button. Power Query sets up these queries, and in this case it named it "sheet one" after the "sheet one" in that workbook that we exported in—I'm sorry, that we imported in—and if we hover over it we can get some details about the columns, when it was last refreshed, it's load status, and even data source. Now, "Connections" over here on the right, right now we have zero connections; that's actually what's controlled by Power Pivot, which we're going to be going over in the next chapter on Power Pivot, but anyway, back to Power Queries itself. Right now we see with "sheet one" that 3,000 rows are loaded, and if necessary we go through and refresh the data set, as showing it loaded the data, and 3,000 rows are loaded again, pretty quick. So let's actually get into manipulating this. Well, it says that 3,000 rows are loaded, but I actually—I can go in and delete this tab, and it's going to give you this warning that's going to—per delete the sheet—do you want to continue? Yes, I do. And whenever I do that, since the data is no longer loaded, it now displays that it's connection only. So we can actually change where we load our data to, if you will, and I can get to this by right-clicking it and then going into here, and we'll be exploring all these other options as we go through, but I'm only want to focus right now on this "Load To," and they have a few different options in here. Let's actually explore them right now. It has "Only Create Connection," so right now it only has a connection. If we go back to that table and click okay, it once again loads it into that table. If we want to actually get into a pivot table, we'll select this "On PivotTable Report." We can also do a pivot chart, and it asks whether we want to put it in the existing worksheet or a new worksheet. And then finally, it has "Add this data to the Data Model." You've seen this one before, and once again we're going to be going over data models more in depth in chapter eight on Power Pivot, so we're not going to be enabling this checkbox just yet. Anyway, I went in the existing worksheet; I don't need that table there, so I'm going to click okay, and it says, "Hey, there's possible data loss because we're going to be basically getting rid of that table and replacing it with a pivot table. Do I want to continue?" Yeah. And now, like we did before in the pivot table chapter, we're now using a pivot table, and so we can put things like "job title short" and analyze it for the count of different jobs that it has within it. There's no change whatsoever in everything we learn in pivot tables; still same application that we're using it here for.
So now let's actually get into importing multiple different Excel files. We're going to specifically be importing all 12 of these, of January through December. This time, whenever I go into the data tab, under "get data," and we want to get it from a file, but I'm not going to select an Excel workbook; instead, what I'm going to do is select a folder because all those Excel files are in the same folder inside my course. I'll navigate into resources, data sets, and then I'm going to select the folder itself and select open. Now you may notice the Navigator window looks a little bit different, and that's because now it contains the metadata of these Excel files itself, such as the name, data accessed, modified, created, and whatnot. And with this one, before we had that load and load to along with transform data, we're just going to go into combining this data set. So I'm going to go ahead and click that, and specifically we're going to use "Combine & Load." Now we navigate to a window we're more familiar with, of combined files. And what this is doing is showing is how it's going to actually combine the files. In that we need to make sure, one, that they're all the same format, but if I actually click "sheet one," of which the sample file is looking at is the first file, this is what it looks like, and we know this already because we looked at the January file. Anyway, if you wanted to, you could also change this to a specific file. I'm fine with just using the first file, selecting a sheet. If I was having errors, I would do "Skip files with errors," but I'm not worried about that just yet. I'm going to go ahead and click okay. And bam, now we have that—once again that table loaded into here, and this has all the data. So I expect it to have around 30,000 results, similar to what we've been working with before, and it looks like it does. And if you notice, we have this new column right here on "Source Name," which tells which Excel file each of these comes through. And just doing a cursory check, it looks like all the different months are in there.
Now, onto this Queries and Connections pane. Up here, I'm going to actually make this smaller so I can actually see it all. Previously, we only had our Sheet One query, but now we have also this Data Jobs Monthly query. And with that up here at the top, because we're connecting multiple different files, we have these helper queries that were created during the process. So you can navigate over these and basically see that, hey, it used the September file as a sample, and this is the steps it took, or this is what the sample file actually looks like. Anyway, I'm not too concerned with those helper queries right there, or with anything underneath this Transform from Files. I mainly care about what's under those other queries. So we have Sheet One and Data Jobs Monthly.
Speaking of which, Sheet One is a really bad name for this, so I'm going to rename this to Data Jobs January. I also rename the sheet, so just to prove with the Data Jobs Monthly that we actually imported it all in, we're going to go in and Load To, and we're going to do this time a Pivot Chart. Going to go ahead and click OK. We're doing the existing sheet with the table. I don't care if I get rid of that table, so I'll click OK. And similar for, I'm going to put that Job Title Short this, we're going to put in the Axis or the rows, and then we're going to want a count of that as well. And then I'll just organize this in descending order based on the count of Job Title Short. So bam, we now connected with Power Query to multiple Excel files and imported in at once. I hope you realize that now this unlocks a lot of potentials because, say you get January of next year's data, you could just put it into this folder here, and then just all you need to do is go back into the Data tab, click Refresh, it's going to go through and refresh all that data set and pull those new numbers in.
All right, in this example, you're not going to follow along; I'm just want to show the power of Power Query. Okay, the pun's getting old by now. Anyway, I have this CSV file, or comma-separated values, basically it has commas separating everything. I looked at this in VS Code; don't worry about any of this stuff. Like I said, you're not doing it. The main point is to show this data set itself right here. It's starting at the top row of one, and if I scroll all the way down, we get to the last entry, and that's 2.7 million jobs that I have here in this data set. We can actually import this into Excel. Now, if you recall, if you scroll all the way down to the bottom of Excel, it only includes about 1 million rows. So how the heck are we going to do this with Power Query? So this is a CSV file. I'm going to go to Data, Get Data from File, specifically it's a Text/CSV, and I'm going to import in this Data Jobs Large file that I have. Reminder again, you don't have access to this file; it's just too big to even get onto GitHub, so that's why this is a demo only. This is the data set itself, so I'm going to go in and actually go and look, load it. Now this has taken a little bit of time, as you can see, it's loading around 100,000 rows as it goes through. Also, it had—well, it has three errors now in here. This usually appears whenever it has a row of data that doesn't necessarily make sense for what it's supposed to import; it'll alert you there's an error. So the 2.7 million rows are loaded, but I get this error message: The query returned more data than will fit on a worksheet. Remember, it automatically by default tries to load it into a table into Excel, and it's telling me that, hey, it's not going to fit, so I'll click OK. Now it's still going to try to load that table, but it's going to cut it off at that 1.5 million. But this doesn't mean we can't analyze it. If I scroll over this query, it reminds me that the results of this query is too large to be loaded to the specified location; worksheets have a limit of 1 million rows. Sure, instead what I'm going to do is go and Load To, and I'm going to load to a Pivot Table. Click OK, and it's going to warn me again about the table loss. Yeah, I know. So once it loaded, like a hot minute to do that, I can actually go through and now analyze these 2.7 million rows. So if I do something like put the Job Posted Date into the rows, and we also want—and we want to get a count of this, so I'm going to put the Job Poster Date also into the Values, so we get this counts. Anyway, reformatting it with commas to actually be able to read this, now we can see that we did actually get in 2.7 million different data points for this. And as a side note, this is all the data that I've collected since I started in 2022 doing this, so there's a lot of different jobs. So Excel is not necessarily limited to just analyzing 1 million rows of data.
All right, so let's finally get into that last example of importing in this list of S&P 500 companies. Feel free—you don't necessarily have to do this—table from Wikipedia, but I'll drop a link below on where this table is located, and you can use that if you want. So I copied the web page of that table, then I'm going to come in here and select this From Web. You can do Basic or Advanced; with Wikipedia, it's perfectly fine to do the Basic version. Putting in that URL, clicking OK, we get into that Navigator window, and there's actually multiple tables inside of here. One is the list of 500 companies, and the second one is a list of companies that have been added and also removed from there. They also just have random tables in there as well, just because in the internet you're going to have random tables like this: One of Main Menu Contents Tools Appearances, not applicable. Anyway, we want to do Table One. I'm going to go ahead and click Load, and now that we have it in here, anytime we do this, probably need to rename it appropriately from something like Table One to S&P 500 in this case. And bam, scrolling down, we can see that we have—um, should be 500—oh, a little bit more than 500, apparently the list has been updated to clear a little bit more. I don't know why that is, but got all the data nonetheless.
Now, quick note on: if you want to actually navigate into any of the files and see what I've done, whenever you go to open it—so in this case, I want to open Power Query Intro—I'm going to open it up, you're going to get this: External data connections have been disabled. Do you want to enable content? In this case, yes, you want to enable all that. Now the problem you now may also have is that it may give you a warning that your data source settings aren't correct. And what do I mean by that? If I go into Data, and then under Get Data, we're going to see this thing here for Data Source Settings, and it's managing settings for your data sources. Anyway, you're going to see these locations here; these are file locations of the data sets, and they reference the files that are on my computer. That's not going to be the same for your computer; it's probably going to be in a different location with a different name. So here I know that this is the Data Jobs Monthly folder. If I wanted to actually go in and update it with the actual location for where it is, I would go down here, select Change Source, and then from there select Browse to navigate to it. You're going to, once again, navigate to your course of excel.analytics into Resources, Data Sets, and then there's that Data Jobs Monthly. Click Open and OK, and then it's going to update. You're going to have to do that for this file and all the files within a Power Query and also Power Pivot because your file locations are not the same as my file locations. Then after you do that, all it should go through and refresh, but if it doesn't, you can manually refresh it underneath the Data tab by clicking Refresh All.
The last item to call out is the Options menu. We're going to be going into the Query Editor in the next video, so we're going to save that for that one. Anyway, Query Options has a lot of advanced details in controlling Power Query. In this case, of showing the Query Peek when hovering on a query in the query's task pane; that's sort of annoying to me; it pops up every now and then. I'm going to go ahead and unclick it, but they also have different behaviors you can control for data load for the Power Query Editor, the security, privacy, and even Diagnostics. So feel free to go through this and navigate and see what is available to actually customize with this. I'm going to go ahead and click my changes, OK, and now whenever I go to the Queries and Connections and actually hover over something like Data Jobs Jan, it doesn't just pop up on the screen and sort of catch me off guard, so I sort of like that.
All right, we now got some practice problems for you to go through and get more familiar with performing Bally ETL with Power Query and loading in some different data sources with it. With that, we'll see you in the next one. We're going to get into the Power Query Editor. Anyway, nothing to be intimidated by, as a lot of the core principles we've learned already in Excel are going to be applied to this new window, so you're going to pick it right up on it. All right, with that, I'll see you in the next [Music] one.
In this lesson, we're going to be continuing on with Power Query, focusing on specifically getting you introduced to this Power Query Editor. And in order to facilitate this, we're going to be going through—or walking through—actually importing and cleaning up our data science job posting data set that has over 30,000 rows of data. We're going to be automating a lot of the steps using Power Query that previously we had to use functions and formulas for, so it's going to be saving us a lot of times in order to actually automate this data. In justest for this, we're going to be starting out with a blank workbook. So I know we do have this Power Query Editor, but I don't want you actually editing from that; that's more for a reference. Now, if you do open this file in order to reference it as we go along, this remember you're going to have issues or an error saying, hey, data source isn't there. Remember, you need to go in and actually select where this data set is. So under the Data tab, Get Data, and then under Data Source Settings, you're going to need to update this link or this address right here of where you're actually accessing the Data Job Salary All Excel file. This is my location, not yours; got to update it. Anyway, like I said, we're not going to be using this, so I'm going to open up a new notebook. And like before, we're going to be importing in that data set, so we'll go to Get Data from File, from Excel Workbook. You'll navigate to the course itself, under Resources, under Data Sets, and then we're going to be using this Data Jobs Salary All Microsoft Excel file. Go ahead and import this in. We're going to select that Sheet One, and this time, instead of doing Load or even the Load To, we're actually going to go into Transform Data, and this is now going to pop open the Power Query Editor, and this is where all the magic happens behind the scenes in order to get our data cleaned up. So let's go over a quick overview of the window itself. It's very similar, laid out to Excel. Up at the top, we have a ribbon with four different tabs: Home, Transform, Add Columns, and View. We'll be walking through each one of these as we go through this lesson. Underneath here, on the left-hand side, we have which query we're selected to. Once we're building multiple queries, they'll start popping up underneath each other. We can close this if we want and make more room. Is right here in the middle is what the current step, or what the current status is, is of our data set. Now yours may look a little bit different right now; specifically, I have this Column Distribution enabled underneath the View tab, which I'm going to go to more in a second, but anyway, it basically outlines all the different columns or where we're at with the data set itself before we finally loaded in. Now, right above this area is a Formula Bar, just like similar again to the Excel UI, and this has all the steps, or all the code, the M language done in this current step, if you will, of actually cleaning up this data set. And you're like, step, like what step? Well, over here on the right-hand side, we have our Query Settings, and in it we have the name of our query, and then we have the Applied Steps. This lists all the different transformations that we've walked through. So just a brief walkthrough: the first step is Source, and if I look at the Formula Bar, basically what it's doing is it's connecting to that Excel file with the file path that it has. In the next step of Navigation, it's basically selecting, hey, out of that Excel file, actually select Sheet One. From there to actually load in, then from there we can see that the headers are actually in the first row and not up at the top. So the next, or third step, is the Promote the headers up to the top, and then finally the last step is Change Type; it actually goes through and assigns for each of these what data type it is. So in this case, Job Title Short, it assigns to type Text, whereas something like Job Posted Date, it assigns to type Number, which needs to be a Date, which we're going to fix that in a little bit. Down at the bottom, there's a few statistics on this; specifically talks about 16 columns and over 999 rows, and it tells you when the last preview is downloaded. Anyway, if I just wanted to stop here with this data transformation, if you will, I would just come up into Home, go into Close and Load. We're just going to do Close and Load To, and in this case, like I'm just going to put in a Pivot Table, specifically analyzing for Job Title Short, specifically how many different counts or that we have of this. We can see, totaling it all up, have around 32,000. Anyway, that's a quick overview; let's actually get into exploring each one of those tabs in the Power Query Editor. So we're going to go back to Data, Get Data, and from there you can just select this of Launch Power Query Editor. Similarly, you can also use a shortcut of just Alt+F12. I'm on a Mac, so I have to press Option, but actually launching this up, boom, it has it with just a shortcut. Anytime you launch it, it may be grayed out here, so we need to make sure that we go in and actually select a query that we want to analyze and transform. For this overview, we're going to start with the View tab because mainly I want to get into actually how we can use the Power Query Editor for EDA and thus save us a lot of time of actually having to analyze it in Excel in the spreadsheets itself. Instead, we can do it right here. So going through this, first thing is you can toggle on and off the Formula Bar. I always leave the form on, so I don't know why that's an option. Next is the Data Preview. I can change the font type. I can also change the Column Quality. So this is telling us if there would be a potential error in here, or if, in this case of Job Location, if there's empty values. You typically have error values whenever the data type isn't being—being understood correctly. So in this case, Job Title Short is Text; everything in there is a Text column. If I were to change this to Number, press Enter to run, I'm going to get errors all the way through here because, well, that was Text and can't convert Text to Numbers. Also, not sure why, but it should say 100% error, but it's not. Anyway, they also have this green bar up at the top, and you can use this; that's what I actually prefer, so I'm going to unclick on the View and changes from the Column Quality because you can actually look up here and see, and then also toggle it. So in this case for Salary or Average, it looks like there's 60% of them are valid and 40% are empty. Now remember, this is only doing the data sets around 30,000 or 32,000 rows, but it's only profiling, so down here on the bottom, Column Profiling based on the top 1,000 rows, so that's all we're seeing right here. If I wanted to see all of the data itself, now depending on how big it is, we may not want to do this. I can select this at the bottom and Column Profiling based on entire data set, and it's going to reload back into here; not sure how long it's going to take. Now going over, I can see there's 22,000 data sets of—for data points of the Salary Year where 10,000 are empty. The other thing that you may have enabled by now is that Column Distribution to be able to see what are the—what is the breakdown of distinct and also unique values. Investigating what actually distinct, unique means, I went back to the Job Title Short; looks like now it's actually picking up on all the different errors. I'm going to actually change this back; we don't want this to be Number for Job Title Short; we're going to change this back to Text, and I'm also going to refresh the preview by going to that Home tab, basically refreshing it to get it all cleaned up. Anyway, if we recall from our previous analysis, there's 10 different job titles of—Sat, Senior Data Scientist, Data Engineers, and whatnot, and so that is the 10 distinct values. They're distinct because they have repetitive values in here, like right in here in six and seven; Data Engineer appears more than once. Now if we go over to something like Job Country, they have 111 distinct, so meaning 111 countries that have multiple different countries and only 12 countries that have one value for it, or one unique value. All right, the last thing in Data Preview is Column Profile, and this is pretty neat. Right now I'm selected on the Job Title Short column; it provides one on the left-hand side, key statistics about the column, and then two, it actually shows a breakdown of the value distribution of it. So this is really helpful in performing EDA. If I wanted to go through here and actually see something, so I can easily go in and even see something like Job Country and see how United States has the majority of the values and then how the different other countries fall underneath that. Now this takes up a lot of room and sort of valuable real estate, so I find myself toggling this Column Profile on and off.
All right, last few sections in this View tab. Go to Column; if you have a large data set with a ton of columns, you can just come down here, select the column you want to go to, and then it will navigate you to it. Parameters; this is beyond the scope of this course; we're not going to be enabling parameters or even using them, so we'll call this na. Next is the Advanced Editor, which allows access to basically the behind the scenes of our AM—uh, M language, which we're going to be breaking down further in an upcoming lesson, so we're going to save that. But you can also access that from the Home menu in Advanced Editor as well. Lastly is Query Dependencies; whenever it gets into complicated ways that you're actually building your different queries and how they're connected to each other, this is going to come in handy. In this case, we're showing that, hey, we connected to that Excel file on my MacBook and we loaded it into a Pivot Table.
All right, next up is Query Settings. I'm actually going to go ahead and close this out for Queries over here. Anyway, with the Query Settings, we can actually change the name of the query if we want to. In this case, is named Sheet One. I don't really like that; I'm going to name it something like J Jobs, and I know it has salary data in it, so I'm going to have Salary. Down here on the Applied Steps, like we mentioned, this is a step-through, walkthrough of each of the individual steps that Power Query has taken.
To actually clean up our data set now, one thing I will call out in this: if I need to modify anything. So, in this case, if I wanted to modify the data source here, I could come inside of here into the formula bar and edit it. I would encourage you, if you're not familiar with the phone of the bar, with using that or comfortable using it, instead click this settings icon over here on the right-hand side, and then typically a window will pop up and allow you to edit it. So I could technically change the location of this or change what type of file it is. The same for navigation as well; I can basically pull back up that navigation window that I had before and change the sheet I wanted to. For the change type, this doesn't really have a gear icon next to it for us to edit, so we're about to go through and actually change it. But if we inspect the job posted date, we'll see that here: one, it has it underneath the type number, but then, actually looking at the column itself, it's a number value. Because remember, Excel stores ex uh dates as number values behind the scene. Well, we could convert this to a date by typing in date here, but you may not be comfortable doing that just yet. Anyway, with that, that's a great segue into the Home tab, into how actually we can change something like a data type with the home typ. We've already seen a lot of things already, right? We saw the close and load too; we also saw that I can go through and actually refresh my query query to make sure that it's fully loaded and up to date. If I have multiple queries, I can not only do this refresh pery, I can go to this refresh all, and it does refresh of all queries. We've already seen Advanced edited before; properties just allows us to actually go in and change the name of this query if you want to, and manage is more advanced; we'll be diving into that in a little bit. Similar to Under The View tab with goto column, we also have this option of choose column, and go do column; we can also just actually select a column if you will. So if I wanted to actually select job post to date or even more than that, I can just do that, and it's going to select it, and it's going to actually remove all the other columns, so which is not what we want to do. Which brings us a good point: if we want to get mid rid of a step, all we have to do is come over to the applied steps, and there's a red x mark that will appear over any step that you do. So I'm just going to go ahead and click X here, and it's going to remove anything that I've done.
Moving on to remove columns, which I think is pretty self-explanatory: if you want to remove a column, you just select it and you select remove column. Additionally, if I want to remove all other columns, so in this case job title, let's say I want to keep that, I could select remove all other columns, and it would do that. I want to cancel this step, so I'll click X. Similarly to remove columns, we have well keep rows and also remove rows, and then we have options for also sorting our values if we want to sort them from A to Z or Z to A depending on a column. So back to job post to date, maybe I wanted them in numerical order; I could just click A to Z, and it would go through and actually sort it. Anyway, I don't really want to do this; I'm going to clear this step as well. This brings us actually into what we want to do: we want to change this job posted date to a date time, and that's we're going to use underneath this transform section in the Home tab. Right now, this data type, as I'm selecting this job posted dat, it notices that it's a decimal number. I go to something like search location; it changes to text. So what I want to do is change this data type of decimal number to specifically a date time, because that's what we have in here; we have date and time. Now this popup is going to come up if you're doing this underneath the step that has changed type already. What it's noticing is that the selected column has an existing type conversion; would you like to replace the existing conversion or basically preserve that as a number and add a separate step? I'm just going to go ahead; we're going to do replace current, but I just want to show what it looks like of adding another step. In this case, I converted it in this step to a number, and then the next step I converted it to a date time. I don't like having a bunch of steps; I want to make this as concise as possible, so I'm going to clear that step instead. And instead, this time whenever we go through it and select date time, I'm going to say, hey, replace current. Now underneath here, it updated that job post to date type to date time, and it's all within one step; love this.
Similarly to that date time, I also want to convert the salary or average and the salary hour average columns. Right now they're decimal numbers, which is nothing wrong with that, but I actually have the option to change it to something like a currency. In this case, once again, I want to replace the current step for that. I'm going to do the same for salary hour average and change that to a currency as well for replace current. Covering briefly these other sections in the Home tab: first up is merge and append; we're going to be covering an entire lesson on this and how we can actually take different Excel files and different queries and combine them together. With this manage parameters is outside the scope of this course; I don't find myself ever really doing this, so not something we need to worry about. Data source settings, similar to what we saw outside of the power qu in Excel, basically the same popup is going to come here to allow you to change where your data source is. And then down here at the very end, if we have wanted to put in a new query, I wouldn't necessarily have to back out of the power query editor; I could just come in here and select a new source, a file or database or other source, and then work through actually importing it in in a query. Sometimes I find myself also using this one of enter data; say I had a simple table that I wanted to input into Power query to have, I could go through and just create that table.
All right, next up is the Transform Tab, and this one I feel is maybe actually, although it looks like a lot of options, it's probably one of the most simplest. As you can see, we have things like text column, number column, date and time columns, structured columns. Basically, if we have a data type of this, we're going to go to, you can go to; if I have a number column, I want to go to this and see what things I could do to it, such if I could do statistics to it, I could do rounding to it, or I could even get information out of it if it's even or odd. I also have this section on any column that basically applies to any column; this is allows us to, one, like we saw in the Home tab, actually convert the data type of something, but also even more advanced Transformations such as pivoting and unpivoting columns, which we're going to be diving deeper into in the next lesson on Advanced Transformations. And finally, we have this section on tables, which just does more of generic things to this data set, such as if I wanted to actually go through and count the rows on this could, and I find out I have 32,000 different rows on this. Anyway, I actually want to transform a column of this, specifically this job via column. As you notice from here, that all these different job platforms have via and then a space right at the beginning of it; I want to actually remove that. So in order to do this, I make sure that one job via column is selected; I notice up here in the any columns it has the data type of text. Now there are a few options in underneath the text column section for like splitting columns; I could split it by this half and then delete that via, but I find actually the easiest way to do this is just go through this replace values, and we're not going to do replace errors; we're going to just do replace values itself, and we find a value in here; in this case, we want to find VIA with a space, and we want to replace it with, well, nothing. If I wanted to go into advanced options, and I have a few different selections available, but neither of these applicable does, so we're going to just go ahead and click okay, and Bam, now we have these job platforms cleared up.
Now we've been going through this and keeping the names of these steps the same, but sometimes I like to be more descriptive; in when it's not a general tyag, now it named this new Step replaced values; I may actually do that a few times, and I want to be able to whenever I go back to this actually be able to identify what steps did what. In this case, change type, promoted headers, navigation, Source, those are all only usually typically done once, so I know what that means. However, however, for this one I don't know, so I'm going to right-click it and go to rename, and I'll say this is replaced via in job via, which is much more descriptive in my opinion. All right, only one more tab to cover, and that is the add column. With transform, we transformed a current column; with ADD column, we're adding additional column to this. Similar transform, it has these options for text, number, and also date and time, so very familiar features with this. So let's say I wanted to extract the month and the year out of the job posted date column; basically, I want a column for month, and I want a column for Year. Anyway, previously we learned with that transform tab, if I were to come into here under date time and then select something like month, it's going to transform this tab, so it's going to get rid of the contents of the job posted date; is not necessarily what I want; I want a new column, so I'm going to actually get rid of this Stu. So with ADD column, what this does is with that job posted date column selected, I select date; in this case, I want month; I could do start a month, end of month, day of month, whatever; I just want the month itself, and then inserted month is pretty descriptive. I, however, don't like the name of this, so I could come in here; this is an option, and change; I double-clicked on this and name this job posted month and then press enter. Now with this, I'm going to get a renamed columns here, so now I have two steps of this month was inserted into this, and then we rename the column. I would encourage you to minimize the amount of steps you have because these queries can get quite long. In this case, I'm going to delete this rename column; go back to this inserted month; if we actually reread this, you don't actually need to understand what's going on much in here, but I can see basically that we have this month in quotation marks, and this is named month, so I basically can reason that this is probably the new column title of this. So instead of using month, I'm just going to edit this in the formula bar to job posted month, then I'm going to click at the end and press enter, and now all within one step I inserted that month and renamed it as well. If you're not comfortable doing that, feel free to go through that next step of actually double-clicking this and actually changing it, but I would encourage you if you can actually try to mess around with the formula; if you make a mistake, it's pretty simple to just X out of that step and then redo it again, so there's no harm to your actual data set now. Similarly, if I wanted to create that job posted year column, I could just go through here, select year, whether I want start year, end of year, year itself; once again, it inserts year, and then I would want to change the name of this and change this to job posted year, and then click enter, and Bam, now we have it. I don't actually need this; all these are from 2023; I don't actually; this is not going to provide any useful data for me, so I'm actually going to delete this Stu.
All right, I want to do one last transformation before we actually load this and going to actually visualize this. So we have our salary year average column, and then also want to compare this to the salary hour average column, but right this is on a yearly basis; this is on an hourly basis. What we could do is do a conversion to our salary hour average column to get it to an equal value or comparable value to our yearly value, meaning we could put the number of hours in a year, multiply it times this value, and from there get what would be the expected yearly salary for this hour data. So I could do this via the transform tab, right, going into that number column under standard; we want to actually multiply, and then there's 2080 hours in a year working hours for 40 hours of work week. I could go through and actually do that, and that's going to update this column itself, but remember we probably want its own column, so I'm not going to use that. Instead, we'll go to add column; with this hour average column selected, select standard, multiply, put in those hours of 2080, and then click okay. Once again, I'm going to rename this; I can see that this multiplication column is titled this via in this step right here, so I'm going to rename it to salary hour adjusted, and in this case, I'm going to also rename this step to adjusted hourly salary to yearly. Now I'm sort of a stickler for keeping my data set in order; right now I have this job posted month, and it's sort of right away from it's pretty far away from my job posted date; I would actually want to move it right next to it. So there's a couple options I can do to move it; I can select the column and then come up here to the Transform Tab and move, go left, right, to beginning, to end, or I can actually just take it and then drag it; and this is taking forever; it's like paint dry, but find where I want it; boom, plant it in, and then inserted the step of reordered columns. I'm going to do the same thing with salary hour adjusted and put it right next to salary hour average, and both of these done with one step of reordered columns, so I'm fine with that. So now let's actually get into analyzing this; specifically, I want to be able to analyze and compare this salary hour adjusted column that we just created compared to the salary year average. So going back to home, I'm going to close and load this in; we have this previous analysis that we did before doing EDA on the jobs; actually want to create my own from scratch.
All right, so back on sheet one, we can see our queries connection, specifically that data job salary; remember the data tab, you can go into that, and it can toggle on all that queries and connections. Anyway, we want to insert; I want to analyze that hourly adjusted salary, so I'm going to come in to create a pivot chart; we also do pivot chart and pivot table at the same time. Anyway, when this pops up for pivot table or pivot charts, we want to; we're not going to select a table AR range because this is a power query connection if you will; we're going to use this external data source, and we're going to say choose connection; what connection do we want to use for this? Specifically, I want to use that DOA job salary, so go ahead and click that and open, and we're going to insert it into the existing worksheet. So now the pivot table set up for us; go forward to do one quick note; you may be tempted, say if we went back to jobs EDA, to right-click this and then go load to, and let's say, hey, I wanted to create a new pivot chart; well, the problem is is going to then get rid of this pivot table that we previously created, so you don't want to necessarily; if you want to keep this, you don't want to actually do that. Back to the pivot table itself, you'll notice now because we have these queries and connections, but you can toggle between the two over here on the right-hand side. Anyway, what I want to compare is that salary hour adjusted to that salary year average; right now it's doing sums; we don't want that; we do eventually; we're going to Value fail settings; we're going to do average here; we're eventually going to do median, I promise you, but we're going to STi for average for the time being. I'll adjust both of these to be of average; then I'm not really liking the formatting here; I know we adjusted it as currency back in the the power query, but this is the one data type that I find doesn't actually follow through in actually making into the correct data type when you import it into Excel, so you do need to go back still and actually convert it into the correct thing. Anyway, we're seeing that the hourly salary is much less than the yearly salary, and moving this over we can also see this via visualization; this doesn't really show as much; I would rather look at this when compared to job type, so I'm going to go ahead and grab job title short and throw it into the axis. Now closing out of this and then closing out of this on the side, we can now get a better view of this. I'm not liking the format of this pivot chart; specifically, I'm going to go in here, design, under change chart type, and change this to a bar chart; I feel like it's going to be easier to read; yeah, it's a lot easier to read. Also for these visualizations, I'm going to right-click this, and I'm going to say hide all field button, so that make this easier to view, and I'm going to go ahead and stick The Legend at the bottom. Okay, we're off to a good start; other things I want to do to clean this up is, oh my goodness, this is so long; I'm going to change these column titles to hourly adjusted salary and then yearly salary. Additionally, I want to sort this a little bit better; specifically, from high to low; so under sort options, more sort options, I'm going to go into sorting this as sending based on the year L salary from high to low; sorry, that's actually descending; selecting year salary, clicking okay; no, it was right the first time; it's ascending. Okay, this is looking good; you know, also I don't like having different colors; I like actually going with a consistent theme, so going into design, change colors, I'll change this to this monochromatic pallette 8, and Bam, we now have our final visualization that we use power query to basically ingest all our data in, clean it up, create this new column of hourly adjusted salary, perform an analysis in Excel to average it, and we can see that consistently the hourly salary is well below that of the yearly salary, so I guess it pays to have a salary job.
All right, we have some practice problems for you to now go through and test out all these different features and get more familiar with the power query editor. In the next lesson, we're going to be going into advanced Transformations and diving deeper specifically in analyzing skills and using power query to actually clean it up so where we can actually analyze skills with that; see you in that one. All right, welcome to this lesson; we're going to continue on with power query, specifically focusing on using more advanced Transformations, and for this we're actually going to get into analyzing those skills and being able to put them on a graph and actually visualize what are the top skills of data nerds. Now, if you recall way back in the functions and formulas chapter when we went over text functions, we did a little bit of text cleanup to clean up this column and then plot it, but we were only able to do that with around 20 rows. Now, with the power of power query, we're actually going to be able to clean up all these values and be able to visualize it for all 30,000 job post. So let's jump in; if you want to, you can continue on from that worksheet that we used in the
Previous lesson and just make sure that you do go through and actually save it before you continue on. However, if you got lost in the way or you just don't have that file anymore, feel free to use the lesson or the file from the last lesson of power query. Eder, once again, you don't want to be using the actual one working cuz that has the final results; we're going to want to work with that one, and this has all the different work that we did. It also has some some additional analysis whenever I looked at plotting it over time to see if how the salary of yearly versus hourly compared. Anyway, let's get into editing this, and we can get to the power query editor by going up to get data, launch power query, or pressing Alt+F12. Once it loads and need to click on the query that I actually want to look at, and I'm going to close this or minimize this.
The first thing that I want to do is start an index column on this data set because in general, whenever you have a source data set or a fact table like this is, you want to have an index associated with it. Yeah, these row numbers are good, but that's not good enough, and we'll be using it more in the power pivot chapter, but it's good practice to start it now. So moving over to the add column tab, I'm going to go to index column. It allows us to start from either zero or one; I'm a coder, so I like from zero. Now, Pro tip: I want this index at the front. Now, I could go to transform and then move and then move this to the beginning, but remember we did this reordered columns right here, so what I'm actually going to do is take this added index, put it before reordered columns. Now that the reordered columns is right there whenever I select this index and move this over to beginning, it's going to be included in part of this step of all of our column reord, so I don't have once again multiple different reordered columns all right.
In order to clean up this job skills column, we're going to end up being putting this uh these skills right now; they're separated by column inside of this list. We're going to be breaking them up into their own individual rows, and because we're breaking this up into different rows, this now is going to put for this Row one value here; this is going to make 1, 2, 3, 4, 5, 6, 7—this is going to make seven different rows of data. This is going to mess up anytime we want to analyze anything because imagine if you have like a salary data; it's then going to appear seven times. So the main point of explaining that is we want a new query to actually populate and actually break these skills out into their own separate rows. So in order to create a query or another query, right now we have queries one to create another query from this; we have two options, and that's underneath Home tab; they have manage, and we can either delete a query, which we're not going to do; we can either duplicate it or reference it. I can also get to this by just right-clicking the query, and it also has these of duplicate and reference. Let's actually look at both of those, starting with duplicate first. So I've created my duplicate query, and as you can see, it basically has a duplicate of the original query; nothing really has changed from it. Now this is cool if I want to walk through all the different steps again, and I wanted to have it in this new query, but I actually like this other option, so I'm going to go to data job salary this CL, I'm going to go down, select reference. Okay, this query, this one named three, is referencing data job seller, and it only has one applied step. If we look at the applied step, all it is doing is referencing the data jobs salary, so this first query right now and populating it for us, and this is really good because say now I make changes to the original query, such as say I want to go through and I don't want any any more of the hourly data in here; I only want the yearly data, so I filter down to only have the yearly data. So now it's filtered these rows for the yearly data. Don't worry; we're actually not going to do this; I'm going to delete this Stu, but anyway, if I go to that duplicated query, the one with the three at the end, this one only has year values in it. This I can verify is 100% yearly by looking either the column distribution or the column profile; everything is your. Anyway, we don't actually want to do that step, so I'm going to go back to this original query, clear the filtered rows, and once again, it's going to just clean this back up to have two distinct values, so compare checking the S rate yearly and also hourly. Okay, so we like the reference for our case cuz I like we may make changes to the original one, so I'm going to delete this number two because remember that was the duplicate, and we're going to keep the number three one which was the reference. We're also going to be doing all our alterations on the skills on this one, so I'm going to to rename this one data jobs skills.
With this new query, data jobs skills, let's actually get into cleaning up this column of data of job skills. Specifically, we're going to be separating this into each of these skills into the new rows by this comma delimiter, but we need to remove a few things from this; specifically, this has brackets around it, and it also has single quotes; we don't need any of that; we need to remove it. So going to that transform tab, we're going to go into replace values, and we've done this before, so for the value defin, I'm going to just start with the first square bracket; we want to replace with nothing; I'm going to click okay. Additionally, we want to replace the other bracket as well; replace it with a blank, and then finally, we want to replace that single quote as well. Also, I'm going to just rename these all. Next thing we going to do is actually split these columns on this delimiter of a comma, so under transform, we can go here to split column; it has a few different options by delimiter, number of characters, by positions; we can go to by delimiter; I'm going to select that. For this, we're going to use a comma delimiter because there's multiple different options you could potentially use for this; we want to split at not just the leftmost, but we want to split at each occurrence; there's no quote characters in here; we removed all the quote characters, so I'm going to click none and then click okay. So now we just split these skills into let's see how many different columns we have here; looks like we have up to 24 skills for all these different skills that we have. So now what we need to do to get all of these, if you will, skills within a single column, we need to unpivot them, but the one issue right now, so I have all these skills right here, but we also have all these other columns right here; I don't really care about all them; just I don't really care about around too much; I want to mainly just analyze job title short and indexed. So what I'm going to do to make this easier because I need to basically select which columns I want to remove or which ones I don't want to remove in this case, so what I'm going to do is go back to source, and this one has before we actually broken up the job skills, so I'm going to select job skills, hold down control, and then from there select job title short and also index, and then underneath the Home tab, we're going to go to remove calls; what we're going to do, remove other columns, basically going to keep those three columns that we have. Now we are doing this in the applied steps after that first step of source, so it's asking, hey, do we want to insert this step? Yes, we do, and so now we've limited it down to those three columns, and Bam, now whenever we go down here, down to that last step of change type, we can see that we have all our different job skills and then over on the right-hand side we have our index and our job tile short, which I don't really like the order of this; I'm actually going to go back to reorder this over here; I'm going to just take these column values and then put them in this order of index, job title short, and job skills. So now we actually get into unpivoting these job skills columns, basically making all these job skills into one column, so I'm going to select instead of selecting all the job skills column; I'm actually going to select the opposite, holding control, select the index and job title short, and I'm going to go into to transform tab into unpivot columns, and for this one, once again, we're going to use the other; we want to unpivot other columns and go ahead and do this all right. So what we do here, we now have this new column of attribute and value; attribute, if we go back, that is just the name of the column that was created previously, and then the value is what was in the cell itself, and that's filled with all the skills. So personally, I don't really care for use of this attribute, so I'm going to go ahead and just remove this column by right-clicking and selecting it. Additionally, I'm going to go back up here, and I don't want this to be named value, so I can go in and inspect this under unpivot other columns; I can see in here that it renames these columns attribute and value; in this case, I don't want to be value; like I said, I want to be job skills; clicking enter, boom, renamed it to job skills, and then in here it is job skills.
Now, one thing that's bothering me real quick before we continue on to actually visualizing this data is this column here; typically, I like to name things something like job uncore whatever it is; in this case, index; I want to Name jobor ID, but if you recall back we created this back in this data jobs salary portion, especially here under the step of added index; I want to change this from index as we've done before, going in and renaming it to job ID. However, whenever I do this, press enter, this is going to break my queries, and this is going to happen to you anytime you're manipulating it, so I think we need to get familiar with it. So if I go to the next step of reorder columns, we're going to have this expression error: the column index of the table wasn't found, duh, because we named it job ID in the previous step instead of index, but this step is still the same, so what I can do is come in here, change index to job ID, press enter, and Bam, that updates, but then now going to data job skills, we're going to have the same thing; you're going to notice with this one right, the column index the tail wasn't found index, so same error message; what we want to do, you can do is go to error; it's going to go to the first occurrence of that error; in this is trying to reference index; we, if you call back from, if we go to the first step of source, we expect it to be called job ID now because we renamed it right, so I'm going to change this to job ID and then scrolling through the applied steps to see whenever we get to our next error if there is an error, and that's unpivot other columns; specifically, they have job title short and index; I don't want index here; I want job ID, and now bam, now we have it cleaned, so I should have done that job ID, but that was actually good troubleshooting to walk through that you may encounter. So let's actually get into visualizing this, so we're going to go to home, and we're going to close, and we're going to close and load. Now it's popping up as a table, but we actually want to analyze this; I don't really care to have it as a table, so I'm going to right-click it, and I'm going click load to; specifically, we're going to go to a pivot chart, and we'll insert in the existing worksheet because we're going to get rid of that data; yes, there's going to be possible data loss; we understand that, so I'm going move this chart off to the side, select inside the pivot table, and we want to analyze the job skills, so I'm going to take the job skills, put them in rows, and then the job skills also in the values to to count up the values, then also I'm going to sort them; I want to sound them from high to low, so I went to more sort options; um, we're doing a descending order count of job skills, so now there's a ton of different skills in here, but want you to inspect this; if you notice one these skills have sometimes have spaces in the front of them; basically, we didn't do a full cleanup of this, so that's why we have python twice in here is cuz this one has a space of it, so opening up the power query editor by playing by pressing Alt+F12, so underneath the data job skills query, I'm going to go ahead and we want to do a text transformation; specifically, if we look underneath this, underneath for format, we can change this to lower case, upload case, capitalize each word; we're going to do trim, which removes leading and trailing white space from each of the cells in the selected cell; from there, we'll go back to home, close and load this, and now it's going to be reloading the data, and those duplicate values are now going to be removed. Now there's a lot of skills here, so I really only want to see the top 10, so I'm going to put a filter on here, go into value filters, and top one; specifically, want to see the top 10 items by count of job skills; also, I'm going to rename this to skill count, and because these are text values down here, I'm actually going to change this from a column chart, going to change chart type into a bar chart instead; clicking okay, boom, and then with this, obviously, it's not sorted from high to low; that's how I want actually to sort it, so I'm going to go in here back underneath our more sort options, Chang this from descending to ascending, and the good thing about this is we still have that top 10 filter on it, so it's still going to apply this and have the top 10 values on there.
First last little clean up, I'm going to hide all field buttons; I'm going to get rid of this Legend right here, and and then I'm going to rename this to What are the top skills of data nerds? Now let's say that I'm frequently referencing the top 10 skills as we have right here, and instead of having to populate this every single time, I want to actually create a own or create a query for this, so opening power query, going to Alt+F12; I could do the same analysis inside of power query query and get this into its own table to be reused, but for this, I don't want to use this data job skills query; instead, like we did before, I'm going to create a new query; we're not going to duplicate this; instead, we're going to reference it, so now it's Unique and distinct, and I'll rename this data jobs skill count because we're get the top 10 and their Associated count. So in order to do this analysis to find what is the count of all these different skills, we want to do a group buy, and it's right here under transform form under that Home tab, and I can do group by, which group rows in the table based on the values in the currently selected column; we're going to be forming a basic Group by; we're using that job skills column; I could change it to another column if I wanted to, and that new column name is going to be skill count; operation, we're going to be counting the rows; we could do any other type of aggregation as well; if we had numerical data, we could do average, median, min, max, whatnot; go ahead and click okay. So we've done this aggregation; now the next thing is I just want to get the top 10 values, but before to do that, I need to actually sort this in descending order; right now, I can tell looking into the numbers, this isn't necessar, although it looks like it isn't right, so clicking the arrow up at the top, I'm just going to say, hey, sort descending, and then we want the top 10 values, so underneath the Home tab under keep rows, I'm going to have keep top rows, and it's going to prop me how many number of rows do I want to keep? 10; in this case, I want the 10 values, and now from here, all I got to do is close and load this into its own separate query, and Bam, here we have it, and so if I needed to reference the top 10 skills anytime, all I would have to do is just reference this query, and I wouldn't have to, like we did last time, go through this full analysis, so power of query is really great at automating some repetitive analysis and having it just ready for you.
All right, last little cleanup; if we look at these skilled names, they're not formatted correctly; specifically, if I look at something like SQL, I expect to be all capital letters SQL, capital letters python; I expected to be Capital At the beginning python, so we're going to go through and actually fix this so that way whenever we present our data to someone, it doesn't look like a hot mess, so opening up the power query menu by pressing Alt+F12, we're going to go into the data jobs skills query; specifically, on that last step on, and we're want to alter the job skills column, so the first thing I want to do with this text cleanup, the easiest thing looking at this is we just need to capitalize the first letter of every single word, and then from there, we'll go through and actually fine-tune it to capitalize in case of SQL, capitalize all letters; we'll have to put in special case for this. Anyway, if you recall from before, we have that transform format, and they have this capitalize each word; we're going to do that. The next thing though, the more complicated one is we're going to go into add column, and we're going to add a conditional column, so what we're going to do is go through; we're going to keep the the name of custom column cuz we're technically going to be since we're adding a column; we're going to have to go and delete this job skills column once create this new one; I don't want to name a job skills right now; going to call MK. Anyway, what we want to do is we want to select the column that we want, so if job skills equals, in this case, we expect to equal something like SQL, we want the output to equal SQL; then if we want to add more conditions or Clauses to it, we go to add Clause; once again, I'm going to select job skills, and I'm going to put something like powerbi; it had a lowercase ey at the end; I want the powerbi to be fully capitalized at the end; I also went through and added some other ones such as AWS, GCP, no SQL, and SAS; most all these required them to just capitalize fully except for the no SQL one; then what do we want it to be if it's not any of these conditions? Well, we'll add this else clause, and we want it to be basically the results of an entire column; we want it to be whatever it is already in the job skills column; I'm going to go ahead and click okay. So now we have this cleaned up data set as well with nice looking names. Now, if you want to if you're going through and finding anything in here that you want to clean up, feel free to add to that conditional column statement; those are the ones I'm just going to go for right now. Anyway, because we added this new column, and I don't really know an easy way to do this without actually creating this new column, we need to now go ahead and remove job skills and rename custom, so going to the Home tab, I'm going to remove column; I'm going to remove the one that's selected, and I'm going to renames custom to job skills, and conveniently, because we're using that same name and just replacing it, if I go to the data jobs skill count, that one because it references this one will also get updated, and all those values in there are updated as well. Anyway, let's go ahead and
Close and load and inspect. This is our previous pivot table and pivot chart that we analyzed. It's now going through and loading all the data, and now we have it updated with all that correct formatting for those different data points.
One last thing before we go: this is generic. These top skills of data nerds—tall data nerds—and that is using the data job skills query, which has the job title short column in it, so we can actually visualize this for a certain job by going into pivot chart analyze. I'm going to go into insert slicer; specifically, we're going to look at job title short. I'm going to put it over here, and then, as usual, I'm going to rename it real quick to job title. Now let's say we want to analyze something like data analyst. We can see that SQL is the top skill, but Excel is in second place, followed by Python, Tableau, and SAS. What about for business analysts? Very similar in that SQL's top, and then Excel is in that second place. So really unique and showing the importance of Excel within these skills, and pretty meta that we used Excel to find this out.
All right, now it's your turn to give it a shot. You have some practice problems to go through and get more familiar with doing these advanced transformations, specifically pivoting, unpivoting, and then also group by.
All right, with that, I'll see you in the next one. We're going to be diving into append and merging queries; specifically, going to be doing this with that skill query that we did previously. All right, see you there.
Let's now get into how to perform append and also merges. And so the first portion of this lesson, the easiest portion in my opinion, is going to be append. Specifically, going back to that Excel sheet where we had all those different uh sheets for the months of the year and they're job posting on each, because all these data sets are of the same format, I have the same columns; we're going to be able to append all these together and get what is our final data set of all 30,000 rows. If you recall, each month had around 3,000 postings, so that's how we get to that value.
From there, the primary focus of this lesson will then shift to merge. For this, we're going to be combining our two queries that we built previously: one which was our original data set, so we titled that one data jobs salary, and then that new query that we created in the last lesson on the skills, so data job skills. We're going to be merging those two together, and this will allow us to do some pretty interesting analysis; specifically, now that we've merged those, we'll be able to see, based on a skill, what is the expected salary, and we're going to build a visualization for that for the top 10 skills.
Now merge, unlike append, is a very complex operation, mainly because there's a lot of different types of merges; specifically, there's six types of merges in Microsoft alone, so we're going to be walking through each one of those so you understand the differences and know which one to use when.
For this first append example, we're going to be using this data job salary monthly data set, and just as a refresher, this contains everything for—in this case, I'm selected on the January sheet down here—and this has all the January data, which has around 3,100 rows for this, and we have each one of the months for the year here. Anyway, let's use Power Query to append all these together, because previously, before you knew about this, you'd have to go through and actually copy and paste all these different options right here and then put it into a new sheet. Doing this 12 times is a hot mess. So since this is only a simple example that we're not going to use later on, I recommend just opening up a new workbook for this.
Now coming into the data tab, I can come down to get data, and they do have this option right here for combined queries, merge and also append, but this is for append two queries from within in this workbook; it's basically assuming you've already imported it in. So instead, what we need to do is actually go to from file and actually start our first query of connecting to that Excel workbook with all those different sheets. Navigating to the course underneath resources, data sets, and then here down on data job salary monthly, I'll select that, select Import. In the Navigator, we can see all the different sheets that are available. We want to actually do enable this of select multiple items and then go through and select all the items. With all these loaded, we're going to then shift into not just loading it; we want to actually go into the Power Query editor, so I'm going to select transform data, and it's going to start by loading each one of those sheets and just going to be naming each one of the queries respectively after those sheets.
With Power Query editor launched, we can see over here in the left-hand pan all 12 of those queries for each of the months. So these are all their separate own queries. Because of that, we need to now move into actually appending them and make it one final query that we can actually export into or import into Excel. So underneath the Home tab, they have the option for combine, append queries; they have append queries, and append queries is new. With the January query selected, I'm going to go to append queries, and for this I can say either do two tables and specify the table I want to do. We're going to do three or more, cuz we want to do all of them. With them all selected, I'll now go through and click okay to append.
Now this inserted a step of appended queries inside inside of that January query, so now that January query is all those different data sets. So I just want to verify that I got all the data in here right now. If we scroll down—well, I'm just going to show it right here—we're only showing column profile based on the top 1,000. The fastest way to actually find this out is just go to the transform tab and go to count rows, which it tells me there's 36,000 rows, which it's a few thousand too many. And if I go back into the appended query option and actually look into it, I can see in the formula bar we have August in here. I accidentally selected it twice, so I'll go ahead and delete it and then look at the counted rows; that's actually what I expect the value to be, around 32,000. Anyway, that was just to count the rows. Additionally, I don't want the append query to be inside of that January query, so I'm going to delete this step as well. Instead, with the January query selected, I'll go back to that home, append queries, and then select append queries as new. This is going to create a completely new query. Once again, we want to do three or more tables. This time I'm going to hold control and select all of them and then move them over at once. Make sure we don't have duplicates this time. So this now starts a new query. Right now it's called append one; I would probably name it something like data jobs all, and then pressing enter it then loads in here, but you can see these queries—like imagine the case where I—right now we have 13 queries. I want to organize these a little bit better, so we can actually group these; specifically, we can group these monthly ones. I selected April and then, holding control, selecting all the other queries as well, then right-clicked it, and I'm going to select this option to move to group. We need to have a new group, and I'll call this real uniquely data jobs monthly and click okay. So now we have these two folders: one with data jobs monthly. I'm going to close that down, and then there other queries which we've seen before, and there's one query inside of this of data jobs all. This cleans it up. Also, you may get this disclaimer up here: the preview may be up to 33 days old. Feel free to refresh it if you've been getting that; should have no effect on your data.
Then, if we wanted to, we could go through and actually analyze this by pressing close and load to. I pretty maturely selected close and load; I recommend you select close and load to. Anyway, nonetheless, I'll go to the data jobs all; we'll go to load to; specifically, I want to look at a pivot table. I know there's going to be some data loss because it's going to remove the data in the sheet, and then I can inspect that job posted date; specifically for the account, dragging job post date into the rows and then also dragging job posted date into values, and once again this is why we double check it. This time it looks like I accidentally imported in January twice. With this, as we can see that it's 35,000. Anyway, opening up that Power Query editor, going to the data jobs all query and updating it to remove that second January that I should have caught from before and then close and loading it, and now it should refresh and update for these these correct values. Boom! So now it's actually aligned with what I expect to see. This is why we always double check any type of query or analysis you do. This double check of the work is going to save your butt.
All right, let's now get into the bulk of this lesson. I'm moving into merge. For this, feel free to continue working with that workbook that you were working with in the last lesson. If you didn't happen to save it or you got lost, you can use the advanced transform workbook from the last lesson; that'll pick right right back up where we left off. And then, as usual, the append and the merge are the final examples that you're going to see at the end of this, which specifically for append you've already saw. So let's actually get into merging those queries. For this, I want to press Alt F12, and right now we have three queries in here: the data job salary, which is basically like our fact table; this includes all of our data. Going into transform and count rows, we have, as expected, around 32 data point points. I'm going to go ahead and delete that step. Similarly, we have this data jobs skills, which has all of our skills in it. Let's see how many rows are in this by going up to transform and to count rows, and this has 167,000. Now it's important to understand these numbers because we're going to be using them or need to understand them whenever we actually get into the joins to see when we have missing or more data. So I'm going to go ahead and delete the step of counted rows as well; we don't need it. Then we have also this final query of data job skills count. This was made as an example only; we're not going to use this any further into the future, so I'm actually going to go ahead and just delete this to minimize my queries. It's going to ask them—I'm sure—want to delete it—yep.
So let's get into merging these queries. I have data job salary selected, come up to the Home tab, under merge queries, we're going to have merge queries, and merge queries as new. Like we learned from the append of append queries and append queries is new, we're going to want a new query so that way we still have these source queries. So I'm going to go merge queries as new. With this, this merge window pops up, and it says select the tables and matching columns to create a merge table. Specifically, we want to go with the data jobs salary, and we want to merge it on the job ID; that's why we created that a few lessons ago. We're trying to connect to the data jobs skills on also that job ID. Now down here underneath this, there's a join kind, and there's six different options from this: of left outer, right outer, full outer, inner, left anti, and right anti. Now Kelly put together this fancy chart that shows visually what is happening with these merges, and we're going to be walking through all of these briefly in order to understand which type of join you should be choosing depending on which scenario you're in. As a quick overview, these circles are signifying the two different tables, so in this case, table A and table B, and the shaded blue area shows what portion of the contents from those tables will be included in the final table.
First up is a left outer join, and with this join what's showing here is that all rows from table A will be included in the final table, and then from that center portion right there where A and B overlap, this signifies that it's only going to keep items from table B that are in table A or match with table A. So what does it actually mean? So if we go here into join kind and select left outer, and then what we get told based on this next to this check mark is the selection matches 29,000 of 32,000 rows from the first table. So what are those missing jobs? Well, basically there's some jobs that don't have a skill. Now this isn't necessarily a bad thing, although we're not going to go with this join; this could be an option we could use. I'm going to click okay to load it in. So right now we have it under this query called merge one, and as you can see there's not repeating any job IDs; basically, we have the original data jobs salary table, and then we scroll all the way to the right, we have the data job skills over here, and if you see each one of these items is a table. If I click on it and expand it to see, hey, what's in this table, we can see that for this one there job posting or job ID of 10,001. This is the table associated with it. So I'm going to go ahead and actually delete out of this step and go back to it. So what we could do is expand it out, and there's this icon up in the top right-hand corner. I'm going to go ahead and click it, and it's going to ask me how it wants to basically expand out, and in this case I already have the job ID, I already have job title short; I would expand it by job skills. So now seeing how these skills are broken over, I can actually scroll all the way over and see that now 10,001 ID is duplicated multiple times. And if actually looked at the number of rows within this data set, this new data set, we have 170,000 rows. Now technically this merge has exactly what we want, but we still need to go through those other merge examples to understand them, so we're going to show them as well.
Now for this, I want to go back to that merge window, and I'm going to click the settings icon. I need to get rid of the step; we're going to be trying out different types of merges, so I'm going to exit out and then go in here and click the gear icon. Now it's popping back up. We did left outer; next thing we're going to look at is right outer. For right outer, this takes all of the rows out of table B, and then from there any that match those rows in table A are included. Now this one, when we look down here, it says, hey, the selection matches 167,000 of 167,000 rows from the second table. If you recall back from that left outer, we had 170,000, so 3,000 higher. Why is that? Well, that table A or data job salary has 3,000 roles in here that don't have any skills listed; hence why 3,000 is less. This provides a similar type of merge that we did before where we need to actually go over to that data job skills and expand it out, selecting the job skills column, and with this table we can just check that we have 167,000 rows, which bam, we confirm. All right, I'm going to get rid of these two steps; we're going to move into the next merge.
Next is inner join, and this provides only matching rows from table A and matching rows from table B. So depending how you're join it, there could be missing data on both A and B. For this one, it's saying, hey, the selection matches about 29,000 of 32,000 rows from the first table, which what we expect, and then basically all of the rows from the second table. So this one, if actually go into it and then expand out those data job skills, looking only at the job skills column with it expanded out, actually counting the rows, we have once again 167,000, so missing that 3,000 of jobs that don't have skills.
Next is left anti, and in this case it checks to see what matches it doesn't have and returns the value for that; specifically for table A, whichever values don't have a match, it's going to return that. So in this case, it says the selection excludes 29,000 out of the 32,000. When I go to load it, I get the rows from table A or data jobs salary, and it still has the data job skills, but actually if I looked into here right, we should be matching on things that don't match or don't have a value; specifically, there shouldn't be inside anything in this table that I'm clicking on, and as expected they're null values because it doesn't have skills. So exiting out of navigation, going back to source, counting these rows, we can see that we have 3,000 jobs basically with no skills.
For right anti, this gets rows from the right table that do not have matches in the left table, and for this, with right anti selected, this selection excludes 167,000 out of 167 rows from the second table, so basically everything from this table is included. We're not going to walk through this in the Power Query cuz this is also not what we want.
The final one we're going to actually use is a full outer join. From this, it takes all rows from table A and all rows from table B, and if there's a match it will join those two; if there's no matches, it's still going to return them in the table; it will just be a null value for where it doesn't match up. And this talks about how basically selection matches 29,000 of 32,000 rows from the first table and all the rows from the second table. Loading this in, once again we have data job skills; we need to expand out, and we only want to expand out those job skills, and then from there just going to do a double check; I'm I'm going to do count rows, and this has 170,000 rows in it, so similar to our left outer, we could have done either of these; these are one the twos that we want, but I'm going to stick with this one of the full outer because I have all the work here, and I'm going to close out the step. And I think that's a great example of sometimes there may be multiple joins that fit the example; it's important that you go through and actually count the rows and understand the data set to figure out which one you need to use and for what purpose. Anyway, one thing I glossed over real quick, going back to source and that gear icon is right underneath this; underneath the join kind, they have used fuzzy matching to perform the merge. Right now we're doing basically exact matching, as the job ID of 10,001, we're matching up exactly with the 10,001 from the other table. Fuzzy matching allows you to connect to tables that have basically non-exact matches, so in this case we have table A with a student ID and a student's name and only their first name, but then in table B we have the student name full, so first and last name and the grade. With the fuzzy matching, we could merge table A and B based on that student name first column and the student name full column. Now what happens if we get to where we have students with multiple similar first names? It's going to create a hot mess, so I don't always recommend using this unless you know the data and you know you're going to cause complications with it. So that was a quick overview of of the different joins within Power Query. If you want a more in-depth tutorial for how this is done, then and you can check out my SQL tutorial where I go through it with all the different SQL analysis that we do in that course and break it down step by step. I'll include a link to that video right here for you to go and see it.
All right, so we have the final table that we actually want for this. Remember these do have duplicate values in it, so you have to keep that in mind anytime you're doing analysis. I'm going to rename this as data jobs merged. One last thing for close and load, we have this job skills column, which is sort of redundant right now because we actually have the data job skills, not job skills, the actual skills itself, so
I need to get rid of this column. I actually want to do this. I'm going to do this in the source step before we even break this out. So I'm going to select job skills and select remove columns. It's going to ask if I want to insert the step, which I do. And then after we remove the columns, we go into expanding it out. And because we did it in that order, I can actually come in here instead of renaming it here; I can just rename it via the formula inside of expanded skills and just change it to job skills. And bam, now I only added two steps vice one. All right, go ahead. Now we're going to close and load two. I'm going to want a pivot table and also pivot chart, so I'm going to select the pivot chart option here. And underneath queries and connections, it's going to show that it's loading this in here under data jobs merged. So let me show you what we're going to be creating with this. I want to build this visualization that's showing what is the salary of the top 10 skills, top 10 skills by count for data nerds. And this is a combo chart. We're going to have not only the salary, or the average salary for a skill, but also for this line portion we're going to have the associated count for the number of skills that appears, or how many jobs it appears in. All right, so I'm going to go ahead and move this pivot chart out of the way and select the pivot table. Remember, we want to use the job skills; we're going to be analyzing that. So I'm going to throw in the rows; the first thing I'm going to look at is the easiest is the count of these job skills. And I'm going to rename this to job count, along with changing the value field settings, going to number format. I want to change the number specifically; I want to use a thousands separator with zero decimal places. I'll go ahead and press okay. So we have a count. Now we want the average salary, so I'm going to take salary, your average, drag it into the values. Right now it's doing a sum, so I'll go into value field settings, select average, and then for number format we're going to do currency with zero decimal places. Click okay and okay again. And I'm going to change this one to average salary and then specify the units of USD. All right, so now X-ing out of this and X-ing out of this, now our pivot chart is sort of all jacked up—well, it is jacked up—mainly it's trying to PR this as like a dual column chart, and that's not what we want. So we're going to change this design of it. Going to design, change chart type. I'm going to go over to combo, and then underneath here for the combo, for the job count I want that to be a line, so I'm going to go up here and select line. And for the average salary, I actually want that to be the column. Now I want the job count on a secondary axis; I don't want the same axis as the salary itself because they're just not proportional. I'm going to go ahead and click okay. I want to clean this up a little bit further by removing the legend and then also right-clicking here and hiding all field buttons on this. Okay, there's—now there's still too many skills on here. Remember, we want the top 10 skills, so going into the pivot table itself, I'm going to come up into the filter, into value filters, and top 10. We're going to do top 10 items by job count. All right, this is getting a lot more readable now because I have the top 10 by job count. I want to order this from high to low by salary, so I'm going to go to more sort options, and we're going to do descending on average salary. I'll click okay, and bam, now we're getting somewhere. So we're seeing things like Spark and AWS have the highest, and Excel did make the top 10, so it's on there at 100,000. Other things I'm going to change, selecting on this pivot chart is the actual design itself. You know how I am about colors, so we're going to change the colors. I'm going to use this monochrom MAAC palette 8. I want the line to be a lighter color than the actual bars itself. I'm going to go ahead and add access titles for primary vertical and secondary vertical. For this, I'm going to just select the box, go into the formula bar and say, hey, for this one, make it equal to average yearly salary. For this one, selecting the box, going into the formula bar, pressing equal, I'm going to make it equal to job count. I'm also going to add a title to this. I'm going to call this "What is the salary of the top 10 skill of data nerds?" And remember, this is for all data nerds, so I want to be able to actually—what's the great thing about this of joining these tables? Now we not only get salary data, but we can get job title information. So I'm going to add a slicer now, but going in pivot chart analyze, insert slicer, add in that job title short only. Going to move that out of the way. Now I'm going to go to slicer. I'm going to rename this to a more friendly title of job title. And now—now let's actually look at it for data analyst. So with this, looks like Python's at the highest. Excel still makes that top 10, and for data analysts at 86,000, it's also—if we look at this, it's the second most important skill behind SQL, which has a value of 96,000. Let's see what it is for a business analyst. Once again, SQL and Excel are two of the highest, and for business analysts, Excel is paying 87,000. So bam, we just showed the power of well append, but also more specifically merge. We can now take this analysis to another level, analyzing skills to other data points from our main fact table or that data jobs salary table that has all of the data in it. So now you have some practice problems to go through and get more familiar with using both append and also merge. After that, we'll be jumping into the last lesson of Power Query focusing on the M language. As I warned at the beginning, don't worry if you don't have coding experience or anything like that; we're going to be taking it nice and easy, and you're going to be able to follow along and fill it out pretty easily. We're going to be doing some final prep before we finally send this data set on over to Power Pivot, which we're going to cover in the next chapter. All right, with that, I'll see you there.
Welcome to this final lesson on the M language, and we're going to be going into some pretty advanced techniques and understanding how to read and better utilize the M language in building your Power Query queries. Anyway, nothing in this lesson is going to be used that we actually go through and do, used to build on our project. So if any time you're not following along or you're not able to do anything, don't worry too much; nothing's actually be used. It's more to inform you about the M language so you get more familiar with it. As a disclaimer, you will not be an expert on M language; you will not be able to code in M language after this. Mainly you'll just be able to look—look at it, understand what's going on there, from there and make slight adjustments if necessary. Feel free to continue working on in that worksheet that you've been using previously, where we just calculated in the last lesson, looking at the top 10 skills and what the salary is for them. However, if you got lost or wasn't able to follow along or just starting over, feel free to use this merge notebook. Don't use—once again—that M language one; that one's going to be what is going to be done at the end of this lesson. So what are we going to be covering in this lesson? Well, if you open up the Power Query editor, we can navigate into it. We're going to be covering three main things: first is the Advanced Editor, actually walking through a previous query and understanding how to read it. And then from there, under add column tab, we're going to go into these different examples on creating custom columns and also custom functions. So what exactly is this M language? Well, if we dive in documentation, we can see that the Power Query engine uses a scripting language behind the scenes for all Power Query Transformations; the Power Query M formula language, also known as M. So although we're doing all these edits inside of this Power Query editor, behind the scenes, if we navigate something like the advanced editor, it's actually using this M language right here to carry out all the Transformations. And it goes on to say, if you want to do advanced Transformations using the Power Query engine, you can use the advanced Editor to access the script of the query and modify it as you want it. It even goes on to discuss that if you're not finding what you need in the actual GUI, or the graphical unit user interface of the Power Query editor, you can use the M language, editing it in the advanced editor for this. So let's go into breaking down this M language more by going to that data jobs merge and entering the advanced editor, and we're going to be just breaking down this simple query right here. Up here on the right-hand side, there's a few different options, display options. I'm going to do this render whitespace; basically, it shows me the indentation that's going on here. Right now I'm seeing that there's four spaces in here. Anyway, the key thing here is we've—have first have this let keyword, and then in keyword. This begins the basically definition block, if you will; this whole portion right here for defining different variables and specifically different tasks. If we look, we have things like source, expanded data, job skills, sorted rows, remove column, remove columns. If I go ahead and move this over to the right, those applied steps are the same thing; those are the variables itself. I currently have enable word wrap enabled, and I'm not liking the format and how it looks. I'm going to go ahead and unclick that. Finally, we have the in keyword, and then this displays the final value that we want to appear for our query. So in this case, we want the final value of rename columns, or the last applied step, to be what appears. Now this Advanced Editor, I'm going to expand it back out again, is also a syntax checker. So in this case, let's say I deleted this quotations at the end of this rename columns; it's going to one—it's going to give me these red squiggly lines to say that hey, there's something wrong here, and two, it's going to actually give you an error of invalid identifier, and so we would probably know that we probably need to fix this. So we're not going to be breaking down much more of the formulas here, but I do want you to spot two main things from this: the first thing is this column names; column names are always put in quotes in here, and conveniently they're also highlighted in here. So if you needed to do any changes to column names or see what's happening, that's one quick way to identify it. The next thing is this: every step that is taken refers to the previous one. What do I mean by this? So this first step is assign the valuable variable of source, and I know it's assign this variable because it has an equal sign right next to it. And then whenever we go to the next line of expanded data job skills inside this function of table expanded table column, it references source, which if I scroll over it, I can see that it's giving me the same formula for source, which is right above it. So basically, it's plugging right into it. Similarly, this expanded data job skills is going to be located in the next one below it on sorted rows, and it's going to be the first value in here for this table sorted. And if you're curious about what these different functions are doing, you can just scroll over it as well. In this case, table.sort sorts the table using one or more columns of names and comparison criteria, and it tells us via the syntax inside the parentheses that the first parameter is table is table, so it takes that previous variable which is a table. Anyway, one minor last thing about this: if you notice these are surrounded by—these variables have a hashtag and then double quotes on each side, and that's because they have whitespace in the actual names that we're doing for this. In the case of source, there's no whitespace; it's only one value with no whitespace, so it doesn't need to have this around it. Anyway, why am I yammering about all this stuff? If you need to understand this M language, anyway, we're going to actually create this data jobs merge query. I'm going to select it all, press Ctrl+C to copy it. Then from there, I'm going to close out of it. We're going to now create a new query. So underneath the Home tab, I'm going to go to new source, and then under that other source, and I'm just going to go into blank query. Okay, right now this is completely blank, but I can go into that advanced editor of Query 1, and it has the let and instill, and obviously nothing going on here. What I can do is just highlight this all and then using Ctrl+V, paste all of that other query into this. Now when I press done, it goes through and actually creates that same exact query from data jobs merged. Now it could—could have gone through and right-click data jobs merged and click duplicate, but this is more of to show that you can actually go in, copy queries or copy portions of queries, and thus paste it into other ones, which we're going to do in a little bit. So let's get into more of learning about the M Language by actually cleaning up this Query 1 that we just created by using this column from example. First thing though, I do want to rename this Query 1; this is the one we're be working with for the remainder of this lesson, and I'm going to call it data jobs clean because that's what we're going to do; we're going to clean it up. So we have four major tasks that we're going to do with this: the first is for job schedule type; I just want to extract out the first value out of here that's full-time out of it. Additionally, we're going to be using the date and date time columns to extract the weekday and also the hour of the job postings. And then finally, we're going to do some data cleanup on this job title column that frankly is a mess; specifically, we're going to move job postings that have this parentheses remote around it. Anyway, let's start with this first one of this job schedule type. If I go into view and then look at the column profile, it looks like we have that full-time, contractor, part-time, and whatnot, but we have a lot of combines of full-time and part-time, contractor and temp work, full-time part-time and internship. I basically want to go through and just extract out what is the first value that appears in here. So in the case of this full-time and part-time, just want to extract full-time, contractor and temp n work, only contractor. So under add column and then column from example, we'll do from selection, and this appears at the top of add column from examples; enter sample values to create a new column, control enter to apply. So I'll first go by entering full-time, and it's already picking it up. I'm just going to type it in first. Okay, and then I'm going to scroll down, but in this case I'm going to put in, hey, I want full-time for this one; this is the example, remember. So now it's cleaning up that—let's scroll down further if it's done this fully for even more—okay, it's getting the first of these, and you might think that this is correct, but the problem we're running into now is if we go down to this one where it says contractor, it's only contract do—and just looking at the formula, this is the formula it's generated so far; it's doing teex start and nine. I don't really know too much what's going on here, but I'm assuming that it's taking the first nine values; that's not I want. So inside this contractor one, I'm going to type in contractor with an R, so that way it hopefully fixes this. So this is good, and now it has text before delimiter and a space, so I'm going to go ahead and click okay to load this in. So let's scroll down to just inspect it to make sure that we have this correct. And an easier way instead of scrolling down and trying to find something, I can just use this drop-down right here and look in here, and it looks like we're good except for we now have a comma here; specifically, I have a full-time and then a full-time comma. So what's going on here? Well, for values that have more than two—so three—they actually insert a comma in there, and when we inspect our formula, opening up the formula bar here, it's only checking for a space. So the easiest way to fix this is actually just like we did before; we're pretty familiar with it. Let's go to the Transform tab and then under replace values, we want to go to replace values; specifically, we want to find commas; we want to replace it with a blank. Bam. So now pulling down that drop-down, we don't have multiple different full times; we just have that single one without the comma; we have what we want. All right, we're going to rename this, and I can just go ahead and double-click this and rename it, but I'm actually going to do something first. I see that I have the step already for renamed columns, so I'm going to take that and I'm going to drag it to the bottom. Now with rename columns as the last step, I'll then rename it to job schedule type first, press Enter, and then it inserts it into that current step as we can see from here, cuz we're now familiar with it, and we don't have multiple rename columns in there. And then finally, you know how I get about column ordering; this job schedule type first, I want it next to the job schedule type, so I'm going to drag this on over here. See how long it takes, and we've moved it over, and we now have this new step of reordered columns. All right, let's look at some other quick examples for column from examples. For this, we're going to be using the job posted date. For this, using column from example, I'm going to select from selection. Now with some of these things, whenever I type in this box, I want to get, let's say the year. In this case, if I were to type in four one, it would pop up that hey, with all these different options we can do. And so this provides a lot of different options as far as okay, I do know if I wanted to do the month, I could do that. And pressing Enter, it's going to copy it all the way down; that's not what I wanted. This case though, I'm going to double-click it again, go 2023, and scrolling down and looking through this, this option here of year from job post to date. So we're going to go with that, then press Enter. And looking at the transform, we can see what is the M language code that it used for this; it used the date and year function, putting in job posted date. This is what we want; we'll click okay. You know, I—I'm with naming, so we're not going to keep this named year, so I'm going to modify this M language to be job post posted year. With that renamed, let's actually move over to our other example, extracting out the hour. For this, we're going to be using that job posted datetime column. Column from example, from selection. In this case, I want the hour out of it, so I'm just going to put something like nine, and we can see that we also have this here for hours from job post to date time. I want that one, press Enter again. Inspecting the M language formula, it's extracting the hour out of this one; I'm good with it. I'm also seeing the other values are updating correctly; I'll click okay, and we have our new column called hour, which, you know me, we're going to fix this, an updated hour to job posted hour. Press Enter. All right, now we got it. So you're probably like, look, I already know how to go something like the Transform tab and already extract out that information using these functions that we used before. Well, that was mainly as a primer for this next example we're going to be doing, and that's that with this job title column; there's some job
Titles in here that have a lot of sort of frivolous information that we don't need, like in this case, "supervisor information technology specialist," and then parentheses, it has "associate director." I don't need anything in parentheses. Similarly, for this, for the senior data engineer, I don't need this "remote" in here. So let's select this job title, go into "add column," "column from example," and "from selection." For this first one with the associate director, I'm going to select it so it appears below, and then just highlight what I want, press Ctrl+C, and then paste it in here. Then, scrolling here through here to do a cursor check, so I'm seeing that "senior data engineer remote" is in here. I could select it and copy this down here. Another option is I just go in here, double click it since it's now populating, and delete out that "remote," press Enter, and it looks like it's doing this; it's getting the text before the limiter "job title," specifically before the parenthesis, and looks like in this case, "University grad data scientist PhD only now hiring," it removed all that. Okay, so this is now doing what we want. Click "okay," and I don't want—I want to call this column "text for delimiter"; I want to call this "job title clean," pressing Enter.
All right, so last thing I want to now clean up these columns, and you know how I get—I want the "year" and "hour" to be next to the "date time," the "job tile clean" be next to the "job tiles." I could drag and drop these; I'm going to show you something else. This "reordered column" step, we're going to be modifying the M language for this, and I don't want "reordered columns" to appear more than once, so I'm going to take it once again and drag it to the very end. Now what I can do is take and modify this M language that we have in here. Now if we actually inspect this "reordered columns," it may do this or may not; in my case, it didn't add anything after "job skills"; it basically let any new columns just fall towards the end. So this "job skills," all these other columns after it aren't included, which isn't a big deal. So what I want to do is I want to move this "year" and "hour" to near "job posted date" and "job posted month," so I'll enter inside of here, put in "job posted year" and also "job posted hour," make sure we're putting commas after both of those, then I'm going to run this to make sure there's no issues with it, and it looks like it moved it over. Inspecting next to "job post a date," we have our "month" and also "year" and "hour." All right, the last one is this "job title clean," and I want this to be right after "job title," so I'll go ahead and put that in right here, making sure to put a comma after that, and then from there, pressing this check mark up here to move it. Inspecting over, we have "job title clean" right next to it.
Our next to look at is "custom column." We'll go ahead and actually just select this, and whenever we pull this up, this tells us—this allows us to add a column that's computed from the other column. It provides a box to basically put in the new column name, but right here, this is where we put in the custom column formula or the M language to maybe clean it up. Now let's start with something simple. Let's say I just wanted to repeat the "job ID" column. I would come over here, select "job ID," click "insert"; it's going to put it in. Notice that the variable itself is inside of brackets, and I'm going to rename this "job ID repeat." Down at the bottom, it's telling me that no syntax errors have been detected. I'll click "okay," and then I get this new step for "added custom," and we can see, hey, it's "job ID repeat." Scrolling over, yep, it repeated it. If I want to go back in to edit it, I'll press that settings icon, and it's going to pull this back up. So let's do something a little bit more complex now, and it's going to involve the "salary year average" column and that "salary hour adjusted" column. Go ahead and cancel out of this. What I want is to create a new column that if there's a "salary year average" value, it will basically be in that new column, and then if there's a "salary hour adjusted" value, it will be in that column instead. Just as a warning, anytime "salary year average" is null, there's always a value for "salary hour adjusted," and vice versa. So like I said, we're not going to become coding experts with this, so I recommend taking use of chatbots like ChatGPT, Gemini, or whatnot; lots of free options available out there. Anyway, we have this prompt of "generate a power query formula for a custom column on building; make the column salary year average if it's not blank; otherwise, it is salary hour adjusted." Now it's giving—do the entire M language. Right, this is what we're providing to the advanced editor, providing that previous step name, what column we're using, everything like that. I care about really this formula right here specifically, everything after the "each." I'm going to copy this from "if" all the way to the end; that's the actual code right here. Going back to the "custom column," I'm going to delete that "job ID" out of there. I want to make sure that there's an equal sign still there, and I'm going to paste this in, and you can see from this—this is just basically an "if" formula; it's doing "if salary year average is not equal to null, then salary year average; else, perform salary hour adjusted." Down at the bottom, we can see that no syntax errors have been detected, so I'm going to go ahead and click "okay." So bam, we now have this. I did—in that "jav ID repeat" value here, so we're going to actually change that to rename that value to "salary year combined," and then clicking the check mark in order to rerun that formula to update the column, and you know, I like to have my steps in order, so I'm going to grab "reordered column," and I'm going to drag it to the very end, and for this one, I'm just going to drag it over to "salary hour adjusted," right after it, to "salary year combined." So now scrolling down just to double check it, it looks like we got 140,000 here, 140,000, 82,000, 82,000 there, so the formula filled out correctly. So let's get into our final task.
So we've been working this "data jobs clean" data set. We made this "salary year combined," which is pretty useful. Actually, what happens now if we want it in something like "data jobs merged"? What do we need to do to actually add it into here because we have everything we need for it, specifically we have that "salary year average," and we have the "salary hour adjusted" columns? Well, we could recreate it in here, going through all those steps, creating that "if" statement, or we could just copy it out of the advanced editor and bring it in here. So I'm going to go back to "data jobs cleaned," and then under "home," "advanced editor," I'm going to go and find the step that's in here, specifically it was this "added column," and I'm going to copy it because I can see that, hey, it has the "salary year combined" in it. I'm going to copy it all the way to the end, and I'm going to copy it by pressing Ctrl+C. Okay, go and close out of this one, and then bring over to "data jobs merged," go into the advanced editor, and I want to insert it in right at the end, so I'm going to go to at the end of this block of this "let" block, going to press Enter, and then from there, press Ctrl+V to paste it in. Now I'm already getting an error message, and it's saying, hey, "token comma un"—basically expected, and it's not getting it. If I scroll over, I can see these squiggly lines right here; basically, there's not—if we can see, there's commas after every one of these variable definitions, so I need to come up here, put a comma in there. Next is this—a comma cannot proceed an "in," so if we scroll over, we can see this is red highlighted, probably wrong not to have a comma here, so we'll get rid of it. Now we're not done; it's going to say there's no syntax errors, but we didn't complete this. Remember, you have to have the name of the—it's got to reference the previous name here in it. So if I tried to—even though it says no syntax errors—if I try to click "done" and go to load it, I'm basically getting an error. I can see this by this basically air Bo at the top of each one of these columns. Also, there's only one applied step, and it's calling it "data job merged" of the actual title itself, but we need to fix this query and actually get it back to where it had multiple different applied steps. So I'm going to go back to the advanced editor; we're going to show what we did wrong here, and that has to deal with—remember we had before where we had something like "remove columns"—you reference the previous column in it. So in this case, "remove columns" right there. Well, "rename columns" is the last one we had. I'm going to go ahead and copy this by Ctrl+C-ing it, but yet we have inex inserted text before "delimiter one," which is not correct, so I'm going to select all of that and replace it by pressing Ctrl+V. So we have the "rename columns" now. One other thing we have to do—this last statement or—and the "in" portion needs to be referencing that last variable of "added custom," so I'm going to go ahead and copy this, Ctrl+C, and then pasting it in, Ctrl+V. Click "done," and now scrolling all the way over, we can see that we have that "salary year combined" column that we created in the last query; it's at the end; we do need to move it over, but it's in there nonetheless. So it helps with understanding these queries.
Now one quick thing before we go—we've gone through basically every single thing in this chapter on Power Query up to this point with the exception of this "invoke custom functions." This basically invokes a custom function defined in the file for each row of this table. This is more advanced and beyond the scope of this course; we're not going to be covering it, but is available for you to dive into. Say you're doing a lot of different imports, and you need to automate the imports that you do, this would be a path you would go, but for beginners like us, I'm going to say stick away from it for the time being. So this now wraps up on the M language, and that was really a crash course and understanding how to use it. By no means do you need to be a professional or be an expert coder and coding the M language. If you got lost at any point in the way, nothing to feel ashamed about; this is a very pretty complex topic. If you would like to learn more, I do recommend this book, which is "M is for data monkey"; it's a good little read, talking about not only Power Query but also how to manipulate the M language. I'll include a link in the description below. Anyway, Power Query, in my opinion, is one of the most important features, the most powerful tools within Excel and also Power BI, and so it's worth your time investing and learning it. And so this all culminates, and we're now finalized covering Power Query in this chapter. In the next chapter, we're going to be jumping into Power Pivot, and that's going to—jumping into actually data modeling. But before that, for those that purchased the practice problems, you have some practice problems to go through and get more familiar with that M language for proceeding forward. All right, with that, see you in the next one.
Welcome to this chapter on Power Pivot. This chapter consists of four different lessons where we're going to go—an intro into Power Pivot and over the window that it actually provides. Then, from there, looking into DAX, or Data Analytical Expressions, which is a formula language very similar to Excel formulas. But before we actually jump into this lesson and going over what we're going for, we're going to focus on what exactly is Power Pivot. So here I am in Excel, and this is meant for me to just go through and quickly explain what is the power—the power of Power Pivot. I know that pun is getting sort of old by now, but it really is powerful. If you're curious of looking at it, it's in the workbook of "Power Pivot Intro Part One"; "Part Two" is what we're going to be using for the actual lesson. So in Power Query, in the last chapter, we end up clearing up our data set to have these two main tables versus "data job salary," which has the complete data set on all the data science job postings, and then "data job skills," which is unique to the skills for a job. We also created a "data jobs merge" table, but that table is actually going to be—well, it's pretty much obsolete, and Power Pivot is going to help replace that, and for good reason. So what exactly is Power Pivot? Well, it's an add-in; we're going to get to adding it in, and it has a few different features that you can do within it, such as accessing the data model, adding measures, KPIs, and whatnot. This lesson is going to be going over this tab. As a quick refresher, Power Pivot is going to be available in basically any version of Windows for Microsoft past 2010, but it's completely not available in either the Mac version or the Microsoft online version, so you won't be able to do this chapter if you have those versions or the final project. Anyway, the core portion of Power Pivot is actually managing a data model, and what's a data model? Well, a data model defines how data is basically structured, stored, and also related. In this case, we have the "data jobs salary" table right here, and we have the "data jobs skill" table. What we can do with Power Pivot besides modeling these tables and showing how they're structured is the more important thing of creating a relationship. In this case, I created a relationship between the "job ID" of "data job salary" and that of "data job skills," and because I created this relationship, I can look at things like the "job title short" column, see how many jobs it has with it, but also I can query across a table over to the "job skills" and see how many skills has with it. In fact, let's actually do that real quick. Here I have my data model itself; I have my two tables which are shown. Anyway, I can look at things like what are the count of the different job titles themselves. I'm going to do that on "job ID," and like we've done plenty of times before, here's the job count with a little clean up of the actual text here. But now with Power Pivot, I can actually reach across to that other table of "data job skills" and drag the "job skills" into here, and this is telling us obviously the count of the skills based on the job title. Pretty cool that we can reach across the tables and do this. Now the other cool thing that Power Pivot unlocks is DAX, or Data Analytical Expressions. Recall previously that we were using the average of the salaries, and like we learned way back earlier in this Excel course, we prefer actually a median salary, but unfortunately, looking at the "value fied settings" window here, there is no option to actually pick median from this, and that's where—where DAX comes to the rescue. With this, I can go to something like the Power Pivot tab and now create a measure, which is where you actually insert in your DAX, and I can create a new one called "median salary," and we're going to be using this DAX formula. In this case, I'm going to use the median formula, very similar to the Excel formula, and I can do it on the entire "salary year average" column here. I'm going to format it real quick and then press Enter. Anyway, bam! Now we have—because of the power of DAX—we have the ability to get the median salary, and those DAX things can do some pretty complicated calculations. So in the case of here, we have this "job count" and "count of skills," and we want to see what were the skills per job, specifically. In this case, what is something like C2 / B2, and then dragging all the way down and filling it for all these—this provides a much better analysis of what's going on with these values of counts and skills here when we get this proportionality. We can create this with measures as shown in this final pivot table that we're going to be creating coming up in the third lesson of this chapter. So in summary, Power Pivot provides us the opportunity to now model our data, which allows us to one, create relationships, and two, allows us on—unlocks these measures that we can create using DAX. All right, so let's get into this lesson. What we're going to be focused on for—well, first thing is we're going to enable the Power—Power Pivot plugin, and then from there, actually getting into data modeling or modeling our data that we imported through Power Query. After we have everything set up with our data model, we're going to then move into performing our first analysis, analyzing based on a job title how many different skills they have associated with it. Like I said, we'll eventually get to that "skills per job" in an upcoming lesson. So for this, you can continue to work in that workbook that we were working with in the last chapter, "EMP power query." We're going to continue work on that because we want to use those queries that we built. If you got lost during the way and just want to start back up, we're going to be starting from that M language workbook back in the Power Query chapter. As a reminder, these lessons or workbooks are what are the completed workbooks at the end of the lesson, specifically for this lesson; "Part One" was just that intro; "Part Two" is what will be done at the end of this lesson. Anyway, here I am in the M language workbook. We need to get into enabling Power Pivot. Right now, you probably don't see Power Pivot up at the top of the tabs, so I'm going to go into "file," and then go down to "options." From here, I'm going to select "add-ins," like we did before, and instead of "Excel add-ins," we're actually going to be using those "COM add-ins." I'm going to click "go," and they have three different ones available: "Data Streamer," "Power Map," and "Power Pivot." We want "Power Pivot." I'm going to go ahead and click "okay." Now "Power Pivot" should appear up at the top, all the way on the right-hand side, and should look something like this. Quick little overview of this tab—"manage" here pops up the Power Pivot window, which we're going to be doing a deep dive on this in the next lesson; we're going to use it a little bit in this lesson, but anyway, that's one way you can actually access it. You can also go to the "data" tab, and then here under "data tools," you should see it also, and you'll be able to manage your data model, and once again, it will pop up the window. Additionally, on this tab, you have the ability to create measures and KPIs, which we're going to be diving deep into in the third and fourth lesson. If you have a table within your worksheets, you can add it to your data model; you can also go about detecting relationships, although I don't find that this feature works that well, and then finally, they have "settings," and "settings"—I don't really touch that much, nor does it have much control here. So let's actually get into—embarking some data into our data model. We're going to do a simple example first. Here I created a new sheet, made three columns of "ID," "name," "salary," and then different values associated with it. One way I can add to the data model is if I have data in a table is to do this feature of "add to data model." In this—my table has headers; I'll go ahead and continue, and then it will pop open Power Pivot; a similar like environment will exist with Excel. I can't actually edit any numbers in here; this is just how you're modeling your data. If you needed to actually edit it, I have to go back to the sheets, and like I said, this isn't a method I typically use; typically have bigger data sets not located in tables. So I'm going to go ahead and right-click this down at the bottom, this table name of "table two," click "delete." It's going to say, hey, "do you sure you want to delete this table," and bam, it's gone. All right, so now there's nothing in our data model right now. Here we are still inside the Power—
Pivot window and if you've noticed from this in the Home tab right here, it has the option to get external data. They have options for you to actually connect directly with Power Pivot to things like a SQL Server, Microsoft Access. You could also get it from some sort of data feed, and then this option would be more probably useful in that it has a lot of different sources you could use, such as other Excel files, text files such as CSVs and whatnot.
Now you may be asking yourself—I'm going to close out of this Power Pivot—why would I import that whenever we just went through with Power Query to get data via this? When which time should I use which? Well, it's very important to remember the purpose of the tool that you're using. Power Query is an ETL tool—extract, transform, and load. We did a lot of Transformations with our data set, and so that's really the power of Power Query. And then it loads it in. Power Pivot's strength is not in ETL or data cleaning; instead, it's in data modeling, creating these relationships and DAX.
Now, now you may be tempted to come inside of existing connections and try to connect to specifically that salary and skills, and if we went through like in the salary case and try to click open, we're going to get an error message. And I'll be honest, this is really confusing because we have this workbook connections; why isn't this working? Well, it really just comes down to naming conventions and the fact that Power Query connections are not the same as Power Pivot connections, but we have a fix for this. We just need to exit out of the Power Pivot window here inside of Queries and Connections. Remember, you can get to that by going to the Data tab and going to Queries and Connections. We can go to something like Data Job Salary, which right now is a connection only, right-click it, and go to Load To. Right now, it's only under Only Create Connection, but we need to check this check mark of Add this data to the data model. I'm going to click OK. It's going to go through this process of loading the data, and now it talks about the rows are loaded, but mainly if I go to the connection, it has this new connection now of this workbook data model, which if I go to and actually open up or manage our data model, we can see that it's inside of here. We have this basically sheet for the table itself of Data Job Salary inside Power Pivot inside the data model.
Now we do need to get that other pivot table or other table into there as well, so I'm going to go to Queries Data Job Skills, right-click this, Load To, and also add this to the data model. Okay, it talks about 167,000 rows are loaded, and another connection—still, it's only going to be one connection because we only have one data model in this case—and now when I go to manage the data model, I have two basically sheets down here, but two tables, and now we have the Data Job Skills in here. Anyway, I want to do some cleanup real quick. I'm going to clean up Power Pivot, but this Data Jobs Merged and this Data Jobs Cleaned—it's going to be very confusing, like I said, we're not using this mainly for the fact that we have duplicate values in here for Senior Data Scientists in this case and then for the salaries, and so if we don't manipulate this in a correct manner, we're going to get the wrong results, so we're just going to get rid of these. So for Data Jobs Merge, I'm going to right-click and select Delete, and it's going to say, "Hey, should you want to delete Data Jobs Merge?" Yes, I do. And then I'm going to do the same thing with Data Jobs Clean, right-click it and select Delete. Also, if you have these tabs down here for Data Jobs Clean or Merge, you can go ahead and delete those as well.
With our models now cleaned up, let's actually get into going over really briefly this Power Pivot window. With this, we have three main tabs: Home, Design, and Advanced. Advanced, we're not going to go into a lot of things inside of this, if any at all; it's beyond the scope of the course. We're going to be focusing mostly on the Home and the Design tab. So with this tab, we've already gone over Get External Data, but we can do things like refresh our data if we know that it's updated in Power Query, generate Pivot Tables and Pivot Charts based on our data model itself, change the formatting of a particular column—in this case, it's noticing as text. If we go to the Data Jobs Salary data, we can actually scroll over and see that for the Salary Your Average column, it knows that it's a currency. We did a lot of this cleanup right in Power Query and setting these different data types, so this saves a lot of steps here in Power Pivot if it wasn't done. Now we have options displaying the table below that we can actually sort it, we can filter it, or sort by a certain column. They also provide options to find a specific value within here, and then these features for calculations—I don't find myself using that much as far as the auto—so anyway, over on the right, the most important thing I find is allows you to toggle on the different views of your data set. So right now, this is the Data View, and if I scroll over here, this is the Diagram View, and this is going to show our two different tables side by side. I'm going to move them over and actually expand this one out to show all the different columns and then the Data Job Skills.
Now back on that Data View, clicking that, we have Data View, but also below this we have this calculation area, which I can toggle on and off. Calculation areas are where we're going to be storing our different measures that we build with DAX, and so they'll be appearing underneath here. Here, if we have any hidden columns, we'll be able to toggle them on and off. Right now, I don't have any hidden columns. Now, one thing to note with this data cleanup—some of that we did before with formatting stuff—some of it's going to be quite limiting. You may not be able to do like in the case of this—so Data Job Skills has this Job Title Short column, and actually if we look at the Data Jobs Salary data set, we have the same repeated column in it. So Data Job Skills, this Job Title Short right here is unnecessary. Now I could right-click it and try to delete the column and ask me if I want to delete it; it's going to tell me it's not going to be able to do it because it was created by a query, i.e., through Power Query, and instead I should actually update it through Power Query, which I would actually argue as best practice anyway. So I could exit out of Power Pivot, launch Power Query by pressing Alt+F12, then go into the Data Jobs Skills query, and if I want, I can just select this column and select Remove Columns, but you know how I am—I like to actually clean up the applied steps because it could, depending on how large your Power Query query is, it could take a long time to load it and unload it necessarily. So if I go to this Remove Other Columns—that's the first time that it appears in it—I can remove this by deleting it out of there, then pressing Enter. We may get an error message, we may not; I'm not sure. Going to the last step in here, I notice there one thing of the table wasn't found specifically here; it's appearing Job Title Short in here, so I can go ahead and delete Job Title Short along with with that comma, and bam, we now have this final table. Just to lean for those two steps, I'm going to go ahead and close and load this, and now going back in to look at our data model and Power Pivot, I can see that it updated for Data Job Skills.
All right, moving into this Design tab within Power Pivot, this has a few different options within it for adding columns, freezing columns, just messing with the columns. They also have different options for creating calculations concerning columns. We'll be getting into calculating columns more in the next lesson, so stay tuned for that. Right, the main thing that we're actually going to be doing in this portion of the video is actually setting up relationships, and that is we could go about creating a relationship here, and right now I have Data Job Skills, and I could relate it with the Job ID by pulling the drop down to the Data Jobs Salary table on that Job ID. Now that's a way I can do it; I'm actually not going to do it this way. I actually prefer going to the Diagram View, and then from there just dragging and dropping the Job IDs across each other, and then it establishes this connection, which we can see through this line through here. Now there's a few different things that we need to notice from this line here: one, this arrow—it's going to come to bite us in the butt later—and that's that that arrow only allows data flow in one direction, and by data flow, I mean filtering. If I try to filter something in the Data Job Skills table, this arrow is only pointing in one direction; I won't be able to filter it back. We'll encounter those problems in a little bit, and we'll talk about strategies how to actually offset it. The other thing to note with this relationship here is you notice right here it says One, and over here it says Star. In this case, this is a one-to-many relationship, and what does this mean? Well, going to our Data View for Data Job Salary, we only have one unique ID for each job, whereas in the Data Jobs Skills, we have multiple different Job IDs, or many Job IDs. Now if we only had one Job ID in there and we actually looked at Diagram View for this relationship, we'd have a one-to-one relationship, but we have multiple skills in there, so that's not possible. Now it's also possible to have a basically as-is or many-to-many relationship, but that causes a mess, slows down your data model, and I don't recommend it, so you should typically see either a one-to-one or a one-to-many.
Last little wrap-up before we actually analyze and use this relationship: we have the options for Table Properties, which we're not going to be able to look at because this was created via Power Query for this connection, and then we have options to create date tables underneath Calendars, which we're going to be exploring in an upcoming lesson, and like always, you have an Undo and Redo. Anyway, let's actually get into analyzing and putting this actual relationship to the test. So what we're going to do is inside the Home tab, go to PivotTable CU; we're going to want to create a PivotTable with this. We're going to insert a PivotTable, and we'll have it insert into a new worksheet. Selecting inside the PivotTable, it's not having the field list come up, so I'll select it under PivotTable Analyze. Anyway, we want to query across this table to show the power of the relationships. So what I'm going to do is from the Data Jobs Salary table, I'm going to take that Job Title Short, throw it into the Rows, and then from there going to come down to the Data Jobs Skills table, and I'm going to throw the Job Skills into the Values. It should be performing a count, and then I'm going to organize this real quick from largest to smallest, and it looks like Data Engineers have the most. So this is pretty neat; we're able now to query across tables. Going back into that Power Pivot window, this connection allows us to do that. I'm going to just show you something real quick by clicking this relationship, right-clicking it, and deleting it—want to delete for model—and I want to show you how these values are basically going to change inside our PivotTable, basically to the fact that they're going to have it to where they're all the same value, and that's how you know that your relationship is not set up correctly whenever you have multiple repeating values and you expect them not to be. Anyway, sometimes you'll see this popup come up of Relationships between tables may be needed; Autodetect, and sometimes it works, sometimes it doesn't. Um, in this case, it looked like it worked, so we're going to go with it, and just double-checking it in Power Pivot, it is set up correctly.
So for this final analysis, we're going to be looking at building this visualization right here, analyzing what are the top skills of data nerds. We're basically remaking what we did in the Power Query chapter. Now that we have that updated data model, anyway, we're going to build this out to see where the skills counts for each of these and also provide filters for Job Country. So back inside the workbook that we're previously working with—if I would actually remember—we did make that sort of similar visualization that I talked about, but however, if I go to Data and actually refresh the data, it's going to give me this error message because once again we deleted Data Jobs Merged. Anyway, I thought this was actually going to go away; it didn't. It is not what we want. We're going to delete this one, and then we're going to do a little bit of cleanup. So that one that we created the job analysis on, I'm going to actually just rename that quick to Job Analysis, and then now in this new sheet, we're going to do—we're going to name this one Skill Job Analysis. Anyway, let's insert a PivotTable in here, so we go to Insert PivotTable, and now what we have the option for is from Data Model, and it's asking if I want to put it in the existing worksheet. Yes, I do. Remember, we want to analyze the skills and specifically how many counts they have associated with it, or how many jobs they have associated with it. So I'm going to put the Skills into into the Rows, and then from there I want to count how many jobs are associated with it, so I'm just going to drag that Job ID into the Values. Right now it's doing a sum; going to click on it, go to Value Field Settings, change this to Count. Now you may be like, "Luke, could we use the Job Skills Count?" And we can, which has the same exact values, but actually closing this out and taking out Job Skills, you're probably more interested in why can't I use something like the Job ID from the Data Job Salary table? Well, if I drag that over, and then I change this Value Field Setting to a count, Count, and click OK, you notice it says 32,672, which is coincidentally the same number of rows of that data set, and this gets into the point of filter direction. What do I mean by that? Let's go back to the data model itself, looking at it in Diagram View. Remember the arrow is pointed towards the Data Job Skill table. Right now I have Job Skills in the Rows, and I'm trying to filter for Data Job Salary based on the count of the Job IDs, but the arrow doesn't flow in that direction; we can't do it. Now in something like Power BI, you can actually right-click this, edit the relationship, and change the direction; that's not possible within Excel, unfortunately. Anyway, we're going to be using DAX to fix this in the future. For the time being, we're just going to go about using, in this case, for this analysis, the same values in the same table. I'm going to remove this other Job ID from the other table. Anyway, we're going to sort these values from largest to smallest. Then additionally, I only want to show the top 10 skills, so I'll go to Value Filters and then Top One… Top 10 items by Count of Job ID is what I want, and so now we have this. So now we have the values we want to visualize. I'll go in and actually insert a Pivot Chart for this. I like the bar because it makes it easier to read the different skills that it has right there, and I'm realizing now the sort order is actually backwards in this; I want it from smallest to largest. I'm also going to right-click and hide all field buttons. We're also going to be adding axis titles for the primary horizontal and then removing that Legend. We'll update this title to What are the top skills of data nerds?, and then the y-axis is self-explanatory, but for the x-axis, we'll label this Skill Count in Job Postings. Okay, the last thing we need to do now is actually add some slicers to this so we can actually control it better. So selecting the table itself, going to Insert Slicers, I'm going to select the Job Title Short, and also we want Job Country right here. Which each of these slicers, I'm going to rename them also. This one Job Title Short, I'm going to rename to Job Title, and then Job Country, I'm going to rename to Country. Now when I go through, I can actually select something like Data Analyst, and it will filter down and actually see the associated skills. I could also do something like look at those in the United States specifically for their counts, and we see that SQL, Excel, and Tableau are the three top skills.
Now you may be scratching your head on like, "Okay, I thought we were trying earlier to actually aggregate something in the PivotTable, and it didn't work." Well, remember this arrow is pointing to the filter direction. So in our case, we have a Job Title Short slicer, because this arrow is in the direction back to the Data Job Skills table, we can filter in that direction, but we cannot conversely filter in the other direction; that's why we can't get the counts from these tables. Little confusing, I know, but I promise you we will work out as we go through this entire chapter in Power Pivot. So bam, we just completed our first analysis for our final project. We have a few more analysis coming up in the next lessons. You do have some practice problems though to go through and get yourself more familiar with Power Pivot and understanding what's going on with these relationships, the one-to-many and whatnot. All right, with that, I'll see you in the next one, which we're going to do a deeper dive on looking into that Power Pivot window. That I'll see you there. All right, let's now dive further into Power Pivot, and we're going to be focusing on the Power Pivot window for this. We're going to be looking at some major aspects of it. For this, we're going to get into using a little bit of DAX to create our first measure, and with those measures, we're also going to be exploring the difference between implicit and explicit measures. Don't worry, we'll cover that in a bit. From there, we're going to move into a feature that's related to measures called calculated columns, and it's going to allow us to, inside of our data model, create different values—such in this case, we can actually create a date column from our date time value. The last thing we'll explore are date tables, which Power Pivot gives with a click of a button and allows us to connect these data tables of these date tables to our original data source and then filter it by a lot of different data. And so we'll wrap this all up with a final analysis where we're looking at job postings based on a day of week using this date table.
Anyway, jumping into Excel for this, we're not going to be using any of the work that we've done previously; instead, we're going to open up a completely new workbook and be working out of this instead. And the reason is all the work that we're going to be doing within this lesson, we're not going to be carrying it on to our project that we're using. This is more—this lesson is more to get us more familiar with the powers of Power Pivot. Oh gosh, this pun's killing me. And so we'll eventually incorporate some of the stuff into our final project, but like I said, we're going to be starting with a blank notebook or workbook for this. As always, if you want to see what the results are at the end of this lesson, you can just go to Power Pivot window, and it will have it all right. So let's actually get some data into here to start working with, and like I said, we're not going to use Power Query at all for this; we're going to use Power Pivot. So I'm going to open up the goto to the manage the Power Pivot data model, and we want to get this external data—specifically, we want to get that Excel workbook that we've been working with of Data Jobs Salary—all so underneath the Home tab, I'm going to go to Get External Data.
And it's going to be from other sources. We scroll all down; we could look at how we can import it from different databases or whatnot. We're going to be doing it from an Excel file. Then from there, we're going to browse the connections, navigating into that data set folder. I'm going to select "Data Jobs Salary." All it PR me if I want to use the first row as column headers. I do. If I wanted to, I could go in and test the connection to make sure it's it's going to succeed, and it does. So we'll go from there to next. It sees that it has one sheet within the workbook; that's the one that I want. I'll click "Finish." Next, it'll go through the import. Looks like it completed; it has a success; got 32,000 rows. I'll click "Close."
Now let's go through and actually clean this data set up using Power Pivot. Now I know in the last lesson I talked about, hey, we're using Power Query for ETL, and that's true. But let's say you have a quick data set you need to connect to and model quickly; in that case, you would do some of the stuff that I'm going to do here in order to quickly model it. If I wanted to rename it, I'd come down to this basically sheet tab down here; it's called "Sheet 1" after where it's at. I'll rename it, and we'll keep a similar naming convention of "Jatt Jobs Salary." Go ahead and click "Enter."
So let's say for this quick analysis that we're trying to do in this lesson, I'm trying to analyze only the yearly salary data. I don't care care about the salary uh hourly data, and I don't even want the data entries in here. Well, I can get rid of that "Salary Hour Average" row by just deleting this column by right-clicking it. It's asking me if I want to delete it; yes. And now there's still blank values in here, right? So I need to get rid of this salary rate values that are equal to "hour." So I'm going to click the filter here, unclick next to "hour," and click "OK." So now we have that out. The other thing I want to do is actually clean up the format of the salary, and I'm going to change that instead to a currency. This talks about how the data is going to be a changed when where it's stored. Yeah, I don't really care about that. No doubt that I care about will be lost; it'll all be here still. And then I'm going to reduce the decimal places by two. The other thing I can do if I wanted to is actually sort this based on that job posted date. Could come up here and sort from newest to oldest, and then it's in order. Sorry, actually want it oldest to newest; got confused on that one. So bam, just did some quick clean up to our data set, and now we're ready to proceed forward.
So let's actually get into building our first measure or measures specifically. I want to analyze this to understand what are the different the the amount of jobs in here and then also what is the average, and then also more importantly, the median salary. Well, there's a few different ways we can do this. We're going to do this first within this Power Pivot window. So in order to do this, I'm going to first first I want to do a count, so we're going to just run this on this "Job Title Short" column. And here underneath on the Home tab under calculations, we have this Auto Sum. I don't frequently use this; I use it every now and then, but I can run things on this like count or distinct count. I'm going to do "count" in this case, and this is going to create our first measure down here. Remember, down below this area is our calculation area. I can toggle it on and off by clicking "Calculation Area" up here. Anyway, I can also make this column slightly bigger, and what's cool about this, keep on scrolling over, is now it tells us the name of this measure, "Count of Job Tile Short," and that there's 22,000. Remember, there's normally around 30,000, but because we've taken out that hourly data, we're down to 22,000. Now I can also edit this measure. If you notice, it appears right up here. Similarly, they have a formula bar in Power Pivot, and to the left-hand side, it tells you what is actually selected, "Job Title Short" column, and then the actual measure itself in here. Now one quick note, there is basically a colon and then an equal sign; that's how we're going to know that we're doing measures, and we'll get to calculated columns in a little bit, and it will only be the equal sign. But this is Microsoft's way of signifying that this we're using a measure so that way you don't confuse with anything else. Anyway, I can edit this the actual title in this case, and I can change this something to more more descriptive, to "Job Count," pressing "Enter." It now runs it, and it's a lot shorter. Additionally, if I want to actually format it, I can have the measure selected, come up here, select comma, and then it formats it with the comma, and then I don't want two decimal places; I'll go ahead and remove it.
Next, let's get into analyzing that salary column with this. Once again, we can click this; I could use that Auto Sum and do something like "average" here. Clicking "average," and below it, it generates that average of salary or "Average 123,000," and I can change it if I want to "Average Salary." But if I wanted to calculate something like the median instead, I would have to actually manually type out this calculation. So selecting right below "Average Salary," and then coming into the formula bar, I can type in something like "Median Salary." Remember, we want to create a measure, so it's going to be a colon and then an equal to, and then for this, we want to use the median function. Now a lot of these functions that are DAX functions are very similar to what we use in Excel, so they have a lot of different similarities. But with this, like we talked about before, this allows us to now put in basically an entire column into it and then perform that entire aggregation on it. In this case, I want to do it all on "Salary Year Average," making sure I put a close parenthesis to close out that function. And bam, right next to "Average Salary," we have this "Median Salary" now, which needs to be formatted, so I'll format it as "English (United States)," stes USD, and remove the decimal places.
Now what happens if I didn't enter that colon equal sign? So here I am selected below "Median Salary," we'll go ahead and paste in that formula, and we'll delete that colon. I haven't run this yet. Now I'm going to run; I'm going to press "Enter," and as you notice by this, it's not actually calculating a value; it actually just converts this to text. So this is not what we want; that's why we have to do the colon equal to sign for entering in the formula bar there. So with measures, it's important to understand implicit versus explicit measures. So let's close out the Power Pivot window and actually get into exploring these different measures by creating a pivot table of that "Median Salary" we just created. So I'm going to go to "Insert," "PivotTable" from "Data Model." We're going to insert it into the existing sheet here. We have our table of data, "Job Salary." I'm going to analyze the salary based on the "Job Title Short" column, so I'll put "Job Title Short" into the rows, and then look, we scroll down at the very bottom; you'll notice that the measures that we created have this F of x; basically, it shows us an equation that it is a measure. So I can take these measures, this "Median Salary," and in this case, drag it into the values, and now unlike Power Pivot where it did in that same column, we're now filtering down to do it by well the appropriate job titles. Now we could also do something like drag the the "Job Count" into the values as well and actually see the "Job Count" there. Now both of these measures are explicit measures because we explicitly defined it; we despine defined what "Job Count" is and what "Median Salary" is. So what is an implicit measure? Well, you actually created this before. So in regards to that "Job Count," we're doing a count of the "Job Title Short" column. If I were to drag that down into here, you can see it says say "Count of Job Title Short." This is an implicit measure. These are great for quick short analysis, as we demonstrated before; you can quickly throw something in and generate it, and you didn't even know your us the measures, and you were similarly with the "Salary Year Average." If I drag that in down here, we previously well changing this up to actually perform an average mov to average from there; that was also an implicit measure. So I think you get the point, but we're going to see the power of this as we go through this when we start to make newer measures that are actually going to use our explicit measures specifically. We're going to be using our "Job Count" in other calculations, and so these explicit measures are going to save our butt and save us so much time and ensure we're doing the correct calculations.
So let's get into our first calculated column, and we're going to be going back into the Power Pivot window for this. We're going to be creating a column that will convert the "Salary Year Average" values into Euro values. So there's a couple ways we can do this or add these columns. We can go under "Design," and right here under "Columns," we can click "Add" to add a column. Additionally, without that that unselected selecting back in into again, you see this "Add Column" up here; we can just go right in and add a column. I feel that's actually easier. Anyway, in this case, in order to get the Euros value of what it is for "Sal Year Average," we need to multiply by a conversion rate. So inside of here, I'm going to put the equal sign, and we see it's popping up here in the formula bar. From there, I'm going to use the value in "Salary Year Average"; I just selected one of the values, and it popped right in. Then from there, similar to how we wrote formulas before, I'm going to put times 0.9. "Enter." Now notice from this one, I didn't use the colon equal sign, right, because is not a measure; it's a calculated column, and it still knew that this was a currency, although I don't like it; it has two decimal places, so I'll remove it. And to me, it knows it's a currency, but it doesn't know that it's a Euro, so I'm actually going to convert it over to Euro and then remove the two decimal places. Additionally, I'm going to rename this from "Calculated Column 1" to "Salary Year Euro." You can identify calculated columns because normal columns are green; the calculated columns are black. Also, if I go to the DI diagram view, we can see that well you can't really tell that we have the calculate column C Euro, but you can see your different measures that you've created. All right, so back to the data view, even though we have this calculated column, we could also create a measure on this calculated column. Clicking in the box below here and then typing in here, I could do something like "Median Salary Euro," and then put in that median function for "Salary Year Euro," and then close the parenthesis. And bam, now we have it. I'm going to spread it out to actually see we have the value of €103. Now going back into here, we can take this and actually if we wanted to, we could put the "Salary Year Euro" into there that column; it's going to aggregate it appropriately. Right now it's doing a sum, so if I wanted to, I could get an average of this of these values, or we could actually go to that measure that we created, that explicit measure, throw it in here, and we get the explicit value of the "Median Salary Euro."
All right, so let's shift our focus on this analysis. Let's say we wanted to analyze more around the date, specifically the day of the weeks for when job postings are occurring. Well, let's go back into Power Pivot and manage to open up the Power Pivot window right now. Investigating our the diagram view of our data model, we only have one table in here, "Data Jobs Salary." Well, if we go under the "Design" tab, talked about in the last lesson, we can actually create a date table. I could also potentially mark this table of do job salaries a; it's not a date table, so we actually need to create one, and you'll see what it looks like after that. And with that, I did click "New" on this anyway. It created this new table called "Calendar" and expecting all of the different values in here. Well, let's actually just get out of this view; let's actually go to the data view one, which is pretty cool with it with it what it created; it created it based on the dates it knew what was in our original table. So from the first of 2023 all the way to the last day of 2023, and with this, it has a year column, month, day of week, and day of week number; so a lot of great values from it. Now we need to actually connect these two; there's no relationship between the two. If we go to that "Data Jobs Salary," so selecting it here here, we only have this "Job Posted Date" column, which is a date and a time, so we need only a date. Because this column is named inappropriately, I'm going to change it to "J Job Posted Date Time." So now let's create that new column with that "Job Posted Date Time." This time though, instead of clicking "Add Column," we're going to go to "Insert Function," and this is pretty neat because it allows us to actually look under different things. In this case, we wanted sort of a text function, and we can look and explore different ones. Specifically, I know we want this one, a "Format" converts a value and text to the specified number format. So I'm going to click "OK," and it automatically fills it in with this colon and equal sign of "Format" equal to. From there, I'll select the "Job Posted Date Time" column; that's the value, and then what do we want for the format? Well, I know we want in the format of basically the year first, then two months or two M's, and then two D's for month and date in order to match. Close that double quote because that's the actual format we're using; that's all we need, so we'll close the parentheses and press "Enter," and then I'm going to take this "Calculated Column 1," drag it over here, and then I can see that it did convert it correctly. So I'm also going to go now and rename this appropriately to "Job Posted Date." Press "Enter." So now let's create a relationship between the two. Remember, we can go to that diagram view, or I can use this of "Create Relationship," go to "Calendar" to match on the date itself. Let's see what it looks like in that actual diagram view. We always want to inspect it to make sure we have this right, one to many or one to one. Anytime we have many to many, you need to start questioning it depending on what the data is. Anyway, we now have a relationship established with this. So let's actually get into analyzing this with our calendar based on this day of the week and seeing what is the proportion that they're turning out during the week for job postings. Closing out the Power Pivot window, I'm going to go in and create a new sheet. From there, I'm going to go go "Insert," "PivotTable" from "Data Model." We're going to do it in the existing worksheet. Underneath "Calendar," underneath "More Fields," I'm going to drag in "Day of Week" into the rows, so it has Sunday all the way to Saturday. Then from there, remember we created that "Job Count" already, so I'm going to take that and drag that into the values. So looking at this, I can see that I think our relationship is not set up properly 'cause we have basically the blanks at 32,000. I think I know what's going on with this. Let's go back into the Power Pivot window. In "Calendar," when we select the date, it's of the time data type "Date"; it also has this format of date and time. I don't that really matters too much, but if we go into "Data Job Salary" and we go to that "Job Post to Date," because we use that format function, right now the data type is auto of text; we need it to be of date. And this now looks a lot more similar to what does on the "Calendar." Now when I close out of this, bam, all the values pop up here. So don't forget about your data types and making sure they're match within the data model. So let's actually visualize this by inserting a pivot chart. And bam, we get this bad boy, which we'll rename to "To When Are Most Jobs Posted During the Week," and it looks like we have well on Saturday, Sunday or the lowest, obviously during the week it's the highest with a basically a higher amount on Wednesday. So pretty cool analysis that we were able to do based on the day of the week; we didn't have to create any additional things. And additionally, we can evaluate based on this "Calendar" table created; we can do other analysis such as by the year, month, day of the week, and whatnot. All right, so that's a brief intro into measures and also calculated columns. Don't worry too much if you're not feeling too confident with them just yet, as one you have some practice problems to go through to get more familiar with it, but the next lesson will be and the next two lessons will be on DAX and DAX advanced in order to explore different formulas that you can also use inside of your measures and also calculated columns. All right, with that, I'll see you in the next one where we're getting into DAX. See you there.
Welcome to this lesson on DAX, or Data Analytical Expressions. We used it a few times before in the previous lesson, but now we're going to go much more in-depth and actually understanding the basics of it. Now, as we've learned, DAX can be used within measures or even calculated columns. For the purpose of what we we going through in the project, we're not going to create any calculated columns, but we will be using it for measures. For this, we're going to be focusing on three major types of functions in this lesson, specifically around aggregation, statistics, and also filter. These functions you're going to notice are very similar to your Excel functions that we did back in Chapter 2, so a lot of those similarities and concepts we've learned already are going to be able to be applied to this, so we'll be able to move pretty quick.
Now we're going to be answering two major questions regarding our final project. The first one involves calculating the number of skills required per job title. We're going to use DAX in order to calculate this, and then we're even going to go on to actually graph this to show how it correlates with median salary. Spoiler alert: the more skills you have, the higher median salary you can expect. From there, we're going to go into a deeper analysis of salary, specifically looking at the median salary and specifically being able to compare it from your home country to the US and also non-US countries. So we're going to use the filter function in order to be able to view these things within a pivot table.
Now jumping right into Excel for this, you can continue working in the Excel file that you have from that first lesson on Power Pivot intro where we created this visualization right here, which analyzes top skills of data nerds and has some filters for job title and country. If you don't happen to have that file anymore or you got lost along the way, you can just use the Power Pivot intro part two file, and you can start from there. Now if you're loading it via the Power Pivot intro part two file, you're going to have two sheets in there, one "Skill Job Analysis" and then also the "Skill Analysis." We're not actually going to be using the "Skill Analysis," so you can feel free to delete this. Or conversely, if you're working from the files that you've been building up during this and didn't necessarily load from the Power Pivot intro part two file, you may have multiple tabs in there. Once again, I only care about this "Skill Jobs Analysis" where we have this this is what we're going to keep for the final project, the "Job Analysis," and and also this other one that we created back in the Power Query lesson. We're actually going to be recreating it with Power Pivot, so both of these I can just delete, or anything else you have in there, you can feel free to delete after holding control and selecting both of those.
I'm just going to delete them all. Right, so we're going to be looking at aggregation functions first. Conveniently, Microsoft has some documentation around the DAX functions and also statements that they have. So I'm going to dive right into the link that's provided on the screen underneath DAX functions. Specifically, I'm going to go into the aggregation functions. They have this page here on aggregation functions overview, and it shows a lot of the different functions they have for this: average, count, Max, Min, sum. Let's look at count real quick. Count is pretty simple; all we're going to do is use the following syntax: count, and inside of it you provide a column. And for this, it says, "Hey, the column that contains the values to be counted." So pretty simple function to use.
Similarly, we have distinct count, which has the similar syntax of: you provide distinct count and the column, and the column that contains the values counted, and it will return the number of distinct values in columns. We're going to use this. So what we're going to be calculating with those functions that we just went over is trying to find out how many skills per job. We're going to first go through based on a job title and find not only the skill count but also the job count, and then we're going to take both these values and divide them to get the skills per job. So I'm going to create a new sheet for this, and inside of here I'm going to insert in a pivot table from our data model. We're going to do in the existing worksheet. For the rows, we're going to go through the do data job salary table, and we're going to put that job title short into the rows. And then now we need the skill count. Remember, we could go in and do something and create an implicit measure by throwing job skills and the values. We want an explicit measure because we're actually going to be using the skill count in a later calculation to find that skill per job. Anyway, how do we do this? Well, we can also not only create a measure by going to Power Pivot and underneath here going to new measure, you can also just select in here which table you want to use. In this case, I'm doing a skill count, so I want to contain it in the data jobs skills table. Doesn't really matter which table I'll put it in, but I just go by my memory of which one I'm going to know to go look at for which in there. It auto-selects that table of data jobs skills. The measure name is going to be skill count, and then for the formula itself, we want to do a count of the job skills column from the job skills table. Make sure it's not from the job salary table. Okay, I'm going to put a closing parenthesis on this, and then for this we do want to format it to use a 1,000 separator and zero. Click okay. And now in the data job skills table, we have this explicit measure. Can drag it right next to it; same values are getting created as the implicit measure. So I'm going to take out that implicit measure.
Next thing you want to calculate is that job count. We're going to be counting it based on the distinct values of the job ID. So I'm going to go to add measure. We're going to call this one job count, and we'll do a distinct count of—we want to do it of the job ID column—and for this one we want to make sure that we're actually doing it from the salary or data jobs salary table because this has all the job IDs in it. Once again, we're going to format as a number with a 1,000 separator and click okay. And then I'm going to drag at the bottom; the measure is going to appear. I'm going to drag it into here. So now we want to get how many skills per job. So we want to take the skill count column and divide it by the job count column. This one doesn't really matter too much because it contains both of them, but I'm going to put this in the data jobs skills table. I'm going to call this skills per job. Now what's great about these explicit measures that we just created is I can go, "Hey, I want to do this skill count, and I want to divide—divided by the job count," and it's right there. So you don't have to necessarily write out every single time, "Okay, I want to do a count of the job skills column and then divided by a count of the job ID column," which actually needs to be a distinct count. Anyway, this is where we run into errors; that's why the explicit meas—are so measures are so great. All right, so I have skill count divided by job count. I'm going to create it as a number, and I want one decimal place for this. Go ahead and click okay, and then we're going to add this skills per job to here. Now I'm actually going to recommend, although we just use the division sign, I'm going to actually recommend this divide function with it, which is a math function. And what would you do in this case is you would provide divide and you list a numerator and a denominator. And the reason why I like this is because it fixes any type or catches any error. Specifically, it performs division and returns alternate results or or blank on division by zero. So we're not going to necessarily error out if we have a division by zero issue, and you can actually provide, as shown down here in the alternate result, the value returned when division by zero results in an error. So you could actually catch that any—So I'm go—going to go back into that skills per job, and I'm going to go to edit measure. I'm going to change this to divide, specify the first and second parameter with a comma, and then click okay. Okay, overall no real change here, but just a best practice to know about.
So now with this skills per job, I want to actually get in and comparing this to median salary. This is what we're going to be building right here. We're going to be comparing it to median salary and then graphing it in a scatter chart in order to see how these different job titles correlate to each other. So first—so to know what the final analysis is going to be of this, I'm going to rename this sheet appropriately, specifically calling it salary vers skills. And this pivot table here, we don't need necessarily the skill count or the job count; we just need the skills per job. Okay, we're going to calculate now the median salary. And median is a statistical function, which is encountered underneath here, but there's a lot of different options underneath here, such as Med—median, finding the different percentiles like we did back in the formulas, looking at things like standard deviation and whatnot. So a lot of good statistical functions that you have access to via DAX. So for this measure, I'm just going to come up here to Power Pivot, go under measures, and select new measure. I do want this in the data job salary table, and we're going to call this median salary. For this, we're going to be using the median function, and we need to provide it a column, specifically that salary year average value. For formatting, we're going to format it as a currency with zero decimal places since it's a salary. So now we have median salary here. I actually want it to appear on the Y axis, so I'm going to throw it over here on the first column. So now we have the median salary and skills per job. I'm just going to rate these or sort these from highest to lowest to see if I can see visually if there's anything going on with a correlation. Right now, I am seeing some higher skills than—with a higher salary, but let's actually visualize this. So I'm going to insert pivot chart and select PIV—pivot chart. For this, we want to enter a scatter plot. And if you remember back from our charts lecture, we're going to have issues with this; you can't create this chart with the data inside the pivot table. Doesn't natively support creating scatter plots; kind of annoying if you ask me. Anyway, let's X out of this. And for this, what we're going to do is we're just going to select this area starting up here; we're going to set it equal to this entire table right here. I'm not going to capture the grand total at the bottom because we're not going to be plotting that. Now with these values, I'm going to select the contents in that—this column F and G—and then from there go insert a scatter plot, specifically this one right here. I can see it already looks pretty good. You can't actually add the data labels in whenever you create this chart. We actually have to go about doing that somewhat manually. Specifically, we have to select on the data points and then right-click it, and we have to select add data labels. Okay, now it's giving us points which—bar—which actually correlate to the skills per job point. It's not what we want; we want to include the job title. We're going to add that. So we're going to do is select one of those values and just right-click it, and then from there select format data labels. Then the pane's going to open up on the right-hand side, and it should pop you up underneath label options—label options—then this label options, and right now we have this Y value selected. That's not what we want; we want value from cell, and it says, "Hey, select the data label range." What we want is right here all the way going down; it's hidden behind here. I'm going to sort of guess, but I know it goes down to E11. Click okay, and scrolling it over—bam—we got all those data labels on there now. All right, so now we need to clean this bad boy up because, well, it's a hot mess that is all up in the upper right-hand quadrant. Labels are overlapping; we're going to fix all of this. First thing is I'm going to correct the axes. So I'm going to click on the Y—or click on the X axis—and it should go immediately to this minimum axis underneath access options, and I can see the first value stops around or begins around 880,000. So I'm going to change this to that and press Enter. Okay, similarly, I'm going to select the Y axis, and if—doesn't go to—it should be under access options inside that format access pane, and I'm going to select this first value that I want to go to is three. I'll leave the default of nine there. Next thing is we need some axis labels for the Y axis. We'll call this average skills requested. For the X axis, we'll call this median salary, and we'll specify the units of USD. Speaking of which, this is not formatted correctly for how we want the numbers. So under that format access pane, under access options, and under access opt—options again, under number, we can go to the custom option. Specifically, you should have this type hopefully appearing up; if not, you can just enter it into this format code below and then press Enter. All right, the last two things to do is rename the title, naming it Do more skills equal more money for data nerds, which from this chart it looks like it does, and we can actually confirm this if we want by adding a trend line. Now there's different options here for trend lines; we've going over linear, exponential, IAL—linear forecast. I feel linear best meets this need here; also like the coloring aspect of it, so we're going to go with that. All right, the last thing to do is just fix some of these names on here. So right now we have the data labels appearing to the right of the data point, and in cases where it's close—so data—senior data scientist—it's too close to the edge, and so it's just sort of over the top of it. Anyway, what you can do is actually select it twice, so click it twice, then you can drag and drop it, and it should have these arrows or these connectors that connect the name to where it goes to. All right, so now we have our final visualization, and I'd say it's not too bad. Some things I'm noticing about this—some correlation—if you notice, yes, we do see the average skills requested are going up with the salary, but those jobs—I mean, if you—you can pretty much see it—they dividing line—those jobs that end in engineer versus analyst or scientist are commanding or requesting more skills, but yet have sort of a similar pay to their data analyst or scientist counterparts. So I don't know; I guess it kind of pays to be a data analyst and not a data engineer. Don't tell my data engineer friends I said that.
All right, last analysis we're going to get into is using filters to actually aggregate. So in this case right here, we're showing what we're going to get to—the final thing of—based on a job title short value, what is the median salary in this first column for the US, then what is the median salary for non-US, and then finally that final column of median salary, what is the median salary of—in this case, the selected column is—uh, Argentina. It's filtered down, basically. I call this filter function. We're going to go over, but we're going to be calculating or figuring out how to prevent filters from affecting a visualization so we can get core values what we may want. So we're going to create a new sheet, and I'm going to call this salary analysis. Like before, we're going to insert a pivot table from our data model, insert it into this new sheet, and we're going to be putting that job title short into the rows. Now we're obviously with this going to be calculating median salary, so I'm going to go ahead and just drag that into the values to start getting those median salaries. Additionally, we're going to want to include a slicer in here, so based on the job country. So I'm going to insert slicer on job country. Click okay. And then with this, we can actually see if we select something like Argentina; it's going to filter down to what it is or what the salary—median salary is in Argentina. But remember, we're trying to add two columns to this so we can compare these values of something like Argentina to US salaries and maybe non-US salaries. So basically countries outside the US. Anyway, we're going to be using filter functions for this. And for warning on this, it says it here: the filter and value functions in DAX are some of the most complex, powerful, and differ greatly from Excel functions. So there's going to be a little bit of complexity here in understanding this. And for this filter function, we're going to be using this one on calculate, and what it does is it evaluates an expression in a modified filter context. Calculate is pretty simple in my opinion. First, you provide an expression, so such as, "Hey, perform a count of this column or a median of this column." From there, you provide a filter or filters, and as it states below here, filters can be Boolean filter expressions, table filter, or filter modification functions. Main thing is here we're going to use things like logical operators in order to compare this to maybe a certain value we're going to expect. So let's jump into creating our first one with median salary, evaluating for median salary in the United States. So I want to create this measure inside of our data job salary column—sorry, data job salary table—and for this we're going to call it median salary US. We're going to be using the calculate function for this, and inside of here we're going to insert the ex—an expression. So in our case, the expression is the median of the salary year average column, and what we're going to do—actually, I'm just going to leave this is 'cause filter is optional. We can tell filter is optional based on the square brackets around it. I'm going to just close out this calculate function, change this to a format of currency with zero decimal places, and then from there take that median salary US and actually drag it onto here. So right now calculate is working by calculating the median salary, and there's no filters applied to it, so pretty simple. So let's go in and actually edit this measure now. Now remember we have an explicit measure of median salary, so I actually don't even need to define it like I did here. I can actually just call out median salary in this case. Clicking okay—still the same value—going back in and actually editing it, we now want to apply a filter. Specifically for this filter, we want to make sure that the job country column is equal to United States. So I'm going to type in job country, and we can use logical operators, so I'm going to use an equal sign right next to this, and I'm going to specify United States. Need to make sure it's spelled exactly right. I know it's that via the column. Okay, so now we're going to leave everything else as is. Click okay, and bam, now it has the median salary filtered by the US. And I can confirm this by scrolling down to the United States, clicking United States, and seeing that these values are the same. But no matter what I actually click, the United States median salary is going to stay the same. Additionally, if you noticed here when I click on something like the US Virgin Islands, would I—am I moving there?—they only have four job titles available. So because of that, they just filter this table down to only show those four that are applicable, it along with their applicable salaries in median salary in the US. So now let's calculate the median salary for non-US countries and actually see how they differ. So come into D job salary, select add measure. For this, we're going to be using non-US values. Once again, we want to use that calculate fun—function on the median salary measure that we created. And for this one, we're still evaluating the job country, but we want it not equal to. So we're going to use basically a less than and greater than sign right next to each other, say not equal to, and we'll say United States. We're going to format this as a currency with zero decimal places. Click okay, and then add this bad boy to the values, and I want to actually see a country with more job postings in it, so we'll go to something like Australia. And now something like Australia, we can see one—comparing US to non-US in general—the US—well, except for data engineers—yeah, it looks like only data engineers are the lowest one in another country. Everything else is higher in the US. But now we can, with this one, compare, "Hey, what does it look like something like Australia compared to US and non-US countries?" So super useful in actually filtering down, providing the right context for what we want to look at. So as a data analyst, median salary is around 100,000, which is higher than US and also any other non-US median salary. So may have to move to Australia. One last clean up right quick—slicer itself; I don't like it to say job country. We're going to name this to country. All right, now wrap up the analysis for this. All right, so you now have some practice problems to go through and test out these different DAX functions that we just went through along with some others. Now in this lesson we just did some basic DAX; in the next one we're going to be moving into some more advanced DAX features that I do find myself using from time to time. But overall, most of the stuff we apply in this lesson I use day-to-day. All right, with that, see you in the next lesson. We'll be wrapping up basically our final question in our project and be done with our project. See you there.
All right, welcome to the last lesson in this course, where we're going to be going over more advanced DAX. Specifically, we're going to be focusing more in-depth on filter and also relation—or relationship type functions. These are going to be needed by our data model in order to calculate what is the salary or median salary for an associated skill. If you remember back to a few lessons ago, we had relationship issues. I know I feel that—with having them—being able to filter tables in certain directions, and we're going to be able to see that and fix that in this lesson. So in this lesson, you can start with some—the workbook from the last lesson, or if you got lost during the way, you can go into the DAX intro workbook. Now let's do a quick overview of where we're at with which analysis we've done for this project. We've identified what are the top skills of data nerds along with different filters to filter for whatever our interest is. In my case, I'm looking for data analyst in the United States, and I can see…
That, SE, SQL, Excel, and Tableau are some of the highest. Additionally, we've zoomed out a little bit and been able to identify, based on job titles, where our job title of interest falls compared to others and how many skills it requires. For data analysts, it's right above business data analysts, and based on the number of skills, it looks like it's appropriately rewarded for the median salary. Then, the final thing we did was be able to analyze, additionally, based on data analyst, we can look at different countries and compare it not only in that country but to within the US and outside the US. So a lot of good stuff related to, well, data analyst that position and analyzing the salary. But what about skills? Well, we haven't done that yet. We're going to get into actually analyzing, in this first portion, analyzing what is the expected median salary based on one of the top 10 skills. We did this back in the power query lesson, but now we have this new data model; we need to recalculate it. Anyway, we're going to run into some issues with the data model, as we're going to find out. Additionally, we're going to be calculating the skill likelihood instead of skill count, basically finding the percentage of a skill in a job posting. This is somewhat complex, so this portion here will be optional, and you'll be able to use job count instead if you don't want to follow along with this skill likelihood.
Anyway, back in your workbook, whether you started from that uh Dax intro or you're continuing on with from the last lesson, we're going to create this new sheet for this, and for this we're going to name this skill salary analysis. As usual, we're going to go in and insert in a pivot table from our data model so we can get into analyzing the skills. Going click okay, insert it in, and so for this I want to analyze what is the median salary for a skill. So if I drag the job skills from the data jobs skills table into the rows, we have all the different skills pop up underneath here. And then if we went up here and then tried to drag, or we will be dragging in the median salary into here, all these values are going to be the same. Addition, we get this popup right here that relationships between tables may be needed. Basically, we're running into an issue with our data model. Even if I click autodetect, it's going to tell me no new new relationships are found. So what's going on here? Well, let's actually analyze our data model by going to manage and then inside of here go into diagram view.
So the air resides with their filtering direction. Remember this arrow right here signifies which way we can actually filter our data. So in our case, we have job skills, which is over here in the data job skills table, and we're trying to find the median salary. The problem is is we're basing that off of that salary or average value that's in the data jobs salary table, and based on the direction of this arrow, we cannot flow in the opposite direction. This is what we're call one-way or single filtering. Now, unfortunately, Excel doesn't support bidirectional filtering; however, in things like Power BI, you can actually go in and change it from single filtering to both or bidirectional filtering. Kind of makes me wish I was in Power BI right now. So back in Excel, we can't actually control this via here and actually click it to change this to bidirectional filters. We can only control the relationship itself, but we we can use DAX to fix this. Now, in order to fix this relationship, we actually have relationship functions inside of DAX. Specifically, we're going to use this crossfilter function. With this function, you put inside of crossfilter the column names. So in our case, we can specify basically the job ID from job salary and the job ID from data job skills, and then from there we specify the direction. Which the parameters under here we can go into what we can provide to directions. We can either provide none, basically don't create a relationship, both, which is what we want, filters on either side, or one way, which is what we have already. We're not going to use this. You also control filters left or filters right; the one way we're also not messing with that. We want both. Now this crossfilter, remember, is a filter function, so we need to use this in an appropriate formula that we already know, calculate, in order to filter. I'm going to X out of this box right here cuz that's not applicable. What we're going to do is I'm going to calculate median salary, or a new median salary if you will, inside of the data jobs skills table, and because it's uh going to use the same name but we're going to keep it in a different table, it'll be perfectly fine. And then for this, remember we want to use still calculate; we want to have an expression in here. In our case, we want to calculate what is the median salary, and we'll just use the explicit measure that we already defined. Then from there we'll get into the filter one of what we want to actually filter; we want to provide for this crossfilter. And for this we're going to specify the job ID of one table along with the job ID of the other table. Then for the filter type we're going to use both. Okay, I'm going to go ahead and close this. Now we're calculating median salary, so I want this formatted as a currency with zero decimal places. I'm going to go ahead and click okay and have an error in my formula. Should have known that by the X; I need to actually put a closing parentheses on here, and I'll lied to you, a measure, a column with the name median already exists. Okay, I thought we could do that. It's silly me, so we'll name it median salary skills. Go ahead and click okay. Okay, now I'm going to drag this into the values, and we can actually see with this one now that the associated median salaries are actually there and it's not all that 115,000, which is basically the median of the entire data set. So I'm going to go ahead and move this other median salary out of here, and from there we're going to also drag skill count into here. I just want to look at the top 10 most common skills in this case, so I'm going to go up here into our filter and go to our value filters for top one dot dot dot. We want the top 10 items by, in this case, skill count, and then from there based on these top 10 skills, I'm going to sort it from largest to smallest. But like usual, this is no good unless we don't actually analyze for the country and also for the title or job title. So if I actually go back into that skill jobs analysis, I can just select these two slices right there, pressing control, then copy it and paste them into here. Now you may notice whenever I'm clicking this, this is not affecting this pivot table right here, so we can actually inspect this by going to the slicer and going to report connections. Right now this slicer is only affecting the skill job analysis tab, so this one right here. In our case, for this job title, we actually want to affect it on this page here of skill salary analysis, which is right down here. Click okay. Looks like the salary is updated. Also, we want to do the same thing for Country, adjusting the report connections for this as well and selecting this one right here for underneath the sheet of skill salary analysis, clicking okay. Bam, it updated as well.
So now looking at the top skill of data analyst in the United States, which I'm pretty familiar with, I can see things like Python, Oracle, and Tableau are top three. Excel does make the list, and it's the second to last at 84,000. Now with this, I do want a visualization with it, specifically I want a combo chart showing this. So I'm going to go into insert pivot chart, pivot chart, and for this go down to combo. For this I want the median salary to be the main focus, and then for the skill count we're going to put that on a secondary axis because right now it's just way too low if we keep it on the same axis. And this has the format that I want right here. Go ahead and click okay. I'm going to hide all the field buttons on the chart. I'm going to add a primary vertical and also a secondary vertical axis along with a chart title, and then for the legend itself I'm going to click it and then right-click it and go to format Legend. And for this it should go under Legend options, Legend options, Legend options. I'm going to unclick this of show the Legend without overlapping the chart, and I'm just going to move it up here. So not bad. I don't necessarily want this orange line right here for the skill kind; I don't really feel like a line is best to signify the count. Instead, what I'm going to do is select the line, and if it doesn't appear the format data series, you can also just right-click it, go to format data series, and then underneath fill and line they have line but also marker. For the line we're going to go no line, and then for the marker we're actually going to change the marker options to builtin. We'll change it to this square; it's going to be fine, or we can change it to a diamond. We'll make it slightly bigger, and I don't really like the color, so I'm going to go into design and change the color to this monochromatic palette 8. Nope, never mind, not that one. I meant monochromatic palette one. I want the bar charts to be more visually popping than the actual markers themselves. I change the title to what's the pay of the top 10 skills and then change the primary access to median salary USD and the other one one to job count. Closing this out and then making some room over here for the actual visualization itself. So now we have our visualization that we want that looks at this and be able to show us what are the top 10 skills for data analyst and their associated pay.
Now one last thing for this regarding slicers, I want to actually make it to where they're connected between the charts. So right now I have it to where this, basically this one for skill salary analysis tab, if I go over to the skill job analysis tab, select business analyst, it will change, then go go back to skill salary analysis, it updated to business analyst. Anyway, I wanted to, if we change a slicer, to make sure that it changes on the appropriate sheets. So the job title slicer is only on these two sheets; actually that one's perfectly fine, but the one we actually have concerns with now is the country, specifically on this one. I'm selected on the United States; the skill job analysis one, it's also on the United States and updates appropriately, but then if we look in the salary analysis, that one's on Australia; it's not updating appropriately. So we need to go to slicer report connections, and we're going to be putting the country one on all the different sheets. So I'm going to go ahead and select all the sheets for this. I'm going to do the same for skill salary analysis country slicer, which it looks like it updated along for the skill job analysis. So what I'm going to do is actually copy this now and put this into the salary verse skills because we're controlling it on this page as well. And so now whatever I select, select something like maybe United Kingdom, it will update appropriately and update on other sheets as well. Anyway, quick one, quick note because we move those titles around that one time, sometimes it's not going to match up exactly how we had it before. If you recall, I'm going to go ahead and select all; we set up these text boxes in order to view them whenever basically all countries were selected. So that is one of the issues about dragging and dropping those titles and making them stick to a certain location; it messes it up, your filters whenever you want to filter down for something like the United States. So this wraps up basically our four major analysis that we did. Now I'm going to take it a step further; this portion will be completely optional, and that's this. Right now we're using skill count in order to look at what is, you know, the skill count of, in this case, for data analyst in, we'll do United States. We see that SQL is around 400, 4,000, and that Excel is around 3500, but what does that actually mean? Well, if we go to the future file of what we're going to get to, we're actually going to be calculating a skill likelihood instead, which in this case is looking at what is the proportion of a skill compared to all the different jobs that are available for data analysts in the United States. And so that 4500 and almost 3500 is equal to, well, greater than 50% for SQL and about 40% for Excel. So that makes, in my mind, a lot clearer how important that skill is over account, in that you probably should be learning SQL and Excel as a data analyst.
So back in our sheet where we're actually calculating with the job count, how do we calculate this? Well, let's actually get to moving this over to here. Go back into our pivot table self, and if we throw up the job count, you may get this relationship between tables maybe needed. Don't worry about it too much. Now these values are all stagnant based on some issues with the filter direction, but that actually comes to our advantage because for our filter right here, specifically data analyst in the United States, the amount of jobs that actually are are 8339. If I actually remove both of these filters, we would expect it to be the total rows of the column, which is 32672. So coincidentally, this is actually doing what we need; we just need to get a percentage of these two values, and that can be done pretty easy. So let's open the show field list and actually get into creating this measure. We're going to create in the data job skill table; we'll call this skill likelihood, and what this will do is take skill count and divide it by job count. But remember, we probably want to use the divide function for for this, so putting in skill count and then job count. Now there's no option to format this as a percentage, unfortunately, so I'm going to go ahead and click okay. From there I'm going to drag the skill likelihood into the values and go through and format this appropriately, selecting that it's a percentage, and then with this I'm going to select something that a value that I know what it should be of data analyst in the United States, and with those values selected I can see that Excel is at 41%, which I know that's what it is, and SE is at 53% for these values. Bam, we have this skill likelihood. Now we can now go in and remove these other two columns of skill count and job count, and then from here actually move this graph back over, and unfortunately with the adjusting to it we actually have to fix this and turn this back into a combo chart. So we're going to design, change chart type into combo, select for the skill likelihood; we want this to be on the secondary axis. Click okay. Go back to format data series, remove the line, and then change the marker option to be builtin and to be that diamond at 6 point, and then finally update that secondary access to basically say it's skill likelihood. Bam, now we have this final visualization. Now there's one more that we actually do need to clean up, and that's this one right here: What are the top skills of data nerds? Right now we're doing a count of the job ID, an implicit measure, which you know how I feel about that; we should use an explicit measure, specifically we're using skill likelihood instead of that and remove that count of job postings. Once again, I need to actually format this as a percentage, so going to home, change it to a percentage, and then from there clicking in it and sorting from smallest to largest. Bam, for this one, data analyst in the United States once again, we can actually see visually what are the top skills for this. So now we just updated both of these charts to have a more representative understanding of what's going on with the data. All right, so you should be super proud of what we just accomplished in this project, going through both power query and power pivot and actually diving deep to understand some key statistics about top paying skills and also top skills you should be targeting depending on what job you're pursuing and what country you're in. Now do have some practice problems; go through and test out some of these more advanced functions, specifically this crossfilter function that we went over. Then after that, in the next lesson, we're going to be getting into how we can actually go about sharing this project for those that purchase the course. Practice ice problems and also certificate; you can now go through and complete that end-of-course survey, and you'll be rewarded this course certificate. Now if you didn't do this, it's not too late for you to go in and purchase the course, so way you get this course certificate, all you got to do is go in and take that end-of-course survey and you'll get it. All right, congratulations on your work so far. See you in the next one.
All right, congratulations again for finishing that last project in this video and the next video, which are the last two videos of this entire course; they're going to be focused on how to actually go through and share your projects in my recommended way. Specifically, we're going to be sharing this on GitHub so that way others can see it. Here I am on GitHub, and also if you didn't notice there where you actually downloaded all those Excel files at the beginning of this course. Anyway, inside of here is where I'm hosting my different projects, and you've gone through and probably seen this, but you may not have clicked on something like the Project One dashboard, and in this case, yeah, I have the Excel file, but that readme in there displays below this, and this is what we're actually going to be doing in the next two videos to set this up and then create this readme, and this allows you to detail all the different skills that you used along with detailing all the different analysis that you did while going through this. Now that was Project one; Project two is going to follow a similar method, in that it has the Excel file and the readme, and then in the readme itself it details all the different work that we did in it. So you may be like, Luke, why the heck am I going to be using GitHub in order to share this project? I'm not familiar with it; I don't know how to use GitHub at all; why am I going to waste my time with it? Well, I think it's useful not only in Excel but also other technologies, specifically programming. Here I have my SQL project for my SQL course, and this this is where I host my SQL code and all the different analysis that I did for it, and similarly for my Python course and the project we creating that I also hosted on GitHub and detailed all the different the steps that we did along with all the different uh Python files associated with it. So more the story is I think GitHub's a great tool to use in order to share your work, not only in Excel but also other tools. Now if you recall from Project one, we walk through the steps to quickly share your project on OneDrive if you had it accessible via like a paid Microsoft subscription, and this provided a method to go through and share if you go up here and actually copy the link, a usable link for others, whether they have Excel or not, to actually go in and then manipulate your dashboards that you have. So you may be wondering why the heck are we not doing this with this second Excel file that we created with all of our analysis and then sharing it via this method? Well, if you're called back to this handy-dandy table of the different Microsoft versions and the different skills or basically technologies within Excel that it uses, Microsoft online, which where we hosted that first project at, doesn't have the capabilities of power query or power pivot. Because of that, I could go through the process of adding the second project to this, which it's this file right here. I'll open it up, then actually investigating it, well it does, if you investigate all the different sheets, does go through and actually show the analysis that we did, but if you actually get into manipulating it, like in this case, let's say I wanted to see what are the top skills of data analyst, you're
Going to get this popup right here that says, "This workbook contains external data connections or BI features that are not supported." Basically, Power Pivot and Power Query aren't supported. It can't actually query the data; it's just showing the basic last snapshot of the data right here, and you can't manipulate it. So, in this case, Microsoft Online becomes pretty useless. That's why I'm recommending sharing it via GitHub, as you can share all the associated files with this. If somebody wants to, they could come in here and download it, along with going through and actually detailing what you actually did. So, basically, controlling the storyline and sharing what the different analysis or insights that you actually found.
Now, this what you're reading right now is a README, and it requires understanding Markdown and how to write in Markdown. So we're going to be covering that more in-depth in the next video when we get into Markdown and creating the README. This video is going to be primarily focused on just getting this project into GitHub. So what are we going to be doing for this? Well, we have five major steps we need to get through. The first thing is installing Git, which is the core technology used behind GitHub. We'll explain more in a bit. Second and third, we'll be going through actually setting up our GitHub account and then installing GitHub Desktop to then manage with Git our different folders and projects. And then fourth and fifth, we'll be basically initializing the repository, which is a fancy term for a folder, and from there getting that folder repository onto GitHub to then share.
Before we install it, what the heck is Git? Well, similar to how they have track changes and stuff like Word and PowerPoint, Git does this. It's a Version Control System. It tracks changes in not only files but also code, and because of all this, it allows you also to collaborate with others when working on a project. Git is the core technology behind managing all these different things going on on your own local computer, and then whenever you make any of these changes, GitHub is where it keeps track of these final changes, if you will, and then displays it for the world to see and also pull those changes.
Here's my Excel DI analytics course right here on GitHub, and I have the same folders or repository on my own local computer. Now, there's actually hidden folders or Git folders in here managing this, and I can do a shortcut on Mac of command-shift-period to show that. But anyway, I wanted to mainly show this of this .git folder in here, and this thing I don't necessarily touch this at all or work inside of it. This .git folder contains all the different revisions and tracks all the different changes within my project. So, in order to get this .git folder inside your project and then also get it into GitHub, we need to actually install Git.
So navigate over to the Git website, into their downloads, select your operating system choice, whether Mac OS or Windows. I want a Windows machine right here, and from there I'm going to select the 64-bit version for Windows and click here to download. Once downloaded, I'm going to open the file. As do I want to allow this to make changes in my device? Yes, I do. And then it's going to walk you through the setup process for Git. All of these things are going to be left as default, so feel free to just go through and select it all. After I've left all the default settings as is and selected that, it then gets into the actual install itself. Looks like it installed properly; we'll go ahead and click finish. We can confirm it's installed by opening something like Terminal, and you should have a Terminal app installed. This is just confirming it; you don't necessarily have to do this. Anyway, mine opens in a Powershell, and you can just type something like "git," and it shouldn't give you an error message. It should instead give you how you could go about using Git via the command line in Terminal. Don't worry, don't be AF of this; we're not going to be using Git via the command line, although I may need to make a separate course on that. Instead, we're going to be using GitHub Desktop to manage Git.
So, in order to use GitHub, you need to have an account. If you already have an account, you can feel free to just sign right on in, but if you don't, go through the whole process of entering your email, providing your different credentials, and then getting logged in. Once logged in, it should direct you to your homepage. If it doesn't, you can come up here to this icon at the top, and from there just select your profile. I would go through at this point and actually customize your profile, specifically adding a picture, your name, a little description, and any social media links over here on the right-hand side. On my homepage, I have some different pinned repositories. Because you just set it up, you probably have none, but this is where we're going to be putting your Excel project when you're complete, so that way if people navigate to your profile, they can see it.
Now that we have this account, we need to actually get our project or our repository onto GitHub. But unfortunately, there's not really an easy method I've found with actually using the UI from the website to do this, and that's mainly because there's a lot of technical things going behind the scenes and managing Git. Instead, I'm going to recommend downloading GitHub's application to install on your computer. They have it for both Mac and Windows. Navigate to this link here, and for this, I'm going to go ahead and just download the 64-bit version of this application. This one's a lot easier to install than Git. From here, once we have it downloaded, I'm going to open the file. The installer should open this window for you to next sign into GitHub. Once you've entered your credentials for GitHub, you'll use this to configure Git, and for this, you're going to basically say, "Hey, I want to use this GitHub account name and email address to manage all this," and click finish. Now it should navigate you to the "Let's get started" screen. Anyway, it has methods for you to go through and create a tutorial repository if you want to. We're going to be doing that, and it has some different options for this that you can also select via the file menu, such as a new repository, add local repository, or clone repository. We're going to be creating a new repository, and as a reminder, repository is basically a fancy name for a folder, but it's a way for us to maintain and collect all of our different files, not for what we're using in our project.
So, for this, we need to give it a name. So I'm going to give it something descriptive like "Excel Project Data Analytics," and for description, I'll just give a simple one of "My project, demystifying my Excel skills." For the local path, we need to actually point it to the folder that has this. So mine is inside my Documents folder, and real quick, inside that folder itself right now, I would expect you to have Project One and Project Two. I also am going to be putting all the different files that I have for the different Excel workbooks that we work through in the lesson. If you don't have them, don't feel like you need it. The main important thing is that you have both Project One and Project Two in there, and I have them conveniently located in different folders inside of here, never getting out of that, so I can select this "Excel Project.Analytics" folder. I'm going to select this folder. It's going to ask if I want to initialize this repository with a README. I do. As far as the .gitignore, I'll put none, and license none as well, and we'll create the repository.
So now you're going to be navigated to this screen here, which is basically the default screen of GitHub Desktop. It allows you to select different repositories. Right now, I have only the "Excel Project Analytics" one. It allows you to select different branches. We're going to stay on one, shifting to another branch is beyond the scope of this course. Then up here at the top, it has something like "Publish repository," which we want to do, but one quick thing, real quick, I can actually investigate what files are going to be pushed up to GitHub by going here into history, and right now it's just the one I selected that box for README, so the README is in there, and the other one's just .git attributes; the other ones aren't in there, and I'm doing this on a Windows machine. Well, if I navigate back to the folder that contains my project, so here I have "Excel Project.Analytics," which I selected, from the GitHub Desktop whenever I go into it, it actually created another folder inside of it, and that has the .git attributes and README that it's talking about. About now I've done this on both Windows and Mac, and Mac doesn't cause this issue of putting another folder inside your other folder. So for Mac users, you may not have this problem, so completely ignore this, but for Windows users, this is a problem because this right here is the project or the folder was going to get uploaded to GitHub. So what we need to do is take all the contents of this by selecting it all and just pressing Control to select it all and then dragging it into that folder. A little confusing, but if we go back to the Documents, we have our "Excel Project.Analytics" folder, then inside of that we have our GitHub repo, and then now navigating back into GitHub Desktop, I go over here and I see changes. We have 85 of 85 different files and folders within there. It's actually picking up on all those different files that I have in there. Once again, if you're on a Mac, you may not see this because it's already in there in history, and you can see it's actually within this portion of the guy. Anyway, the thing now is if we go ahead and publish this repository to GitHub, it's only going to have what's inside of our history right now under this what we're calling a commit, and a commit is a snapshot of your repository at the time that you're basically committing it. So we need to do a commit in order to get all these different changes into a repository, 'cause technically right now they're in an area called a staging area or the working area. Anyway, we need to provide a summary that's required, and I'm going to add something simple like "Add all Excel files." Doesn't need to be super descriptive, and from there I'm going to click "Commit to main." Now if I go into history, I have this initial commit that it did, but then that "Add all Excel files," it's going to then have in all those different Excel files that I added into it. So now that our local repository on your machine is up to date, we need to then publish this repository to GitHub, and we can either click this button or this button here. For this, we're going to keep the same name and description that we have before. We don't want to keep this code private, so we're going to uncheck that box, and then from there we're going to click "Publish repository."
So my repository has quite a bit of Excel files, and the memory size of it is pretty large, so it is taking a little bit of time to do this. So now we've completed pushing our local repository to our remote repository on GitHub. So inside of GitHub, I can navigate up here to the right-hand side, and I go to your repositories, and here it is, the "Excel Project Data Analytics" that we made public, and it's all in here. So now somebody can come in here and see our different work; in this case, our Project One dashboard is inside of here; we have our Excel file in there, and bam, we've set up Git and also GitHub, and that was a push. So now we need to demonstrate what is a pull, and so in order to do that, a pull request, we need to actually make changes on our remote repository, so that on GitHub, and then pull it into our local repository.
So here's what we can do for that. I'm going to just go in, and we created this README.markdown file upon creation because we selected that checkbox. You can actually come in here and edit this README by clicking the "Edit file" button, and I'm just going to come in here, and I'm just going to say, "Hey, I added this on github.com," adding it in the bottom. Now we're going to go into Markdown formats and stuff, as you can see we have this hashtag here; we're going to go all that in the next lesson, but anyway, I made this changes to here. So we need to, like we did on our local repository and making a change, we need to commit those changes here, and conveniently it just gives us a commit message of "Update README," confirm the correct email, and it connects directly to the main branch. We're just staying on that branch; we're not shifting for this course at all. From there, I'm going to commit changes. So now if I go back into the project itself, scroll on down to see the README, I can see that I have "I added this on GitHub," whereas on my local machine, if I go into look at the README.markdown, it doesn't have that addition that I added to the README file. So we need to pull those changes. Going back to the GitHub Desktop app, I'm going to come up here, and you notice that it says "Fetch," or this isn't going to do anything; this is just going to fetch origin, basically the main branch, and pull it in. This isn't going to make any changes to your file; it's just going to update it of what's on GitHub, and we can see based on this that we have basically one change here by this one and this down mark. And so in order to get these changes, we need to pull the origin, pull it, and so I'm just going to click it to pull. And now when we go into the history, we now have this new one of "Update README." We can see that this README has this addition because it's in green of "I edited this on github.com," and then inspecting this in the README itself, it now updated to say, "Hey, I added this on github.com." So bam, we just demonstrated how to push and also pull from our local repository and machine to our remote repository.
So now that we have GitHub and Git all set up, we now need to get into actually building out those READMEs and explaining what we did in our project and demonstrating those skills that we gained in this course. So that's what we'll be doing in the next lesson. If you're getting stuck at any point during the way, I highly recommend that you take use of something like ChatGPT or even Gemini or whatnot and actually paste in your error code, and it will help you with troubleshooting it. It's a lot quicker than posting a comment in here saying that you had an issue. All right, with that, see you in the next one. We're getting into the README; see you there.
Welcome to the last video in this course, and in this we're going to be going over how we're going to actually document all the different work that you did for Project One and for Project Two. We're going to put this into our Markdown file or our README, and then from there getting it onto GitHub and then finally going through how to share it on LinkedIn.
So right now, navigating to our GitHub repo with our project in it, you should have at least two folders in there, one for your Project One dashboard and one for your Project Two. If you have your other folders for all the work that you did for all the other lessons in this course, that's awesome too, but not required. Mainly just have your project work in there. Anyway, we have this README for the entire project itself, and right now it's pretty bare bones. And if we navigate into that Project One dashboard, right now you should only have a file in there, specifically that Excel file, but we need also a README in here as well, so we can add a description of what we did in that dashboard. Similarly, Project Two doesn't have a README as well. Now we have demonstrated in that last lesson how we can actually go into something like the README and then from there edit it inside of your web browser by just clicking this "Edit this file" icon. It shows not only the edits for you to actually go through and maybe type something but also the preview itself, itself of what the file is going to look like. Don't worry, we're going to be going over Markdown syntax in a little bit, but anyway, that's how we're going to be doing all these different changes to the files. For this, I'm not going to do these changes; I'm actually going to cancel these changes.
Now, an alternate option to making edits to something like a README is using a text editor or IDE, integrated development environment, such as something as Visual Studio Code, which is completely free and is—I have it launched here in my app—um, is an app that I use in order to edit and manage my different files. I can also go through, if I'm editing the README itself, I can type inside of here and edit it, but also during that I can actually go in and view what's going on with the actual README itself off to the side while I'm typing here in this other window. Anyway, I just want to make you aware of this that is an option for you to go through, but it does take some experience with knowing how to use VS Code, setting this all up. So based on the complexity we've already built up already, we're going to stick to just editing our READMEs inside of github.com.
So before we get into building our project READMEs, we need to understand some syntax here. Specifically, if you notice this "Excel Project Analytics" is capitalized and everything else is lowercased, and if we actually go in and edit the file, we can see that we have this hashtag at the front, which translates this into a heading. So they have special characters that you can actually use in front or around text to manipulate text, and the team that created Markdown conveniently created this cheat sheet, which I'll link here, and it shows all the different methods that you can actually use to actually manipulate and make different things happen inside your Markdown file. So let's actually look at a few here. I have a heading one, heading two, and heading three, denoted by how many hashtags and a space, and then if I preview this, heading one, heading two, and heading three. Next, we can either bold or italicize text by surrounding it either double asterisks or single asterisks, and the final results right here is bold and italicized. Notice how the bold text and italicize are on the same line. It's important that after you go to a new line you actually put two spaces in there. Now that I have that in there, it will actually shift it to the next line. We can also do things like an ordered list or an unordered list, which would be like bullet points, and it conveniently indents that and makes it look a lot nicer. We can surround something by a backtick, which is located up at the top of your keyboard, or you could do triple backticks at the top and bottom for if you have multiple lines of code, and if we actually go to preview this, we can see that the single line of code was just surrounded, whereas a multiline creates this entire coding block. The final two worth mentioning are links and also images. For the link, for the text that you wanted to appear for the link, you'll put in square brackets, and then for the hyperlink itself, you're going to put that inside a parentheses right next to it, and then actually changing this to a real-world example of something like google.com. If I go to preview and then I click this link, it's going to ask me if I want to leave the site and go to Google. I'm not going to do it because it's going to mess up all my changes, but you get the point. For images, it's very similar, but the text you provide in the square brackets is just your alternate text, so whenever you scroll over it, what the text is displays, and then from there is the actual image location. However, this isn't an actual image location, so I have this error message that goes on with this alt text, hence this broken file. You're going to notice that if any of your files for your images are broken, anyway, github.com actually makes it pretty easy to get images in. In this case, I have a GIF of the dashboard; you could also use an image file, but all I have to do...
Is take it and drag it into here. And if you notice, it automatically formatted it with alt text and then the actual link location itself. So saving the file itself, and it puts that exclamation point at the front signifying that it's an image or in this case, GIF.
If I go to preview, scrolling down, we can see that we have our image. Once again, you need to put spaces after that other one to make sure that you're not having it all in the same line, but you get the point anyway. Let's actually get into creating this readme that's on the homepage, if you will, of our actual project.
The main point of this one is I want people to be navigated to the appropriate project depending on what they're looking for. So I went ahead and put in some text already for how I want to break this down. I'll break uh I'll shift over to preview, and I'm going to have a title such as My Excel Analytics Projects. From there, we're going to have the Salary Dashboard project and the Salary Analysis. Right now, the image that I have for the dashboard is in the wrong location; actually, shift that up. Now I went ahead and added the images also for our Salary Analysis while cleaning up where the Salary Dashboard is, which I included only just two graphs here, but I just want to give a sneak peek of what's going to be inside of those other readmes that we're about to build out.
Now you may be wondering how the heck do I get screenshots of graphs in my different dashboards? Well, depending on if you're using Mac or Windows, they have software installed already, and so these shortcuts should work for you in order to perform your appropriate screen capture. I primarily use, on a Mac, command Shift 4 to select a certain area, and it allows me to basically just hover over something and snapshot it. This same thing can be done on a Windows machine; you're just going to press Windows Shift plus S.
So I went through also and just added a quick description to each section. I'm going into preview because it's a little bit easier to read there. Anyway, underneath this, I just detail, hey, this contains all my Excel files to follow along. In my case, my free course of Excel for data analytics. I would word it differently for you if you're actually providing all your different project work in this repository. Additionally, I provide a short description for the first project and then also a short description for the second project. Make sure, in this case, you actually are putting spaces after those lines so you don't have those images overlay on top of it.
Now the last thing I would do, as you see here, I link to my course, but I think more importantly what you need to do is actually link to the appropriate files within this repository so people can quickly get to the Salary Dashboard or the Salary Analysis. And so I'm going to add this link of connecting to that appropriate project by first adding this text of Check out my work here, and then inside parentheses, I'm going to list the folder of Project One Dashboard. You have to make sure you spell it exactly like the folder that is inside of your repository or the link's not going to work. I'm going to do the same with the Project Two Dashboard as well, and going to preview it, I can see it's all there. I probably want some spaces in between this, and so just put an extra enter in there. Okay, that's good enough. I'm going to get into committing the changes. This is Update my readme; that sounds good. I'm going to commit them.
So now on our home folder of our repository of excel project.analytics, scrolling down, I have my readme here. It tells me about it, and then for the Salary Dashboard, it says, hey, check out my work here. When I click on it, it navigates me into this folder for the Salary Dashboard, which you need to now create a readme for also. It's just good practice to make sure that you check to make sure that other link works as well, and in this case, it didn't. It's a good thing we checked it. I had Project Two Dashboard, and instead it was actually Project Two Analysis. I'm going to commit changes, and then now when I actually try it out, bam, navigates me to the right location.
So now you have now the basics to go through. You understand markdown enough to edit it. I'm going to walk through how I built out the Project One readme and also the Project Two readme so that way you have some understanding of what you should do going forward.
With the Project One, I recommend including a picture of the dashboard to start and then a brief intro detailing why you wanted to do this project. Underneath this, make sure you include a link to the file itself, which is conveniently right here, and then inside of here detailing the different skills that you use with building this. This is really important for job seekers, that way if a recruiter comes and looks at this, they see what the skills are you used in this. And then from there, I talk about the data set itself, talking about what we were trying to get or extract out of the data; so basically all the foundation they need in the introduction portion. This portion I recommend keeping the similar format. The next portion you can feel free to go about however you want. Specifically, I go into the dashboard build, breaking it down into three main areas of focus. On first is the charts itself; I highlight the different median salaries, all of the different job titles themselves. I go into some insights from that. I also talk about the country map and the insights from this as well. Next, after charts, I move into functions and formulas, detailing one of the key functions that we used using median and then an if statement in order to build out an array formula. So not only breaking it down but also explaining what insights we're able to get with this formula. And then the third skill I talk about is data validation, talking about why it's used, a GIF of how it's actually applicable or how it's actually visually seen in Excel. And then finally, I just wrap it up with a conclusion.
So to recap, for the first project, you need an intro statement describing what we're doing and why you did it and what skills you used, then then from there on the build itself, explaining what you actually built, how you use those skills, and what insights you got out of it, and then finally wrap it up with a conclusion.
For the second project, mine is very similar formatted in that I have an introduction, Excel skills used, the data set, and then since this one was primarily focused on analysis, I included the four questions that we went through and actually answered for our analysis. So then with the template of these four questions, I broke each one of those down with those questions primarily focusing on one: what skill did I use to help answer that question? And then two: what is the analysis insights I got out of answering that question? I repeat the same thing for the second question, specifying the skills that we use for this and then the analysis or what insights we got out of it. After going through questions three and four, we then get to our final thing of a conclusion of what you actually learned and extracted from insights for this. So it's really good to put all this stuff in it. I wouldn't be overwhelmed and think you need to include everything in it. Think about a job recruiter themselves; they don't have a lot of time, so keeping it as short and to the point as possible is going to be best for you.
Once you're done actually gone through and built out your repo with all its associated readme, it's time to get into actually sharing this on social media via LinkedIn. I recommend the same approach that we used back in Project One of listing this down in your project section by going through and actually clicking the add icon and adding the projects. If you did go through and actually add that Salary Dashboard already, I would just focus this one on the Salary Analysis. So I'd put in something like a name of the Data Science Job Analysis, a description, add any appropriate skills. There's a ton of different skills you actually select for what you use. I would focus on primarily these: Microsoft Excel, Power Query, data modeling, ETL, and pivot tables. For the media, in this case, I would include a link to your repo and paste it on into here and click add. It will then provide this snapshot thumbnail of what's going on here and a title. I like it all; I'll click apply. Now if you recall back from that first project, we tried to provide the link of that OneDrive link for Excel and it didn't work. So if you have that project on LinkedIn, I would go through and also attach this link as well to that so that way they know how to navigate to it. Finally, select your start and stop date. If you have any contributors are associated with, I don't have in this case, and then from there save it.
The last thing I recommend doing is making a post telling others about your project so they can come in and see it in it. I would definitely include something like a link and feel free to tag Kelly or myself in it. I love checking out your projects and seeing the different work that you've done for it. So once again, congratulations for finishing this course. Been nothing short of your hard work. Excel was the first skill or main skill that I learned in helping me land my first data analytics opportunity, so I feel the same can go for you as well. Now after you taking a short break and you're ready to get back into learning more skills, I do have a sequel course that I recommend you taking. As you've learned from analyzing this data, Excel and SQL are two of the most top skills of data analyst, so it pays to know it, and you can basically learn it in a weekend. All right, with that, I'll see you then either in the next video or in the next course. See you there.