Transcription
Okay, um, so I hope you guys can hear me very well. Uh, so my name is um Siro. I am one of the team members at Zindua. Uh, I'm more vast in beta science actually, so I was the technical metaphor designs for the longest while, although now my focus, I'm on the Ziploc team, is on operations, um, which is fine. Um, so data science is my core focus, and I'll be taking you through um uh the data science pathway today. Uh, so we get to understand just um what are the skills you need to learn, how can you break it down, what exactly is this in the first place, uh for those who are getting started as well, uh and uh we just want to get a general breadth of the exams, how you can start learning these types as well, um, and then you can have the freedom to explore. Um, uh we have a free intro course in data science, and for those of you who want to take it later, I want to explore more advanced designs if you want to go in um for the core program as well.
As I'm recording this, I'm struggling to find a way to make um this screen full screen. I know in your end it is full screen, but for the sake of the recording, I am struggling a bit. I don't know how this is going to be; that recording would have a very, very small screen on the side, so let me see what I can do on this front. Um, anyone with ideas on how to do uh on how on how on how to structure it such that they the the presentation only one person is being seen? I think you pin it. Let me see if I pin it. Okay, uh okay, that that is I think I'll learn the breadth of it as uh as I keep doing um such Google Meets. I've I just want to find a way to make sure it's full screen, but that's that's the best that we can go for right now. Um, all good, we can't complain about it. Okay, so let us get started. Um, today uh huh, okay, I'll just skip there, and then we can get started proper here. Uh, how do we go more full screen of all this? That is fine. Uh, just bear with me. Um, even though the the sort of recording would not be as full screen as we expect it to be, but those of you are in the workshop Hub should have a really clean uh view of everything uh in that trigger awesome.
So let me just go through what you're going to cover today. Um, I've already introduced myself. Uh, so today we're going to first of all look at a general introduction um to data science. Uh, so we'll be covering some of these questions over here: what is data science? Uh, what are the pillars of data science? Uh, what is the typical data science life cycle? This will be more of like the structure of typical science projects. Um, and then we'll explore some applications of data science. Uh, in fact, I'll I'll just start preparing uh more possible applications we've had of data science right now; outside when we get there it becomes uh it makes sense for you guys as well. And then the most important part of this particular uh session is going to be going through the designs roadmap, and you've broken down the roadmap into three things particularly. So the first layer is going to be data analysis or the data analysis layer. Um, so if you want to become a data analyst, um what are the core things you need to learn to become a data analyst? Uh, the second layer will be now for those who want to go beyond data analysis and they want to become data scientists for machine learning Engineers; what must machine learning concepts are you supposed to learn are to be able to get that sort of Next Level job that involves more Predictive Analytics or rather machine learning modeling on that particular front. And then I'll also expound on additional thoughts is that you can explore Beyond machine learning, and this I will tell us the advanced um data science layer, um and this will be for people who want to do further specialization Beyond just typical data science on what they can take it, and after that I'll be able to share the next steps for most of you, whether it's indoor school and what you're able to offer on that it sounds curve, and we'll also share more about the free foundational costs that we offer for data science covering Python that can allow you guys to get started in data science for those who are ultimate beginners on this particular front.
Okay, so far so good. Um, if there's no question, we'll get started. Um, if you feel like there's something um that you'd suggest me to go through in the course of that uh of this call as well, feel free to um type it down on that meet as well, and then we'll give it some time later on in the call as well. Uh, once again, uh because uh I am not on the Google screen, um just let me know in case I go off. Uh, though I'll just put uh my phone on this particular phone so I can make sure that I'm always I'm in the call at all times, so that should be okay. Okay, uh how do I go full screen on this one? Ah, that's fine; that's as full screen as I can go. Okay, so let's start off with um what is data science? Anyone wants to try on what they think it says is anyone any attempt? Okay, Everyone's an ultimate; feel free to type down on the chat on what you think it stands is. Uh, try and change the layout; it's not sure if it will work. Uh, thank you very much, Victor, for the suggestion. I think on the on the recording now I have something that's slightly more full screen, so that should be okay. That's fine, thanks for that. Um, any ideas on uh what is the distance um according to you, what do you think the design says? What does it involve just based off of your exposure? So far so good? How about Stephanie? I don't know if we can hear from you. Uh, any exposure did Sam? So do you think there's a science is uh before I introduce uh my very nuanced uh uh uh explanation? Okay, it is it is a silent call today, sir, so that's a bit sad. I hope we can get energy as we explore more of these things um later on, uh but I'll go ahead and uh state that um uh okay, so I don't want us to focus on the definitions; yeah, I just want us to focus on the understanding uh of of the concept in the first place. So just from the name data science, you can see that this already involves data, um and uh in the realm of this sense we're going to uh it's going to uh it's going to it's going to focus so much on complex analysis of data, um even simple analysis we can call it the science, but it's going to be more of the more complex analysis of data, um and we are doing this analysis for Three core reasons, and I think this is very, very essential to understand uh so that you understand why we have data scientists in the first place. So the first thing we're trying to look for in this data is to figure out if there's any unseen patterns in the data; so that's one particular thing. Uh, I'll try to give examples when going for it. The second layer is we're trying to derive useful information; that's why we're seeing data science being applied a lot in businesses. Businesses are trying to derive useful information from the data the business has collected over time so that they can be able to make meta decisions for their business, right? So if I'm collecting data on my phone um from time to time, so let me let me walk with number one, right? So let's say I'm using my phone, um this point of mine is collecting so much data on a day-to-day basis, minute-to-minute basis. Uh, I have like 50 applications, 80 applications on my phone; I do different things. Uh, probably when I do my complexity analysis of this, I'll be able to see unseen patterns such that from this time to this time I am avidly a consumer of YouTube, or this is the amount of time I waste um typically going through social media and the likes. Uh, so it was to just look at the data analysis of my phone, given the different ways uh it kind of collects data from me, which is pretty much my usage of the phone, I can start seeing some unseen patterns which I usually don't know of how I use the phone or how much time I use on a typical day. So that's a good example of how data science, so this is a very small scale; I'm just looking at my phone um can detect unseen patterns in my own personal life. The second is now to derive useful information; if we have a business like Java that has opened coffee shops all across Nairobi, let's think about it. Um, they're collecting sales data of these coffee shops all through, right? So they know in CBD, This Place Monrovia street, we have a coffee shop that does this amount of sales on a day-to-day basis; this number of customers come in at this particular time, and then they leave and all that kind of stuff. So if you do the analysis of their sales data, which is grouped per branches, um we look at the times, we look at the frequency, we look at the revenues we're getting off of this and the cost of rent for this particular things, then you can start deriving useful information by trying to look at the profitability of each and every single branch and determine where we should either open the next Branch, so what's this high traffic areas that have a lot of business um that should be closer to where we should open our next Branch, or decide what branches we need to close because they are doing extremely poorly as compared to the other branches. So that could be useful information that Java would want to reduce if it was able to analyze data collected from all its um our stores or cafes across across Nairobi, just as an example.
And the final thing is now uh to make predictions and decisions. Now I'll use a slightly more complex example over here. Let's say someone is a Trader, right? They've been tracking stock markets for the past five years. Uh, they've been able to first of all find unseen patterns; a second layer they've been able to derive useful information such that when this company gets to uh summer, so this is an ice cream company, so whenever it's hot, right, summertime in the US, they have an influx of sales, and this is now Q2 um the summer is about to come in, so we expect an influx of sales to come in. So they've been able to drive this useful information by analyzing the past data of this particular company. Now what they're going to use data science for is to make a prediction on what the stock price of this company will be based on their previous data analysis. Now this is a more complex example that we see applying to let's say financial markets, but it doesn't always have to be financial markets. A good example is you could be a real estate agent, right? Someone has given you a house, and you've inspected the house, and you've written down the features of the house, like it has three bedrooms, um it has really good lights, there's a high ceiling. So if you turn down all these things, you realize that when you write down the features of the house, you've come up with a page of about um uh let's say 200 different features that we look at, but because you've been in the real estate industry and you've done data analysis on this, you already know uh the typical features of a house, and you've just been ticking check boxes for that particular house, and now using machine learning uh in previous house data and the prices that you've been able to quote for or you've been able to see being quoted for this house is different features, you can now use the features of this house to go ahead and determine what will be the selling price for the house that you've just expected, and this is one good practice project now we see a lot of data scientists work on as well, which is the house pricing data available on cargo; that is a typical example of how the science is being used to make predictions whereby we analyze features of houses and the prices of these houses over time such that the next time I'm just given the features of a house, I'm able to deduce what is the expected price of that particular house. So quite essentially in that regard, and of course um this is going to be way more profound; I will talk about Healthcare whereby after analyzing a lot of X-rays of patients in two months, doctors are able to find a way where to predict whether someone has a tumor or let's say a cancer by looking at an x-ray and able to detect this slightly earlier because at least it decides able to see patterns that your clean eye was not able to see before, right? As you start using a deep neural network, which is now that visual intelligence layer, um some patterns that typically doctors couldn't see with the usual eye uh now the data can deduce, right, um because uh you can see the pixel change in this x-rays even to the smallest degree such that they can make these predictions early on when some of these things are slightly more curable, so you can start seeing it even for more complex Industries; it's a very, very essential yeah, so that's what with summarize it sounds are so one data is very, very important here; second layers we are analyzing it, and these are the three reasons we're analyzing it. Um, any questions so far so good?
Now to better understand the designs, the best way to look at it is to explore what you call the pillars of data science, um and we have three particular pillars of data, and I'll explain how they all blend it uh into I don't know what this this is what people call a Venn diagram; I hope uh I hope I'm not massacring the word van on this one, but let's call this a Venn diagram over here, and uh from the Venn diagram you can see that there's three fundamental layers that matter to data science. The first layer is mathematics and statistics. Um, there's no data science you uh if if you're not doing math and starts in the first place. First of all, this analysis of data is based off of statistical Concepts, right? It's statistical methods that you're using to analyze this data uh whereas that we we can make hypothesis and see whether our hypothesis check check out or do not check out; that's the first fast layer. Second layer is um we are using mathematical algorithms; so some of the things we're running on machine learning models, they are based off of mathematical concept calculus, linear algebra; at the end of the day we are applying matrices and vectors on this sort of data, applying a bit of linear algebra of it on it, and calculus are to be able to run this machine learning algorithms. So when you have a good understanding of both the mathematical side and the statistics side, it becomes very, very easy to understand um how some of this machine learning algorithms work because they're extremely mathematical; how they are optimized uh to find the best feature to make the best prediction is entirely based off of math and definitely statistics. So people in the math and statistics background highly recommended start thinking about um data science because this is pretty much statistics on steroids, and some people like I like to think about it. The second pillar is now the computer science layer. So in typical statistics, um we know our statistical methods and then we go and use our statistical tool like Microsoft Excel or startup for those of you who are more into Academia uh and then in the recent past is starting to see are coming up, um but R is a programming language though statistically focused uh because we are trying to find a way to deal with mass amounts of data, and we're trying to do very, very clean automation, um the computer science layer becomes extremely important; we have implemented all the statistical Concepts right and mathematical algorithms in a programming language so that things can become more customizable and more automated and so that computers can do a lot of the learning by themselves. So the second pillow of data science over here is the computer science layer where you need to learn a programming skill because you're using a programming language to effect or do the modeling or do some of this analysis over here. The two most common languages you use for it will be Python, which is by far the most profound language of this particular front, and R programming, which is more used in an academic setting. Awesome. Now when you just have math and statistics and computer science, that's pretty much machine learning; the entirely of machine learning is whether you've been able to embed programming Concepts onto statistics Concepts so that we can make predictions and detect this unseen sort of patterns in data as fast as possible; that's the machine learning layer of it. Um, however, without the third layer, which is domain expertise, then there's no data science whatsoever. So domain expertise, understanding that we are not analyzing data in uh in a black cave; we're analyzing data that expect to a particular domain. Doctors are trying to analyze data from X-rays of patients, right, to be able to predict whether someone has a particular disease allele, let's say brain tumors for example, and they want to say whether that tumor is uh malignant or there's another word for it; is it usually malignant versus what uh something else? Once again, uh not not so much into the medical staff of it, but uh there's also a really good test project on that particular front. Yeah, by now you know malignant tumor; a good example, it's a typical question doctors are trying to analyze a typical day-to-day basis. If I'm the best person in math and statistics and I'm very, very good at programming, right, there's nothing I'm going to deduce off of that data if I'm given a couple of x-rays if I do not have a medical background of this particular thing. Now I know with some of this practice data like a lot of the heavy lifting has been lifted for you; why they've been able to do the matching, they did a lot of the thinking in the problem set up early on, uh but in the initial phase where we were even trying to think about how we collect data, right, what type of data we'll be collecting, how we'll be collecting it, what data can we do without and the likes; if I'm not coming to a medical background, there's nothing I can do there. So the best I can do is just be given better and try to optimize my algorithms as much as possible to make a good prediction, but without that domain touch, there's no much more I can do it in regards to improving that data to get a better output. That's why we're saying data scientists should also have a domain forecast or domain expertise to make their skills make sense. Now uh this is extremely easy for some because uh business is by far the most common skill for a lot of people, so a lot of people end up practicing data science in a business setup, things such as sales analysis uh for people with a financial which is more of an in-depth Finance background, then they'll end up applying that science in more Finance cases; that's why you end up seeing a quantitative analyst, people who are practicing data science deep into our consecutive analysis of stock markets and the like so that they're working for hedge funds to be able to do some of these predictions. So those are people with in-depth knowledge on that particular thing. Um, software developers are a good example of people with domain knowledge and product development, so they can apply data science on trying to do the product development to understand what features work for their users based and not and all that kind of stuff. Now let me just summarize this particular pillars of that science over here: when you only know math and computer science, um that is machine learning; that's pretty much what what machine learning entails, right? Um if you're good at math statistics and you have a domain expertise, that's what we see in Academia a lot. Um you can end up doing a very good traditional research whereby you collect your data, you run statistical models or for videos tools such as Excel and starter, right, um because you have domain expertise you know what to do with that particular data and you know what questions to answer with the data and uh how how to go about the data cleaning process it likes. So let me just say that uh once again a lot in Academia if I'm very good in a very good political analyst
And I know Martin starts, then I can easily do polling. Right, a good example during election relax when you only know domain expertise and computer science, then that's where we end up having people who are called Product developers. So you can do product development. A good example is MailChimp. MailChimp is a is an email automation tool, right? That um, if uh I am just a programmer living in a silo, there's no way I can build an email automation tool. I need the domain expertise in marketing and particularly marketing automation to be able to build a product that solves this email automation problem from these methods. Now, if I do not have that domain expertise, then it's paramount that I'm working with team members who have that particular domain expertise, such that the code I'm implementing is coming from an expert perspective. We're building features that are informed by people who understand what domain is all about. That's product development. I'll start seeing how domain expertise is even very important product development in itself, right? Uh, you don't build an insurance product if you do not understand Insurance in the first place. Uh, that same applies to data science; the only thing we are adding here is the math and statistics. So when you combine all of them, that's now the perfect time of data sets.
So it does not matter what career you're from. I think the biggest idea of this Venn diagram mean uh is is to state that regardless of what career you're from, data science could play a role for you because data is essential for almost all fields in the wild today. And by getting a data science skill, it means you can start doing more complex analysis within your particular domain or start improving your breadth of understanding of your particular domain because now you can do more complex analysis that's better than typical traditional research, or you can get more advanced roles within your particular role. It's a good example of being a Trader in a Wall Street Bank for five years, learning data science and then realizing you can start building algorithms to train for people, so you could speak more value to your company based on your domain expertise. So this is not a way for you to run away, right? Uh, this is a way for you to improve more on your domain. Okay, but it could also be a way for you to run away to a different career if you just want to go ahead and do data science in a different purview, right, and the likes, but the domain expertise will still help you to a significant degree. That's probably the thing I stress a lot when I'm talking about the pillars because in a typical program we'll go through computer science, math, and start machine learning pretty well, but we don't really focus on domain expertise because we can't teach people how to be doctors, to be salespeople, and the likes, and that definitely comes from their backgrounds. Um, okay. Um, any questions so far? So good, as I look at the time now. Um, have a good time. That's fine. Um, if you have a question, feel free to drop it on the chat as well. Uh, I'll be looking at the chat uh once in a while. Uh, you don't know what happened to my present. Okay. I hope you guys can still see my screen. That is fine. Um, once again, let me just open the chat one uh once over here. Uh, so no question so far. So good, but if you do have a question, feel free to cut me short. I mean, ask me the question now. Let's talk about the life cycle and careers in data science, and I'll focus more on the life cycle of data science in the first place.
So someone has given me a project, um, and I'm going to use a project that's uh more uh in line with what you have right now. Let me let me let me think about it. Uh, so let's say I have a project such as um uh I I have gotten uh uh trying to use okay, so we're just from an election period. Let me use that one; that's where low hanging fruit in my head right now. Um, someone has been able to give me polling data, right? So we've been able to do a random sample of votes across different constituencies all across the entire nation, right? And now we are supposed to predict who's going to be the presidential election winner at the end of the day. What is the life cycle of this project look like? Awesome. So let's talk about it here. So the first thing is we need to understand the business and problem understanding. Very, very important layer. So when you're trying to predict the president of a place like Kenya here, let's say during elections such that you already have some data coming in or before election where you're entirely based off of polling data, um, we need to understand the problem. And when you talk about understanding the problem, you need the domain expertise to give you the the right assumptions to make in this case. So a good example will be understanding that Kenya votes in blocks is a very fundamental thing to understand so that you make sure as you're sampling out for the entire nation you're able to put into account the different blocks and how this blocks how big this block are, or probably if you're not going to put Kenya into regional blocks, you're probably going to put it into counties, right? And you know like certain counties will have a struggle for different political parties, so you'd want your analysis to be based off of that particular business or problem understanding in the first place. If you do not know the political landscape of Kenya, and then that will not come easy, right? But if you do, then that will be a very fundamental layer of your problem understanding, and you'll be setting up your data collection to understand that you want to collect samples of data per County, right? So that you can be able to make predictions per County, right, extrapolated to let's say the population in a particular County because some counties have a million people, some have a hundred thousand people, so that also matters. So I'm trying to understand as well there's this sort of County variable, so that's also a data point I need to think about before I can even think about making the prediction at the end of the day.
Now, to understanding the problem, we now get data collection and sources. The moment you understand the problem or the business understanding, it's very easy to know what sort of data you need. Now you're just worrying about where you're going to get this data from. Uh, this data could be readily available; probably you just have downloaded. Probably this data is not readily available. Let's say I'm doing an analysis of Jumia. I want to figure out I want to start a business tomorrow. I want to see what are the most on-sale things, like things that go the most on Jumia, right? Uh, I don't know how I'm going to get that. Jumia is a very robust site, but because I'm a program on this front, I know I can scrape Jumia and output an entire database of 10,000 to 100,000 products in their catalog, do an analysis in regards to how many sales they've made, how many reviews they've made, so that I know the products that have the most reviews have the most sales. Probably I'm going to boil down and say these are the top 10 products or other categories I need to be focused on. That's just a good example whereby my data collection I can't download the data, but I need to think about where I'm going to get this data from. So I want to start a business tomorrow. Um, I know Jumia is a place I can scrape to know what sales or rather what things people are buying the most that I can start a business on Jumia tomorrow. So that's a that's a good example. So you you can download data either CSV, JSON format, whatever format it is, uh, or you can go ahead and script this data yourself from the website where you have to code an automated way to get this bulk data because you can do it manually. And thirdly, is you can connect to APIs. So there's people who already have databases of data that they have an open API access that you can access, of course using code as well, right? So that you can be able to access bits of this data. Usually they give you a bit or a segment, of course a certain rate limits based on that particular application. Uh, if I just do analysis of Twitter, a sentiment analysis of how people feel about a particular, let's say about the election problem, how they feel about a particular candidate, probably I'll go ahead and get into the Twitter API, sign up for the API, um, go ahead and collect data, um, which is tweets that have mentioned the name elections or they've mentioned any of my presidential candidates so that I can do sentiment analysis of them. So I'm using APIs for data collection. So there's different ways to do data sourcing and collection. Um, some of them will involve you programming so that you can be able to get them; some of them will be the easier router; you already have that available for you; you download, and then the next layer will now be cleaning and preparation.
How do we make this uh data ready for our models? I feel like someone has a question. If you have a question, feel free to ask me. Uh, if you have a question, put it in the chat as well. So regardless of where I get this data from, it usually does not come in the best format, right? Um, if I'm scraping an entire Jumia website, um, probably someone has pricing errors somewhere; my scraper was not able to put prices in numbers uh at some point. Let's say you realize that I I scrape data; it also went to the Nigerian version of the site, so some of the units are Nigerian uh Naira. Gardens are Kenyan shillings as well, so that's a bit of a problem for me because I want this analysis to be standardized, right, in one particular currency. So the first layer of my cleaning would be to look at that pricing column and figure out what sort of prices are in Kenyan shillings, what sort of them are in dollars, what other ones are in United's. Can I go ahead and convert them to our standard kinds in the likes? Um, what if I have missing data on that particular front? So I'm doing my election analysis; I realized that after doing all this particular data collection, um, Garissa County, I did not get the number of registered voters in Garissa County. What do I do about it, right? Those are questions we need we're trying to ask ourselves. Of course, domain our knowledge helps you answer some of this question. What's your programming fields that allow you to clean this data in the most efficient way possible? And then we go ahead and realize that we have categorical data, uh, and we know that our machine learning models cannot we cannot just feed um categorical data to it, right? We probably need to encode these categories. So if I was trying to classify wine um and beer, and my categories are B and Y, I know very well there's no mathematical way that can interpret B and Y, so probably I need to encode this data zero one, so that when it goes to a machine learning model it's able to output whether it's predicting it to be zero one because the zero one now we can run a mathematical concept on it, let's say we can run a relief function, um, right, um on on the on the features to be able to put a zero or a one, and then that could be our classification. Once again, these are mathematical things that we get to cover when we get to machine learning; go not for this call; I'm just mentioning an example for those of you already in depth into mathematics, uh, for example. Awesome. So we're going to prepare this data set up um to make sure one we don't have missing values; if there's a class on our data, we need to find a way to fix it; um, if we need to encode the data, we need to do it; if we need to standardize the data, which is bringing the mean to zero for those of you guys who are familiar with uh statistics as well, standardization, you also have normalization as well. These are things that we do on that particular front. Now I'll skip data maintenance for now, but I'll go to the next one. As soon as my data is clean, I can now start exploring this data, doing quick analysis. I don't want to run a model or start doing predictions also the data I've not been able to analyze. So on that exploratory data analysis, that's why I start doing a lot of visualizations. I do summaries; what's the mean and the like. So this is typical data analysis, and then once I understand my data well enough, uh, usually you output visualizations; utilize oh, there's a gap here; you need to add more of this data; probably my data does not answer this question so well; you don't know why this you start discovering there's our class you're left out as well. So this is what you're going to do for the outliers and the likes. And as soon as you go through a cycle of data cleaning and exploration, data cleaning and exploration up until your data is perfect, then now we can go to the data modeling. There, this is whereby we start applying machine learning algorithms to be able to make predictions for us. We'll get to this; I want to expand more on this. Uh, this could either be we are classifying the data; probably we are running a regression; probably we just want to cluster and create segments of this particular data, so now we know what to do best with it. That's a good example. And then after we do our modeling, I want to do our evaluation, right? How well is our model performing? What is the accuracy, right? At times you have training data that you want to evaluate to that particular training data before you push your model out to the world and he starts making predictions. So if you already had um house price data and the price of the particular house I was able to model and get a regression line that can predict um the price of a house at any given time, so when I'm modeling this data, I'll take a small percentage of the data I already have and see if my model will predict the prices accurately and to what degree, and this is wherever you start doing mean squared error analysis; we start looking at accuracy scores; if I'm doing a classification on the likes, all all of things we cover when you're going through different machine learning algorithms as well. And then the final layer now is um we've been able to create a really proper model that's able to do a proper prediction. Let's imagine we were working uh in YouTube, right? Um, and in YouTube we are able to create a model that's able to understand the interest of a customer; whenever they go to YouTube, they're able to understand if someone who watched this video then they'll probably be interested in this, and we did a lot of that, and we we did a lot of testing there when you're doing the modeling, and we served this other video to other customers of people leading the same segment and realize that ninety percent of them end up clicking it, right? So after evaluating our model, we were very accurate; like if someone watches a video by Citizen, they're likely to click um NTV, right? Whenever whenever whenever they're looking at it as well. Uh, so now the next layer for us was now because we know the accuracy of our model, we now want to push it live. So I'm working in the YouTube team now, and we now want to push our model live, integrated with YouTube so that it does it automatically, such that next time a user comes in, it's able to understand their preferences and predict the next video for them um using the model that you had already created. So now because you evaluated our model with data that you already have, now we're going to deploy our models that it can start making predictions for new data or not. We'll talk about that later on, right? And the final thing that's very important, uh, whether you're going to be doing the model deployment or whether you're not going to be doing model deployment, at times people just want some data analysis; they want to see the unseen um patterns in data. Data communication is very very important; it is essential for any data analyst, any data scientist to be very very good at communicating because you're not working in a silo; you're working with a team of people, and at the end of the day, after doing all your coding, you've been able to create a Jupyter notebooks; everything seems okay and perfect; the question is people want you to present to them; understand what insights were you able to get off of it; what what decisions are we supposed to make as a company; what are things that we still need to do research; and then the presentation skill and communication is very very important. And when you're done with all of it, right, uh, the moment you're done with the cycle, you always go back to the beginning again. So I've done all this model deployment and the Lights; I've realized oh, the accuracy could improve; oh, this data is not that perfect; um, there's something we could improve over here; uh, we get certain insights off of it that we use these insights to frame the next business problem understanding, and it becomes a continuous life cycle of what it sounds like a project that we keep improving over time, uh, understanding more and growing it over time. A good example of people who work like this are driverless cars, where you see they create a model for driverless cars; they call it version one; they get to the end; they realize that when they are driving these vehicles, uh, it could not detect a zebra crossing, so they go create a second problem understanding; they improve the model to detect uh zebra crossings, and then they go test it; they realize uh when someone jumps on the road accidentally, they might be knocked, right? That kind of stuff, and it becomes a constant life cycle where we do all this collection, cleaning, exploring, modeling; we communicate; we deploy; and then we throw it out there; we identify another problem, and then we keep improving it. Of course, this will be based on whatever project you're working on, so this might be different from time to time.
Now let me talk about data maintenance and model deployment. But certain companies we are working with large sums of data; let me call it Big Data that uh at times we have to automate the cleaning and preparation process because you want to do real-time processing of this data or rather batch processing of the data. So if I created data cleaning, right, and I was able to put it into a pipeline such that next time, let's say let's say I know how phone data comes, so I've been able to do cleaning of previous phone data, but I want for the future phone data as I'm using my phone on a day-to-day basis; at the moment I use my phone again, this data is cleaned and prepared automatically and sub to my model so that I can I can be able to output a prediction, right? So that I can be able to let's say predict the next YouTube video that I'm going to watch. So we needed more real-time, or we need it slightly more automated; that's where the data maintenance layer becomes very very important. This now we're talking about data engineers on this particular front. So while you have to figure out the collection layout, you also need to think about the entire pipeline. Can I automate the collection, right? As soon as I collect the data, can I automate the cleaning and preparation, right? After cleaning and preparing this data, where do I need to maintain the data, right? Should I save it in a data lake, in a warehouse so that other data scientist or machine learning engineers can have access to it and do predictions off of it, right? Or am I supposed to serve it to a place that a model can work on it immediately as well? So we're trying to think about the data maintenance layer of this; this is usually for more real-time applications on that particular front. And then you also have to start thinking about the model deployment so that things can work automatically, right? Right now when I go on YouTube, it's not like someone is taking in my data and then going behind a Jupyter notebook and trying to make a prediction; they already automated that entire pipeline all through, and that's the big value of data engineers. So off of this life cycle, you realize that you don't have to know everything, right? Though it's good to have a good understanding of almost everything, but you have to specialize in certain layers. What I've just talked about on focusing on the pipeline of data collection, data maintenance, and model deployment, that means you could be a data engineer; those are the
Red zones, right? If you just want to become a data analyst, you're mostly good at the data cleaning and preparation. There, the exploratory data analysis and the data communication, so you move the person who analyzes data and outputs insights, right? You don't worry so much about the pipeline of that particular data or how you can automate the collection. Learn the maintenance layer though; at times it's good to have a good understanding of that because not always will you have a data engineering uh, and then the next layer is those of you who decide to focus on machine learning.
Engineers, engineering, which is whereby you already have your data cleaned, or at times you have to clean your data, and then you go ahead and model it and evaluate it. At times you also have to deploy it, but if you have a data engineer on the team, you probably don't have to to uh, to deploy it as well. So these are people who are extremely vast on that um modeling layer, right? You're able to create models that uh have very very high accuracy, right, uh or have very very low errors that they can make as good a prediction whenever they're deployed or pushed out to your data.
And then if you land pretty much everything um that's in-world that's from data collection to data communication, of course you can always keep data maintenance and model deployment because it's really a domain for these engineers. Do I recommend knowing them? You can always be a general data scientist um, and then you also have more specialized careers. So a good example is a quantitative analyst. So for those of you who want to understand this entire stack might want to work mostly in the financial market uh, so you could be a quantitative analyst, right? At times you also have actuaries. So anyone who's an Actuarial scientist, whereby because data is a very significant layer of the typical risk departments in today's world, you also have to understand it stands for different to a significantly degree to get certain actuarial science terms as well. So I'll just keep that in mind as well.
Um, any questions as we now get to the applications and go to the roadmap? Uh, questions from anyone? Uh, once again, uh, yes, yes, Jesse, yeah, I can I can see a question on the uh chat box. I'm not sure if you've covered it before. I think it's Victor. I'm not sure if you're still on the call, um, hope so. Yeah, he he was asking a cool question, what would be more marketable? Yeah, he's here. Would be more marketable: computer science or data science? That's a very very good question, um, and it's very hard to answer that question because I always feel like uh both of these sort of career pathways are in demand in the market today. Uh, I would say uh computer science could have more demand because it's profoundly used across more companies, right? Whereas data science is more in demand in bigger companies because it's a shortage of talent in that particular space. It's a it's still an area that a lot of people do not understand fully, so people are trying to look for experts. So in the corporate space, I'd say uh the sales is becoming more of a need for them. It's very easy to get banks that are saying we don't have enough data scientists, right? Uh, it's not a common degree in a lot of universities as well, so that that becomes a problem to a degree.
However, if we try to look at the sheer number of job opportunities out there, just knowing computer science, being able to do mobile development, software development, almost every company needs it, right? Even someone who's starting an e-commerce business can I need like a web developer to help them do that even before they think about the Sans layer. But then because not many people are training that science, there's a significant need of it. So uh, once again, the answer I can give you, and this is very subjective, uh, probably we can do scraping that's part of the collection layer of job platforms in Kenya to determine which ones more on demand right now, but I'd go ahead and say follow your interest. Follow your passion because when you follow your passion, some of the things you do here will not feel like work. When things do not feel like work to you, you'll end up becoming top 10 percent of one percent in that particular career pathway, and whenever you're top one percent or top 10 percent in any career pathway, it's very hard to uh, to miss out on a job. So just keep that in mind. I would say don't focus so much on what's more in demand uh, even if you want to become a geologist, probably started on demand as computer science, but if that's something that comes easy to you and you feel like it does not feel like electric work to you, you enjoy it, right? Um, it's you are able to obsess over it more, it's very easy to end up being top one percent and you get a job off of it. But these are both in-demand uh layers, right? So if you have the technical skills uh you're going to get the job; you have to be good at it in the first place, but choose one that's more in line with your passion. So if you're a designer on this call and I'd probably say if design is more of your thing, probably attend our HTML CSS Workshop after this uh because that might be better suited for you because now you'll be learning code and how code can be used to design websites and the like, so it's within your domain. If you're more into the mathematical layer of things, data layer, or the type of people to analyze news and say ah, then this could happen here or like this news happened blah blah, so you already have domain expertise, you you love the whole data thing and the likes, right? And the type of people to enjoy Excel sheets, database and the like, data science is perfect for you. So once again, pick it on interest, but both of them are quite on demand or in demand. But if you want to really get a proper answer to it uh probably we can do a scraping all right and then we can see that that could actually be your first instance project if you don't mind. And thank you, thank you awesome.
Okay, so I'll brush through some applications of details uh so so far so good. You've said uh people can use it for business intelligence. So when people just want to know uh how to back uh to make better decisions for their business, I gave the example of Java coffee shops in the beginning. Uh, product development is probably by far one of the most profound use cases of data science; actually these two are the most profound here uh and uh in big tech companies in almost every team of developers working within a feature, there's a data scientist because they want to make sure that uh the product or feature that's being built is quite in line right with the market. So the scientists are there to analyze the data to look at usage uh to see whether it's speaking value to people and whether it's making business sense to that particular company in the first place. So being able to analyze product features and the likes, the very significantly of analysis uh most people in the banks will end up getting jobs in these two areas, so business intelligence which will be based on your domain, so you're probably doing business intelligence in a bank, you're doing it in a supply chain company, onion they like, so the domain could be profound, or you're doing product development and this could be a product that's either a web app, a mobile app, or it could be a typical logistics application that you've built, so it doesn't always have to be the typical software application you're talking about. So uh those are by far most profound because almost all companies need a bit of product development or business types. At times for physical products probably on product development you're doing more of like um uh industrial analysis, so you're trying to look at machines' efficiency and the like, so people who are more into operational management as well might also understand that as well.
For those of you in marketing, data science has become one of the most important levels of targeted marketing. Not only is it being applied in some of the tools you use, so uh how do I get my blog listed on the first page of Google um, right, uh so there's some sort of targeting, they look at that based on such and such terms. How do they choose what article to appear there as well? Um, there's a lot of designs are working on the end goal there and algorithm to be able to predict what article would resonate with someone who searched this the best. But I think uh what applies to most people now is targeted marketing whereby we do a lot of our analysis, we segment people into different interests and you're able to target them and say that on Facebook I want people who are interested in this this and that because you've been able to predict that those are the people most likely to end up liking your project or your product, so it saves you so much on marketing because the more targeted your marketing is uh you don't have to be in front of eyeballs or people who are less likely to buy your goods, so you optimize your your advertisements and promotions better or your marketing initiatives. I've talked about health and medicine and how it can be used for predictions. One that I profoundly love is sports and gaming where a lot of people are doing a lot of people are now applying data science uh particularly teams to analyze players and predict whether they'll be a heat in the future or not. Uh I know in the US there was something called the Moneyball method. Now when we add um this has to the Moneyball method, right? Moneyball is uh I want to talk about Moneyball, our time is not that good over here, but uh if I'm a big club like Arsenal right and I've just come to Kenya right and I'm trying to find the best talent that could be the next Oliech or Mariga uh probably I'll go to a couple of academies, I give them wristbands, try look at movement um of players on the ball in different positions with the likes, and because I have previous data of really good players now a decade went young, I can be able to predict which which sort of two three people out of this 10 000 people should I pick because they have a high chance of becoming the next Superstars right at the end of the day, or when I just analyze the data that clubs collect on performance like the assist mini like heat map of people. Think about the data people will collect; for those of you who are into sports you can imagine when you watch Sky Sports, the analysis they do, there's a lot of data they collect, so it's very important that data is also analyzed to even uh even by people who are into better to predict what sort of uh to go ahead and output what sort of odds they need to put in a game. In fact uh the people who do the best analysis are usually better in companies because they don't want to lose anything, so whenever you see the odds of a team winning is 1.1 and another one is let's say 3.0, keep in mind those people have done extreme deep analysis and they know the likelihood of this guy is winning is this percentage to the T for them to do that. Now of course because the game of luck at the end of the day another team could end up winning, but you can bet if the game was replayed a thousand times right the Betting Company would be right with that particular percentage that they were able to talk about. I think this one we've also seen the people who predicted almost all Euro games um with 90% accuracy; the only games they missed are the ones that that prediction are very close at like 60% this team to win and 40%, those are the only ones they lost. But once again uh betting is also a really good layer to get into.
Transport: we talk a lot about driverless cars; it's not just about driverless cars, it's also about transport optimization. So I know in Kenya we don't have the best traffic light or traffic control system, right? But in other countries that science has been has been used to optimize right um the transport system because we can already predict the number of vehicles that will pass in a particular road at any given time so that we know how much to improve the flow of that particular road, or even we know how when to increase the number of lanes or expand a particular road because this is a typical data collection of a target.
Finance: very very understandable risk and fraud; for those who are into it I don't need to explain more on how that can be used as well, but we can use data science to detect fraud automatically without necessarily having to go through a million transactions per second to know what could be the fraudulent ones; we can just know that as well now automatically. Speech recognition: these are now the more niche areas. Image recognition, right? When you're starting to think of AI as well um, even natural language processing: how can I be able to analyze a paragraph featured by someone and say that it has a positive sentiment or a negative sentiment to other particular thought as well? So these are more of the AI focus layers that come right after it.
Um, any questions as we now get to the roadmap?
In the next five minutes I think we'll just keep this up, and I hope no one on this Workshop uh is looking to get to that, but if there's anyone in this Workshop who was also joining the HTML CSS Workshop uh just bear with me, we'll finish in the next 10 minutes. Um uh so you can always feel free to make the switch uh but yeah the invite should not be in uh the the invite should already be on your email for those those of you who uh who signed up for both data science and software developer, just keep that in mind. Now let's go to the roadblock now, so how should one typically land data science? So the first layer will be the data analysis layer, so what should I learn from the data analysis layer? So first layer is going to be learning Python programming. Python is by far the most recommended language here, uh programming language, yeah, R is also used more in academia, let's keep that in mind as well, but the tooling that Python has been given to be able to do data analysis and machine learning is quite profound today that it's by far the most used language by data scientists across different industries. So when you want to get started, the computer science layer, you can learn Python or R, but Python is the recommended language in the structure. The second layer you want to become a data analyst is understanding um SQL data analysis. What is SQL? SQL is a querying language that is applied on relational databases. Let me explain better: most companies today are using a relational database which is SQL based to store their data. When you go to the website of the indoor school at the moment, right? When you interact with that website, the content on that website is stored on MySQL, let's say database, whether whether you signed up onto a learning management system, you finished course X or less on X and the likes, it's all in a, you know, MySQL database. So if we wanted a data scientist in the team, the first layer of it would be an apt understanding of SQL because the data collection layer needs to come from that particular database, and because a lot of companies use SQL databases or relational databases, it is essential to understand um SQL, the querying language, and how to be able to do analysis of the SQL in the first place. For companies that use NoSQL databases, you should already know how to import JSON data into your Python um sort of um uh into like into Python extremely easily, which is really in the form of JSON, which usually explore when you're doing Python, not not not a way because even when you consume APIs at times the data comes in as JSON, but the SQL layer is very important to understand the query language for SQL, whereas for NoSQL databases you should—that's a skill you get as you're doing Python programming.
Now once we are very good at the programming layer, we're very good at the SQL, well now we can do Python data analysis. What does this entail? The first layer is being able to do data sourcing with Python. Now while we can go ahead and download data most of the time, at times we have to consume data from an API, so how can you access APIs through Python? Second is we have to do web scraping. Whenever you want to get data from the web, unless I just want to collect data from all a million Jumia products, of course I'm not going to go page by page and write down this data; I can create a script to go automatically and get that data for me, so web scripting is a really really a good example of this. And the final layer will now be how do you integrate your SQL databases to Python, right? So MySQL queries are here, I've been able to approve the sort of results, now how do I bring this sort of database outputs um into into the Python environment? Okay, now once we have sourced our data, the next layer is now data analysis, and for data analysis there's two libraries we use in Python, that's NumPy and Pandas. However, it's highly recommended to also think of SciPy because it's a core bonus that covers some things; a good example is something like outliers or whenever we have imbalanced data, uh SciPy is the better library for it as well. So we do explore some things on the SciPy uh side to a significant degree as well. So NumPy, Pandas, SciPy will be some of the libraries you learn to be able to do this analysis with Python, and then the final thing now is we've been able to analyze our data, right? Pandas, this like the analysis like you've been able to analyze it as if you're in Excel, now we want to visualize it, right? So how do we output graphs and the likes, and we have libraries for this, so some of the ones we recommend would be things such as Matplotlib, Seaborn; those are probably the two most core ones. If you want to be able to do geo data such that we can put data in a map, then GeoPandas is highly recommended for you as well. So that's that's something I do it as a bonus. You can also learn Plotly and Bokeh. Once again, we've underlined uh the most important layers of things you should be landing over here. Now not every time will you be doing things in Python, and you need to understand because companies work with tools, right? You don't always have to overcomplicate things and have to code up your visualizations; at times, at times a company already is using Power BI or Tableau, and these are probably by far the most common data visualization tools, and any data scientist you need to understand how to use these tools in the first place because not always will you be required to go ahead and do a visualization in Python, though you do at times. Just the easier route is to go ahead and use a data visualization tool like Power BI and do it for you or Tableau and do it for you automatically because everything is all structured for you. So if you have a tool that's going to ease in your work, go for it and do it. In the case where this tool cannot serve you, then now your programming skills apply which can allow you to make more custom visualizations or make a slightly more advanced analysis or way more customized, let's say data manipulation in that particular front. But because a lot of companies use this when I leave, you want to get into business intelligence and uh product uh development analysis, so those of you who are going to work in typical defense jobs in companies as well, Power BI and Tableau things you cannot run away from. I'd say learn one of the two; Power BI is starting to become way more popular because of the impact of Microsoft across different companies, Microsoft Teams and and the final one is now Microsoft Excel, right? So Microsoft Excel is probably the the bread and butter of any business uh and because the bread and butter of every business and almost all businesses have their data in Microsoft Excel to a degree, there's no way you can do without it as well, so keep that in mind. Uh don't run away from Excel; in fact learn it to a significant degree that you can even use macros, you can even do uh what do you call it uh uh there's uh pivot tables and the like, so it's not just addition and subtraction of this particular front, it's running Excel to...
A degree that you can do proper visualization, proper slicing, proper data analysis, right? Because in certain companies, you'll end up doing a lot of edit analysis within Excel and not even ever touch Python. Probably will end up doing the analysis where you just start SQL and Excel, right? So you also need to make sure that you don't ignore the skills that you think are like, ah, you're only Excel, it's okay. No, no, um, it should be standard for everyone. Okay? If you have a question, just drop it on the chat. I'll be looking at it.
Uh, Stephanie is asking, uh, "First, I'd like to apologize if it is out of context, but my question is in terms of career growth, market need, and opportunities, what would you advise a person to choose between data science and software engineering?" I think I answered that question, right? Okay, I did answer that question. Yes. What is an API? I'm having trouble understanding what is an API. Okay, awesome. Yes, uh, okay. So an API, uh, in full is called an application programming interface, right? Uh, usually it's a way for you to interact with data from a particular application. So let me just give you an example. Um, if today I'm the creator of YouTube, right, um, uh, and on YouTube, I have a database of videos, I have a database of users, and I have a database of how users interact with these particular videos. Most of the time, um, I want to make my data available to other people, right, such as yourself. Probably you want to do YouTube analysis of your own account, right? Or probably you want uh YouTube to integrate with a different application. So a good example is YouTube integrates with a couple of YouTube analysis applications that do like, um, uh, let's say listings of of uh the best YouTube channels, right? Or they're able to also do like search engine, like let's say I want to do search engine optimization on YouTube, right? So I'm building a tool that's able to look at YouTube SEO to see that when someone searches this search term, what are the top five videos as well, so that we also know what these top five videos entail. But because you want people to build and integrate with your tool as much as possible, so they want to access some of the data in your tool or they want to post, right, uh, so they either want to write, which is like the post layer, or get data from your particular database, and you want this as a possibility, but you cannot give them database access. Right? We cannot add them to our database and say, "Um, here's access to my database." You can't do that. The way we do that is we do that using APIs, whereby we create an API. So a good example, it could be something like a REST API. So a REST API is one one example of a particular application programming interface such that using the REST API endpoints I've given you, you can get some of the data points that you want, right? So when you Google developers.youtube or something of Google developers and you look for YouTube, you will see API endpoints for YouTube such that you wanted to access, right, uh, the top five videos on this search term, I can be able to get it. And this allows people to build and integrate with YouTube because our applications don't live in a silo; they have to interact with a couple more applications. The same way when you have to log in into different applications with Facebook, the reason why you can log in into different applications with Facebook is because there's an active API connecting Facebook to that particular application such that when you log in with Facebook, um, the API integration is able to push in your email, right, to the application that you've logged in to, to enable it to also detect that you've logged in without necessarily you having to input your email again. So you can see like we're not living in a silo; that's a good example of it. So through an API, you can either post data onto a database, right? Um, you can either get data from a database automatically without you necessarily having access to that particular database. And through an API, a company is also able to limit how much access someone has uh to their particular database or their particular content, right? Uh, so at times they might restrict it with an auth, so you have to authenticate yourself. At times they have rate limits on the API, right? So that's the other example. A lot of times they've just defined certain endpoints which can only give you certain data points so that you have access to not the full database but areas of that particular database. So that's what I'd call an API, very, very essential because that's the best way to interact with some of these tools because you'll never ever get a database. I hope I've answered your question. Yeah, thank you. Once again, when you're learning APIs, you'll get a more in-depth analysis of understanding of it when you're actually doing it. I think it becomes more profound when you're doing it.
Okay, so for machine learning, the first layer is you need to understand the prerequisites of machine learning. Now, most data science programs, and I'm talking about the data science programs outside a typical degree program, will not go through these prerequisites because it's extremely difficult to go through a six-month program that starts doing mathematics, right, and the likes, right? It's it's extremely difficult. Um, so what what do we do on this particular front? It's recommended that you either have a background in mathematics and statistics, right? So that you can be able to do machine learning to a very, very good degree. And if you don't, it's recommended that you can learn these things by yourself. I know it's not easy to learn math outside a typical school setup, but it's not something you can't learn. In fact, this is probably the most free content you can get out there, whether you're doing 3Blue1Brown or 3Brown1Blue, something like that. So there's a lot of sites, right? Um, I even see Imperial College has a lot of these things available to the world as well. But mathematics and statistics, that is a prerequisite. So what type of math are we talking about? So definitely differential calculus will be at the core of it, linear algebra because we'll be working with vectors and matrices when you're building some of this, optimizing some of these ML algorithms, and then the other one will be convex and concave optimizations as well. Uh, that that's more of the advanced math as you start thinking about things such as gradient descent, which is a very key concept of machine learning. None of the statistics, of course, the descriptive, usually almost everyone understands it. You also need to understand inferential statistics as well as probabilistic statistics. I think for those of you in the statistical field, you should be able to understand this. So things such as Bayes' theorem there and the likes. Um, do you have a proper understanding of Z-scores and the like? So, uh, so statistics are very key concepts because some of these concepts are applied, right, or almost automatically in some of the machine language. I won't talk about that, but let me go straight up to the machine learning, yeah.
So in your learning machine learning, the first thing you learn is supervised learning. So what does supervised learning entail? Supervised learning is machine learning whereby we are training our model based off of data that has both the feature and the label. Let me explain better. Let's imagine I have a data set of house features, but in that data set it already contains the price of the household. So I have the features of the house, which could be it has three bedrooms, four bathrooms, um, then you need the tiling is wooden and the like. So I've written all those things, this is, and I've gotten the price of each and every single house, right? Now I can go ahead and model my data using this so that you can be able to detect a pattern between the features and the price so that I can get a regression line such that the next time I only have the features of the house, it can predict the label, which in this case is the house price. Uh, some of the typical supervised learning algorithms whereby we train with both features and labels that it can make a future prediction of the label whenever we don't have it in the future is classification, where we are classifying things, or regression, where we are predicting more of continuous data. So regression would work for things such as pricing; classification will work for things such as I'm trying to determine if this drink is beer or wine, and these are the features of it. So a good example, and the tool we use mostly for this will be scikit-learn, um, for it. Okay, any questions on um supervised learning there? Okay, no question. Now to understand supervised learning better, it's good to understand unsupervised learning. Unsupervised learning is that sort of machine learning layer where we are trying to make predictions here when we do not have—we're training our data without the label, essentially. There's no label. Imagine I've just been given a house with features; there was no price, right? What what can I do with that? Imagine I've just been given a couple of customer profiles, right, uh, I right, there's that just a customer, a batch of customer profiles, and how they interacted with the with the tool over time. So for unsupervised learning, we are trying to optimize for certain for different things. At times we are either just trying to cluster out these people. So let's say an example, if someone gave me data in a particular supermarket, right, um, I could easily just do an unsupervised learning of it and then go ahead and realize that every time someone bought milk they also bought bread, then I can tell the supermarket for you to optimize your sales—put the bread at the door of the supermarket and then the milk at the exit of the supermarket. So bread at the entry, milk at the at the exit because people who buy milk always have to buy bread. By them picking this bread before they get to the milk, they'd have gone through the entire supermarket so that their eyes can see a lot of products and so that they can make some other random sale. If you want to optimize ourselves, so I can only understand that he was able to do clustering where there's no particular target labeled, but we're able to start detecting and seeing patterns in data and start getting insights off of it. That's a good example of what you call association basket, right? That's used in supermarkets, and they use that to optimize the arrangement of goods to optimize sales in the supermarket. The other thing would be something like dimensionality reduction. So I have a thousand columns of data, right? Probably these are a thousand columns of data, but my uh my my rows are going to be about a million, which is fine, but then I have a constraint: it's computing power that I need to use over here to make some of this sort of predictions, uh, or rather to run a regression from this. It's going to take me a really long while; it's going to consume a lot of money for me to run it on the cloud. But it's something like dimensionality reduction; I can easily run a model to detect how much does each column speak to the variance of the data, like how how how much additional value does each column speak to this particular data set, right? You'll always realize that at times, um, some some some some columns are quite correlated that uh by by like you don't need to have both of them; you can just combine them into one or have just one of them as well. So by doing dimensionality reduction, there's no clear label over here, but I can go ahead and output how much variability each and every single column adds into my data. And then when I go ahead and realize that 95% of the variability of this data set uh is uh is spoken by only 50 columns, then I don't need that thousand different features; all you need are these 50 features. So this is also a way for us to optimize our algorithms as well, and you can see there's no label as well, right? You just you're just able to see an unseen pattern. So that same pattern is: viability of this a thousand-feature data set is only contained in 50 features. So I'll maintain that million rows because I need more data, but I'll only use 50 columns instead of a thousand columns, and that saves me a lot of cost, right, and time when it comes to modeling. Cool. After that, it's hard to avoid time series data. So time series analysis with a library called Statsmodels is highly recommended as well because most of the data we work with, if you're working in sales, you're working in finance as well; it's quite um time-bound as well. And then you have things such as model evaluation, model deployment, very, very important to learn model evaluation, model deployment. How do you know whether your model is working well or not, whether it's accurate or not, results? So being able to evaluate a model is a very, very core skill to understand. And then for those of you who want to focus on machine learning engineering or even data engineering as well, model deployment is a very important here. So how do I push this model to a web app so that other people can be able to make predictions off of it? Okay, any questions um on machine learning? Okay, no questions. Now let's go to the final thing, which are now the advanced concepts. So if you just do step A, you can easily become a data analyst. If you do step A, step B, now you can start applying for data science jobs as well, but you have a chance to go ahead and specialize. So one area you can specialize is on Big Data engineering, where you focus more on databases, you focus more on building Big Data pipelines to do real-time or batch processing of data. And the good example is the YouTube case. Let me give an example of the Uber case: if I ask for an Uber ride right now, um, the moment I step onto the app, it detects my location, right? So it's taking in data real-time based off of my location; it goes and tries to map out the closest Ubers to me before it matches me to someone. So this this is now real-time optimization of a particular application whereby it doesn't matter how good you are at machine learning; like without the sort of layer of making it come real-time, probably someone built the model to match users to closest drivers, right? But how do we push it to production in a way that it's more real-time? So so you'll end up learning more about data engineering here; you learn a lot about databases because you'll be working a lot on—on it; you learn a lot about the tools that you use on how to optimize these data pipelines. Some of them include Apache Hadoop, Apache Spark, or for real-time data streaming, Apache Kafka as well. The second layer will be deep learning with TensorFlow. So for those of you who want to get even deeper into machine learning, right? So for Big Data engineers, you enjoy more of the data pipelines; for those of you who want to do even more advanced modeling, you can go ahead and start learning AI now, which is a deep learning layer. The library is recommended for this will be TensorFlow, PyTorch. I'm starting to have a really good um uh let's say preference of PyTorch, right? I've always been a TensorFlow person for the longest; I've got—PyTorch feels slightly more popular in the market right now. Um, so some of the things you learn here will be things such as neural networks as well, right? Um, how can you start doing image classification and recognition, um, right? Um, how can you start doing natural language processing, right? So uh generative models as well by—how can we start getting um computers to generate real images, you know? Right now we're we're seeing we're seeing tools such as DALL-E whereby they're replacing designers. We tell it to design uh an image of a submarine with uh with with with uh with a rhino on red, and this is a Nancy image, and it's able to generate it, and it looks very realistic as if as if it was an actual photo as well. Those are some of the generative models; a really good concept to understand. I think that's probably one of my most interesting AI layers that I I spend a lot of time on trying to just uh just feel amazed at what we can do at this particular moment. To generative models or reinforcement learning, which we're now starting to borderline on the robotics layer or even when you're trying to train computers to start playing games, a good example like with playing chess against each other and the likes, it's quite interesting reinforcement learning there. So I don't want to delve deep into the explanation of this because of the time right now, but you can always go ahead and explore this once you have an understanding of machine learning if that's where you want to kind of expand onto. Um, though getting deep learning jobs, uh, most of the time you need to have a graduate degree, so that's why you don't see it being covered a lot in boot camps, but you can always like if you work a lot in a machine learning role, it would be very easy to upscale into this because you'd understand the layers; you'd already have interacted with some of this sort of let's say neural networks as well to be able to to make that sort of prediction. But most importantly, this is not even an advanced concept; it's a concept that everyone should learn, whether you're a software developer or a data scientist: cloud computing. At the end of the day, we are deploying all these things onto the cloud, and the three most popular cloud platforms are AWS, Google Cloud, Microsoft Azure. Regardless of what development you do, you need to know how you can push it to the cloud. So if you are a machine learning engineer, how can we deploy this to the cloud? If our data engineer, how is my data pipeline, my Big Data pipeline, how is it working automatically on the cloud as well, whatever platform it is, or even if you're a software developer, how are you making sure your web application is on the cloud? At some point, companies are using this; there's no way to avoid it; it's the way to go for it.
Okay, so let's call it a day now, with a final question. Uh, any questions so far? So good on the pipeline? So I'm going to finish it up. So I know I've thrown a lot of—juggled at you today, and I don't want to leave you out for that, and that's why we're going to go ahead and talk a bit about in your school. Um, so why we do this is because we are a fast, part-time, virtual coding school, right? And we offer programs in data science and web development. Um, the thing that sets us apart is: one, we champion personalized learning. So you're not going to be learning through workshops like this in a factory; we only do workshops like this for free courses. But you get to work one-on-one with a technical event. The second layer is project-based learning, whereby you get to work on a daily challenge, a weekly project, and a custom project. So you can imagine how much projects you have built um in either route, right? If you are part of the Zimdua ecosystem, right? Apparently the weekend Capstone projects tend to be the more weighty projects that are able to build your skills to a particular degree. Keep in mind, in this industry, we are on Parallel Tech. As much as people are saying they look at degrees, most companies are impressed more at what you can build more than what degree you've got. In fact, you could have the best degree whatsoever, but if your building skills are off, there's no way you're getting a chance at a solid job because it's very easy—the good thing about such such fields is there's no gatekeeping based on "I know this person or this other person"; it's all about your skills. If I look at the product you've built, your exchange, and the likes, we're able to see whether you're going to be a fit for the team or not. That's a good thing about this, both software development and data science. 23. And the final thing is our programs are based off of industry-standard certifications by Microsoft Azure and AWS, so students sit for these exams when they graduate um right after. So let's talk about it. Um, the most important layer for us is we believe that income should not be a barrier to education, and that's why we offer the best fees in the coding market in Kenya at the moment. So the first option people usually have is you can either pay for the program upfront or in installments. Uh, so because you break down the program into five-week modules, right, you can pay in five installments, right? Or you can pay it all.
Upfront and save 30% of the cost. Uh, once again, this will share with you the—also on our website as well—um, but for students who cannot define it, we have something called an income share agreement that allows you to learn now and pay later. So you can defer up to 50% of your fees if you cannot afford it. For an income share agreement, that's option two there. However, for people who pay 100% of the program, um, because in an income share agreement you pay nothing if you don't get a job, so uh, we also try to give the same advantage to the people who paid upfront, which means you have a 50% money-back guarantee if you do not get a job uh in six to 12 months after graduating as well.
Once again, it's pretty much based on how much you're looking for a job. Uh, some people we work with, because it's part-time, are probably still trying to finish school as well, so it's going to be valuable. But we go ahead and say that um, after finishing the program and you've done the certification and it's been 12 months, the pattern job will give you 50% of your money back offer, VTP, paying upfront. If you are part of an ISA, if you don't have a job, then you're paying—you're not paying an ISA because it's pretty much better. Okay. I've added this link to this sort of slides that you can always follow in which is: go to the website and learn more. You can learn about these financing options and the terms that go off of it, and you can also explore the breakdown with it. Science Program, which should be pretty similar to what you've gone through today. Uh, pretty similar to what you've gone through uh in the call today.
But I have some good news for all of you. Um, probably you guys already on it. Um, you don't have to make a commitment already if this is something you're still exploring. You have a free Python credit science foundation's course available on this link: theindoorschool.thinkyfick.com that you can join in anytime, start learning a bit of Python programming and decide whether you want to expand it to that sense or not. Uh, and this has no strings attached. You have to pay nothing; you're not forced to join the core programmer for it. Uh, this course is actually, in fact, a prerequisite to the core program, so you can't join the core program if you've not done the free course in the first place. So you have it out there for you as well, such that those of you who are still new to this and are still thinking about it, um, you can explore it. For those of you who are not new to this and want to join the core program, uh, I'm available as well, so you can talk to me after this particular call and you can see the way forward. And I'll share the slides with you um so that you guys can be able to apply if you want to join the core program, or you can join the free course if you want to.
Also, in fact, to make life easy, I'll just copy-paste this link as well on the chat so that you guys can be able to join in. And for those who have not—so if you want to join the free course, that is the place to join the free course. So that's the link I've shared right. Uh, if you want to learn more about the data science program, the coded science program and apply, um, this is the link to to our website page that's focused on data science. So for those of you who want to go straight to the core program, however, because you're going to share all these resources and we have an active Slack community, I recommend all of you to join Slack as much as you can. So let me just add the Slack link as the final thing: join Slack. I'm going to send these slides on Slack actually um as well. And on Slack you can also DM me at any given time if you have questions and if you need um support off of it. So my name is Cyril Machine on Slack. You can always DM me over there. When you join Slack, make sure to join the #forum data science. Okay, so I'm going to share this slide, and that's where we're going to keep having discussions on data science. Um, whether or not you want to join the program, you just keep that as an open forum for everyone um who wants to chat, exercise.
Now, with that said, let me take questions um from the from the from the crowd. Yes. Um, how can one reach you uh maybe after like, do you have an email address like maybe a professional one? Yes, yes, yes. So personally, you can reach me at uh cyril@thedoorschool.com. Um, sorry, uh so I've placed my email over there uh school.com. However, um if you just want to reach out to any one of the team members just for general support, reach out to me for data science related issues. They're not so sure that you want to take this or whatever, but if you have general support you can just do hello@thenorthschool.com straight up to the person who's um a better place to answer you. So that's what I would say in regards to emails. And these are also very, very available on our website. We also have a phone number on our website just to make it easier, so you can always WhatsApp messages as well. All right, thank you. Awesome. Yes, Eunice, you're saying something. I have a question. Uh, yes. All right. Okay, Umi, I can't hear you so well, but I heard you say, "Thank you." So welcome very much. Thank you for attending as well. I, I hope this talk was valuable to you guys, um and I hope for those who are beginners the free course will be a good starting point. I don't know that you have anything else to add because I can't hear you as clearly. Just okay. Uh, I don't know that the problem is my connection or Umi, but you guys can tell me. Um, but what I'd go ahead and say is uh um if you can just type down on the chat so that I can understand exactly what you're saying, that should be good. Or you can stay behind after this this call as well so that we can we can chat if it's more possible. But if you can type it down if it's going to help everyone, please do, because I can't tell you so clearly um as you're waiting for that. Any other question? Okay, no other questions. Uh, thank you very much guys for attending this. I hope this was valuable to you guys, and I hope that I'll see some of you on our core data science program, or I'll see some of you uh learning Python programming on the free course as well. Whatever thing. Um, even if you just want to go ahead and get started online by looking for resources, you now have the curriculum to follow. So if you want to learn by yourself, well and good. You want to learn in a community with the benefits you've talked about over here, you can join us. Um, if you just want to see whether this is for you, you can join the free course. The links are there. Please join the Slack channel because that's going to be the easiest way to reach me. Um, right, uh, because I'm usually on quicktail over there, so you can always message me and get my response on a particular thing or get quick support from me or get forwarded to someone else because everyone's there, everyone in the community is there, from students, alumni to team members. So it's a really good uh portal to be on as well. So thank you very much guys. Enjoy the rest of your night. Um, if you are part of the people who are joining the HTML CSS Workshop, so you have taken 30 minutes of the workshop, but I can still feel free to go during it. Um, I will—I think the link should be on your calendar if you're here, uh, if you registered. If you don't have the link, just click behind and I'll send you—I'll add it to the chat. Um, thank you very much. Have a nice time. I will send a follow-up email with all this stuff, but the best way to get the follow-up is on Slack, which I'm posting immediately. I, when I leave this call, I enjoy the rest of your night.